All articles

Process Control

Tenderness Measurement and Our Endpoint Correlation Data

Shin Hae-won

A prediction model for beef aging endpoints is only as good as the ground truth you use to train and validate it. For us, the ground truth is Warner-Bratzler shear force measurement. This article describes the measurement protocol we use, what the correlation data from our first six months of pilot operation looks like, and where the predictive model is accurate versus where it still struggles.

I want to be direct about scope: our pilot operation is small, our dataset is not a published study, and the numbers I will cite are from our internal work rather than a peer-reviewed paper. This is an honest account of what we have found building something real, not a claim about statistical significance across a large population.

Why Warner-Bratzler shear force

Warner-Bratzler shear force (WBSF) is the standard mechanical measurement for beef tenderness in meat science research. It measures the peak force required to shear a core sample of cooked beef at a standardized geometry, typically reported in Newtons or kilograms-force. Lower values indicate more tender product.

WBSF is not a perfect proxy for eating-quality tenderness. The relationship between WBSF and consumer panel tenderness ratings is strong but not linear across the full range, and the cooking method, core diameter, and resting time all affect the result. Despite these limitations, WBSF is the most consistent and reproducible instrumental measure available to us outside a full sensory panel, and it has the most published reference data to compare against.

We use a standardized cooking protocol, a uniform core sample diameter, and multiple cores per cut, averaging across positions to reduce within-cut variance. The cuts we test are separate from the product going to the operation, pulled from the same batch at the same aging duration and aged under identical conditions.

The structure of our pilot dataset

Over the first six months of operating our pilot system on controlled aging runs in our Seongnam facility, we collected WBSF data across a range of aging durations, cut types, and chamber conditions. The sample sizes are small relative to academic beef science studies, but they span enough variation to start identifying where our model performs and where it does not.

We run parallel batches: one batch with our sensor network active and our model predicting endpoint, alongside a conventional control batch aged to a fixed calendar duration based on the operator's standard practice. At the end of the run, we pull WBSF measurements from both batches. The question we are trying to answer is whether the batch terminated at our model's endpoint prediction has WBSF values closer to the operator's target than the batch terminated at the fixed calendar duration.

The short answer is: yes, in most cases, and by a margin that matters. The longer answer requires understanding where the improvement comes from and where it does not.

Where sensor-guided endpoint outperforms calendar aging

Calendar aging fails at the extremes. When a batch encounters abnormal chamber conditions, whether a warm event, a humidity excursion, or a cold chain history with unusual temperature variance, the fixed-duration endpoint becomes wrong. The batch has either over-aged or under-aged relative to target tenderness, and the fixed calendar cannot see that.

Our sensor-guided approach catches these deviations because we are integrating cumulative enzyme activity across the actual temperature and humidity history, not the nominal aging duration. A batch that experienced a two-day period of higher-than-target temperature may reach the target WBSF threshold 18 to 24 hours ahead of the calendar endpoint. Our model flags this. The calendar does not.

In our pilot runs, the batches where calendar and sensor-guided predictions agreed most closely were also the batches where WBSF differences between the two approaches were smallest, which is the expected and reassuring result. The batches where the approaches diverged most were precisely the batches with the largest WBSF variance in the calendar-controlled group.

Where the model still struggles

The model's weakest area is high-variance batches with unusual starting conditions. If the cold chain history of the incoming product is not logged, we are missing a significant portion of the cumulative enzyme activity that occurred before the product entered our monitoring system. We currently prompt operators to provide slaughter date and any known transport temperature data, and we use this to initialize the model's starting enzyme activity estimate. When that upstream data is incomplete, the model's endpoint prediction is less reliable.

The model also struggles with mixed batches, where cuts from different animals are aged together in the same chamber. The between-animal variance in enzyme activity rate is real, and a chamber-average prediction will be accurate for the average cut but will over-predict tenderness for slower-tenderizing animals and under-predict for faster ones. In a mixed batch, 15 to 20 percent of cuts may be outside the target WBSF range even when the batch-average prediction is correct.

We do not have a solution to within-batch animal variance yet. The honest position is that the model operates at the chamber level, not the individual-cut level. Per-cut instrumentation would require a different sensor architecture, and we have not built that. For premium single-source programs where all product comes from controlled provenance, this matters less than for mixed batches.

Correlation values and what they mean practically

Across our pilot runs, the correlation between our model's predicted endpoint (expressed as estimated cumulative enzyme activity progress toward target texture) and the measured WBSF value is in the range where we can make useful predictions but not precise ones. The model is better at identifying when a batch will clearly not be done yet versus when it is close to target than it is at calling the exact moment of optimal WBSF.

What this means operationally is that the model is most valuable as an early warning and extended-aging flag, not as a precise endpoint timestamp. If the model says the batch is on track to hit the target window between day 16 and day 18, an operator can plan sampling and quality assessment for that window rather than working from a fixed day 21 calendar. If the model shows the batch is running behind due to a cold event in week one, the operator can extend the program with visibility into why, rather than discovering the texture shortfall at end of the fixed aging period.

That is the practical capability we have right now. We will keep publishing as the dataset grows and as the model improves. The number that matters most to us is not the correlation coefficient in isolation, but whether operations using sensor guidance are pulling consistently better-texture product than those using calendar aging alone. On that question the pilot data is encouraging, and we are building toward the answer with more careful measurement at every run.

See it in your operation

Request an early-access demo

We are working with a small group of premium beef processors. Tell us about your setup and we will show you what SaltbyPep looks like on your specific aging operation.

Request Demo

More from the blog