A controller emails Ray a number at 4:40 on a Thursday: setup time on the press line is up 22 percent. Ray forwards it to sales without reading past the subject line. By Monday morning, three quotes have been re-priced off that one email, and a fourth is sitting in draft, waiting on his signature.
Ray is a composite of shops we have looked at, not a named client.
Nobody rebuilt the number. Nobody asked where it came from. It just moved through the building like it was already true.
Here is what a second source would have found. The 22 percent can be real without being setup time. A field tech reclassifies two long-changeover jobs as "run time" the week before, trying to make his shift's numbers look cleaner. The ERP (the system that runs your orders, jobs, and inventory) takes his entry at face value. Nobody cross-checks it against the MES (the shop-floor system that logs what the machines actually did, timestamp by timestamp). One typing mistake, dressed up as a trend, re-prices four jobs before lunch on Monday.
That's the whole problem with plant data. A single report is a story someone is telling you. It might be true. But nothing inside the report itself can tell you whether it is.
TL;DR
- One source is a story you are telling yourself. Two independent sources that agree are evidence.
- Independent means different writers. Two reports off the same table are one source wearing two hats.
- Pick your tolerance before you look at the answer. Deciding after you see the gap is not a test, it is a negotiation.
- A blank is not a zero. Label unmeasured lanes UNMEASURED so they never get averaged into a decision.
- Outlier trimming is a decision. Write down what you cut and why, or you cannot defend the number later.
The trap: two dashboards that are secretly one source
Ray's mistake isn't trusting the ERP. It is trusting a *number*, when what he needed was a *comparison*.
Here's the trap most plants fall into without noticing. Two dashboards can look independent and still be the same source wearing a different skin. They're one source if any of these is true:
- Both pull from the same job-history table.
- One is a saved view of the other.
- A person keyed both from the same nightly export.
That last case is one source with two chances to make the same typo twice.
Real independence usually means one of these:
- Different capture moment. A machine timestamp versus a hand-entered start time.
- Different writer. The scheduler wrote one, the operator wrote the other.
- Different unit entirely. Hours from the labor ledger versus pieces from the scrap and yield log.
- Different system. ERP job records versus the time-clock or badge system.
The best pairs feel a little awkward to compare. That's the point, because awkwardness means the two numbers had no shared path to be wrong the same way.
What the second source found
Ray is a composite of shops we have looked at, not a named client. In this illustration, when the plant runs the second source against that 22 percent, the gap is ugly and instructive at the same time.
**Source one:** total labor hours booked to the job, straight from the ERP job record, the number the field tech touched.
**Source two:** the gap between the previous job's last good-part timestamp and this job's first good-part timestamp, pulled straight off the machine controller. No human typed it.
The machine says something different than the ledger. Two of the four "22 percent slower" jobs have setup times that haven't moved at all, because the field tech shifted minutes from setup into run time to make his changeover numbers look better on paper. One job really did get slower, by a tooling issue that had nothing to do with setup at all. The fourth is a wash.
Three of four quotes get re-priced off a number that is wrong in three different directions at once. That's not a rounding error. It's a coin flip dressed up as data.
Pick the tolerance before you look
This is the step Ray skips, and skipping it is what lets a typo travel four quotes deep before anyone catches it.
Decide up front how close two sources have to land. Write it down before you run the second query, not after you've already seen the gap and started rationalizing it.
Tie the tolerance to the decision, not to your comfort. These are our recommended starting points, not observed industry norms:
- Changing a quoted price: the two sources should land within about 5 percent.
- Ranking jobs worst to best: start at 15 percent, because the ranking holds even with the gap.
- Buying a machine: tighter than either, and you get a third source.
If you decide the tolerance after you see the answer, you're not verifying anything. You're looking for permission to believe what you already forwarded to sales.
When the two sources disagree, you already won
Most people treat a mismatch as a failure. It's the opposite. A mismatch is the only free information in the whole exercise, and it's what would have saved Ray's Monday.
Work it in this order:
1. **Same population?** Same date range, same lines, same shifts, same part numbers. Most gaps die here. 2. **Same definition?** Does "hours" mean clock hours or paid hours? Does a job spanning midnight land on one day or two? 3. **Same filter?** Did one version silently drop cancelled or reworked jobs? 4. **Real difference?** Only now is the gap telling you something about the plant instead of the query.
Step four is where the money is. A real gap between a hand-entered field and a machine timestamp usually means somebody is rounding, and rounding always leans the same direction, toward whoever's numbers get reviewed.
Zero and UNMEASURED are different numbers
This is the blind spot we run into most often.
A **zero** means you measured and the answer was nothing. Zero scrap on that run, zero downtime that shift. An **UNMEASURED** lane means nobody recorded it. The field is blank because the process never wrote to it, not because the value was nothing.
Average them together and you get a number that's confidently wrong. Downtime looks better than it is. Scrap looks better than it is. And the lane with the worst problem is usually the one most likely to have no data, because nobody was watching it closely enough to log it.
Our rule is simple, and it's the whole ANVIL method in one line: **Actual Numbers, Verified, In the Ledger.** If the data can't support a claim, we mark the lane UNMEASURED and say so out loud. A method that never refuses is a sales deck, not a measurement.
Outlier trimming is a decision, not a cleanup step
Somebody always says "let's drop the weird ones." That sentence changes the answer, and it's exactly what almost buried the fourth job in Ray's mess.
Drop the 14-hour jobs from a 6-hour part and your average gets prettier. It also hides the exact events eating your week. In this illustration, that event is the one job with a real tooling problem. Sometimes trimming is right. It's never neutral.
Write it down every time, in one line: what you cut, how many rows that was, and why. If cutting 3 percent of rows moves your headline number by 20 percent, that number isn't stable enough to act on yet.
The five-line note that should travel with every number
Any number that changes a price, a schedule, or a headcount should carry this with it. Five lines, plain language, no dashboard required.
- What it claims: setup averages 1.4 hours on this line.
- Source one: ERP job labor records, Jan through Jun, press line only.
- Source two: machine timestamps, same window, same line.
- They agree within: 6 percent. Tolerance set at 10 percent before the check.
- Excluded: 11 jobs with no end timestamp, marked UNMEASURED, not counted as zero.
If a vendor hands you a dashboard and can't produce those five lines for any tile on it, you're looking at a story. Ask them which lanes on their own dashboard are unmeasured. A vendor with no refusal cases hasn't looked hard enough.
Where to find honest outside baselines
You don't need a benchmark to run this test. But when you want to sanity-check your own scale, use published sources instead of a number somebody remembered from a trade show.
The Bureau of Labor Statistics productivity program publishes output-per-hour measures for manufacturing industries. The Federal Reserve's G.17 release publishes manufacturing capacity utilization every month. The Census Bureau's Annual Survey of Manufactures collects shipments, inventories, and hours by industry. If you want hands-on help, the NIST Manufacturing Extension Partnership works with small and mid-sized manufacturers in every state.
None of those will tell you what your press line does. They'll tell you when your own number is out by an order of magnitude, exactly when you most need to know.
We run this test on our own numbers
LegacyForge AI is an AI-run web agency. We build every site with the same AI systems we sell, which means we get to eat our own cooking on data honesty.
Our sites can ship with a 24/7 AI receptionist that answers the phone, chats, captures leads, and books jobs while the owner is on a machine or on a roof. That receptionist produces numbers: calls answered, leads captured, jobs booked. We apply the same two-source rule to those numbers that we would apply to a plant's setup time. If we cannot rebuild a figure from a second writer, we do not publish it as a win. We'd rather report a smaller true number than a bigger comfortable one.
That's also why the pitch here is a site you can open, not a claim you have to take on faith. Start at legacyforgeai.com and look at the work yourself. You're not evaluating our description of it, because you're generating your own first source on the spot.
What Ray did the next morning, and what you can do tonight
In this illustration, Ray pulls the fourth quote back from his desk before it goes out. Then he runs the two-hour version of everything above, on the one number that is still steering a live decision.
Pick one number that's currently steering a decision. Just one.
- Write down what it claims, in a sentence, with the date range and the lines it covers.
- Find a second table that was written by a different process and rebuild it.
- Set your tolerance before you run the second query.
- If they agree, act, and staple the five-line note to the number.
- If they disagree, work the four checks above in order.
- Label every blank lane UNMEASURED, never zero.
Two hours of this will usually find something that's been quietly re-pricing your jobs for years, the way one field tech's habit would re-price three of Ray's quotes in a single weekend. It costs nothing, and you keep the answer whether or not you ever talk to us.
Sites we build ship in days rather than agency-months, because the AI does the heavy lifting and a human does the taste. Same discipline, different ledger: we verify what the site actually did before we tell anyone it worked.
If you want the two-source treatment run on your ERP data, or a site that answers the phone while you're on the floor, start at legacyforgeai.com.
Questions people actually ask
What is the fastest way to check if an ERP number is real?
Rebuild it from a second table that was written by a different process, then compare. If the two versions land within your tolerance, act. If they do not, you have a data question, not an operations answer.
What tolerance should two sources agree within?
Pick it before you look, and pick it from the decision. Our recommended starting point for a number that moves a quote is that the two sources land within about 5 percent. For ranking jobs from worst to best, we start at 15 percent, because the ranking survives the gap.
Is a blank field the same as a zero?
No. Treating them the same is a mistake we see wreck plant reporting. A zero is a measurement that came back empty. A blank is a lane nobody measured, and it should be labeled UNMEASURED so it never gets averaged in.
