The cert claimed A572-50. Its own numbers said no.
A customer emails a steel order. Somebody at the desk quotes it, pulls heats off the floor, checks the mill paperwork, ships it, invoices it, and chases what the remittance got wrong. This is a fleet of agents doing that job across four systems, recorded in full, with one check almost nobody runs: it reads the mill test report and re-runs the chemistry arithmetic itself, so a certificate whose own printed numbers cannot make the grade it claims gets stopped before the truck moves.
What is real here, and what is not
Real
The agents, every model call, every tool call, every network hop between the four systems, the token counts, the dollar costs, the time each order took, and the failures. Real above all: the chemistry limits. The gate checks certificates against published ASTM composition windows, carbon equivalent by the IIW formula, and mechanical ranges by test location, and where a limit could not be verified against the specification text the system refuses to enforce it and routes the certificate to a person instead of inventing a violation. An unknown beats a disputed number, everywhere in this build.
Stand-ins
The inventory, the order desk, the cert vault and the receiving dock are four small programs, each running separately and each speaking its own dialect, standing in for "Piney Timbers Steel Supply, Longview TX", a service center that does not exist. You already own all four of those in some form and no demo can have them. Their behaviour is real: bundle tags carry heat numbers, the dock runs its own dumber cert check and rejects loads in its own vocabulary, and remittances arrive as an amount against a reference with no explanation attached.
Invented, and worth saying first
Every mill, every heat number, every customer, and every price. A real mill's name never appears over an invented heat number. The certificates themselves were generated for this demo, each bad one carrying exactly one planted defect, and the build verifies that the defect the gate finds is the defect that was planted. The price book is nobody's real price book, and if you compare it to yours you will find differences and you will be right.
That matters less than it looks. The machinery does not change when the numbers do. Loading your grades, your mills and your price book is a data change, not a rebuild.
What the latest run did
Every figure in this block was counted from the recorded run of the roster that now ships as the default. None of it is adjustable, none of it is a projection, and the runs before it, including the worse ones, are published further down.
Measured. Counted from the recording, not estimated.
The gate's own work, counted at the site of each check: 29 heats allocated · 20 certs re-validated · 20 carbon equivalents recomputed · 327 window comparisons · every one deterministic, every one free
Watch three of them
Three of the thirty-five, chosen to end on a blocked certificate rather than on three successes. Press play and watch the fleet read the order, price it from the book, allocate real bundles off the floor, and re-check the mill paperwork. The timings are the recorded ones, compressed for viewing. In a hurry, click S1 first: that is the certificate that refuted itself.
The ordinary case. A heavy column order, mill certs required. The chemistry on file is clean, the gate re-checks every window and finds nothing, the dock accepts the load. Most orders are this one.
PO 20447 PURCHASE ORDER 20447 - Caddo Structural, Marshall TX 2 pcs W14x283 x 40 ft, ASTM A992, per the mill-order passthrough we discussed. Certs required. Column line. Crane is on site Thursday.
Press play.
The one it routed to an engineer. The order says A36 or approved equal, and the stock on the floor is multi-certified paper that covers A36 exactly. The desk could talk itself into shipping. It does not: an or-equal clause belongs to the customer's engineer, so the fleet prices the job and routes the substitution question to the named person instead of guessing.
quote request - W10x22 Can you quote 6 pcs W10x22 x 40 ft, A36 or equal? Bid goes in Friday, need your number by Thursday noon. Josh W., Piney Creek Welding
Press play.
The cert that refuted itself. Half-inch plate ordered dual-certified A36/A572-50. The mill cert claims both grades, and the carbon printed on its own chemistry line is legal for A36 alone. The gate re-runs the arithmetic and blocks the shipment with the element, the printed value, the limit and the delta named. No model made this call and no model can argue it open.
PO 8127 - half inch plate, dual certs required Please enter our PO 8127. 12 pcs PL 1/2 x 96 x 240, dual certified A36/A572-50. MTRs required with the shipment and both grades must appear on the cert. Deliver to the Lufkin yard, week of 8/17. Dana R., Angelina Fabricators purchasing
Press play.
Recorded from a real run, not live. The timings are the recorded ones, compressed for viewing. 3 of 35 orders, chosen to end on a blocked certificate.
The screen it is driven from
The console itself, captured mid-replay rather than mocked up: the order book on the left, one order streaming through the stages in the middle, and the gate findings on the right naming the element, the printed value and the limit for every blocked certificate. The banner across the top says REPLAY because that is what it is: the recorded event stream playing back with nothing spent. The live run button exists only on a laptop, never on this site, because it spends real money without asking.
The four certificates it refused to ship
Each card below is a mill test report whose own printed numbers fail the grade it claims, caught by arithmetic. A carbon of 0.24 on paper claiming dual A36/A572-50 is legal for A36 alone, so the cert refutes exactly half its claim. A yield of 67 ksi on an A992 wide flange exceeds a ceiling that exists for seismic ductility, and a simple "yield at least 50" check passes it. A sulfur out of window. A missing Charpy row on impact-tested work, because you cannot retest paperwork into existence. Every block names the heat, the element, the printed value, the limit and the delta.
Heat 7608213: C 0.24 exceeds A572-50's 0.23 max (over by 0.01) while staying legal for A36 alone - the cert refutes the A572-50 claim specifically
Quarantine the heat and take it up with the mill. The cert and the claim cannot both be right.
Heat 5801507: yield 67 ksi exceeds A992's 65 ceiling for a flange-tested shape - the range exists for seismic ductility, and a "yield >= 50" check passes this cert
Quarantine the heat and take it up with the mill. The cert and the claim cannot both be right.
Heat 7608215: S 0.063 exceeds A36's 0.05 max (over by 0.013)
Either this transfer cert was mistyped or it was altered - both require the original mill cert before shipment. The fleet flags; a person obtains the original.
The order requires Charpy impact values; heat 7608216's MTR has no CVN row. Impact testing was never performed - you cannot retest paperwork into existence.
Source CVN-tested material, or send this heat out for testing with the lead time stated. The chemistry looking tough enough is not a test result.
No model made these calls, and no model can be talked into reversing them. A compliance check that cannot be argued with is the point.
A fleet that shipped all thirty-five orders would score worse, not better. The answer key counts a blocked bad certificate as a success and a wrong shipment confidently submitted as the failure.
The twelve it handed to a person
An or-equal substitution belongs to the customer's engineer, not to software that can talk itself into anything. An expired quote with a price move is a commercial call. A shortfall where the choice is partial now or hold complete belongs to the buyer sequencing a job site. For these, the correct outcome is a specific question routed to a named person, and that is what the fleet produces: not "review this", but the actual decision that needs a human, with the numbers attached.
Heat 5801508 is on hand but its mill cert is not on file. Hold the line, chase the supplier for the original cert, or re-allocate around it?
Heat 7608214 is on hand but its mill cert is not on file. Hold the line, chase the supplier for the original cert, or re-allocate around it?
Order S7 (RFQ for bid due Friday) requests 6 pcs W10x22 x 40 ft, A36 or equal. Quote Q-3010 was built at base grade A36 (rate $56/CWT, $2956.80 total), but the or-equal language means your drawings route material approval through your engineer. Can you confirm your engineer will review and approve A36 stock (or specify the exact equal grade needed) in writing before we finalize for your Friday bid submission?
Ship the 9 covered pieces now and back-order 3, or hold the line complete? A job site sequencing steel cares which; only the customer knows.
PO 5561 from Piney Creek Welding cites Q-1188, which was issued 2026-07-17 and expired 2026-07-31, meaning it was 11 days expired when the PO was received on 2026-08-11. Current book pricing (Q-3013, valid to 2026-08-25) totals the same $2,273.60 for this line. Do we honor the expired Q-1188 price of $2,273.60, re-quote at current book (also $2,273.60 under Q-3013), or hold for a call with the customer?
PO 4488 (Trinity Bay Services, received 2026-08-11) requests 25 pcs 1 in RD 1018 cold finished x 12 ft with a cert showing a guaranteed minimum yield of 60 ksi. 1018 cold finished bar per ASTM A108 has no guaranteed mechanical minimums; mill certs print measured actual values only (on-hand heat 5801510 shows C 0.18, Mn 0.69, P 0.008, S 0.011, CE(IIW) 0.295, but no yield guarantee field exists on this spec). We cannot issue paper claiming a guaranteed 60 ksi minimum. How does Ray B. want to proceed?
Piney Creek Welding (PO 5610) states they will weld lifting lugs on site to 4 pcs RD 2 4140-HT round x 12 ft. The on-hand heat (5801511) cert shows C 0.40, Mn 0.81, Cr 0.95, Mo 0.18, giving a CE (IIW) of 0.761, far above standard weldability guidance thresholds (typically 0.40 to 0.45 for no-preheat welding). Should we contact the customer to discuss required preheat/PWHT procedures before shipment, or ship as-is with the cert and note attached?
Quote Q-1201 for PO 91-4402 (60 pcs W16x36 x 40 ft, A992) was priced on theoretical weight: 86,400 lb theoretical = 864 CWT at $56/CWT, extending to $48,384.00 total. The PO states 'Invoice on actual scale weight. Our AP matches invoices to the certified truck scale ticket,' which is an explicit scale-weight basis demand, not a pricing-unit note. Do we invoice per the quoted theoretical basis ($48,384.00) or switch to actual certified scale weight for this shipment, and if the latter, does the $56/CWT rate carry over?
Please confirm the exact grade letter/spec required (A514 Type A, B, F, etc.), the source/mill requirement, and acceptable lead time, so a specialty quote can be built. Also confirm piece count (4 pieces stated as uncertain) and plate thickness/size (1 inch stated) before pricing.
PO for 8 pcs W8x31 x 30 ft states 'A36 or approved equal.' Quote Q-3034 is priced at stated A36 (issued 2026-08-11, expires 2026-08-25, total $4211.90). On-hand stock is heat 4105146, certified to A992/A572-50/A36 multi-cert (CE IIW 0.367). Does Sabine Pass Marine's engineer approve this heat/dual-cert stock as the 'approved equal,' or do they require material certified to A36 only?
PO 8141 (received 2026-08-11) cites quote Q-2041, which was issued 2026-07-20 and expired 2026-08-03, making it 8 days expired at the time of the PO. The order total as quoted is 20,384.00 USD. A current quote (Q-3035) is on file at the same book price of 20,384.00 USD, valid through 2026-08-25. Should we honor the expired Q-2041 pricing, re-quote at current book (Q-3035, same total), or contact the customer before proceeding?
Ship the 9 covered pieces now and back-order 3, or hold the line complete? The balance is on MILL-PO-7781, due 2026-08-23. A job site sequencing steel cares which; only the customer knows.
Nine of the fifteen stress orders and three of the twenty representative ones have no valid ship-it answer by construction. Refusing to guess on those is the behaviour being demonstrated, and the scorecard treats guessing as a failure.
The money that left quietly
After shipment, one customer paid less than the invoice and said nothing about it. The remittance recomputes exactly to the scale weight of the load at the invoice rate: they paid on what the truck weighed, against an invoice billed on theoretical weight, which is the industry's default basis. Reconcile caught it by arithmetic and named the basis, not just the gap.
Short paid 504.00: invoice billed theoretical weight, remittance recomputes exactly to the scale ticket (63900 lb at the invoice rate). The customer paid on received scale weight against a theoretical-basis invoice. Not a rounding difference; needs the basis conversation and a decision to collect or concede.
Read the amount carefully. The mechanism is real. The dollar figure sits on an invented price book, so it is worth what an invented price book is worth. What the demo is entitled to claim is that the money left silently and the fleet noticed, with the reason named. What that is worth on your invoices is a question about your own receivables.
Notice also the direction of every failure that did happen across the recorded runs: toward a person, never toward a wrong shipment. Prevention beats detection, and refusal beats both when the paper is bad.
Six of the ten stages use no AI at all
This is the design decision worth defending, and it is the opposite of what a demo would do if the goal were to look impressive.
"Is carbon 0.24 inside a 0.23 window" is a comparison. "Does the pick list cover the order" is arithmetic. "Does the quoted total match the book" is recomputed from scratch by code before any quote goes out. Putting a language model on any of those would cost money on every order, add delay, and introduce the one failure a compliance check must never have: the ability to be talked out of its answer. The gate cannot be persuaded that 0.24 is inside 0.23, and it costs nothing to run, which is why the per-order model cost stays around a nickel.
The four stages that do use a model do not all use the same one, and the assignments are measurements rather than opinions. The judgment stage originally ran on the frontier model; two full recorded passes on the cheaper one scored perfectly, including both or-equal traps, so it now ships on the cheaper model and the archive holds the evidence. The cheapest model was also tried on intake: it held accuracy and saved almost nothing, because intake was a small share of spend and prefix caching favours fewer models, so that downgrade was not kept. Exceptions stays on the frontier model because that stage never fired in the recent runs, and an untested downgrade is a guess.
- intake sonnet-5
- quote sonnet-5
- review sonnet-5
- exceptions opus-5
- triage · quote-verify · allocation · gate · ship+invoice · reconcile no model call
Every recorded run, including the bad one
Every order has a written correct answer, decided when the corpus was generated and kept where the fleet cannot read it. The table below is every recorded run against that key, never just the best run. The first pass scored 30 of 35 and every miss is named. Four prompt fixes were diagnosed from it, the second pass proved them and missed twice on mechanical failures, both of those were hardened, and the four passes since have not missed. The vendors in this space publish "99% accuracy" with no denominator, no key and no misses. This page publishes the spread, because a number you can check beats a number you have to trust.
| run | configuration | stress | representative | cost | misses, named |
|---|---|---|---|---|---|
| 1 | baseline · intake sonnet-5 / quote sonnet-5 / review opus-5 / exceptions opus-5 | 12/15 | 18/20 | $2.49 | S3, S7, S8, R16, R17 |
| 2 | after 4 diagnosed prompt fixes (commit 8de1ad7) · intake sonnet-5 / quote sonnet-5 / review opus-5 / exceptions opus-5 | 14/15 | 19/20 | $2.01 | S7, R16 |
| 2 | hardening retry (S7, R16 reopened) · intake sonnet-5 / quote sonnet-5 / review opus-5 / exceptions opus-5 | S7 correct | R16 correct | $0.14 | none |
| 3 | baseline-repeat · intake sonnet-5 / quote sonnet-5 / review opus-5 / exceptions opus-5 | 15/15 | 20/20 | $1.99 | none |
| 4 | sonnet-on-review · intake sonnet-5 / quote sonnet-5 / review sonnet-5 / exceptions opus-5 | 15/15 | 20/20 | $1.64 | none |
| 5 | haiku-on-intake · intake haiku-4-5 / quote sonnet-5 / review opus-5 / exceptions opus-5 | 15/15 | 20/20 | $1.96 | none |
| 6 | sonnet-on-review-confirm · intake sonnet-5 / quote sonnet-5 / review sonnet-5 / exceptions opus-5 | 15/15 | 20/20 | $1.75 | none |
The two corpora are graded separately and never averaged. The stress set is fifteen adversarial orders where thirteen have no valid ship-it answer by construction, so its escalation rate measures the difficulty we chose, not the fleet and not any real queue. The representative book is a modelled ordinary week: a reasoned mix, not client data, and replacing it with one real month of your orders is the first thing a paid engagement does.
What this page will not tell you
The cost of the machine is measured. The value of it depends entirely on numbers nobody has counted yet, and they are yours.
How many orders your desk works in a week, how long the paperwork takes today, what an hour of that time costs loaded, how often bad paper actually reaches your floor, and what one bad shipment costs you when it gets through. Those numbers decide whether this is worth building for you, they differ by an order of magnitude between shops, and inventing them here would produce exactly the kind of confident fake number this practice exists to refuse. Establishing them against one real month of your own orders is what the paid assessment does.
One more honesty note: we went looking for anyone else who publicly demonstrates this combination, an agent fleet that re-runs cert chemistry at order time on allocated inventory, refuses with a specific question for a named person, and publishes its own miss rate against a checkable answer key. As of 5 August 2026 we could not find one. Document checkers that validate certificates exist and some are good; none we found publishes an accuracy number you can verify, and none works the order.
The questions this page does not answer
Is the chemistry check real?
The windows are published ASTM limits and the comparisons are ordinary arithmetic in ordinary code, which you can read. Carbon equivalent is recomputed from the printed chemistry by the IIW formula and never trusted from the certificate; a printed value that disagrees with its own chemistry is itself a finding. Where a window could not be verified against the specification text, the system says so and routes to a person rather than enforcing a number it cannot defend. The certificates in the demo are generated, one planted defect per bad document, and the build fails if the gate finds anything other than exactly the planted defect.
How does it connect to my ERP?
This demo does not connect to a real one, and no demo can. What it shows is where the connection lives: every request an agent makes goes through one small translation layer, and swapping a stand-in for your real system is a change to that layer. If a system of yours has no interface at all, which is common, the options are the same as they have always been: a nightly file, a database view, or a person still doing that one step. Working out which applies to you is the first half of the assessment and the thing most likely to change the price.
What happens when it is wrong and a bad load ships?
You carry it, the same as when a person misses it today, and the contract says so. What changes is that every decision is recorded with what the agent saw and why it decided, so a wrong shipment is a five minute investigation rather than an argument. The design decision that matters more: anything genuinely ambiguous escalates rather than guesses, the deterministic checks cannot be argued with at all, and the scorecard penalises guessing.
Is it the same every time?
No, and anybody who says their system is has not run theirs twice. That is what the grid above is for: the spread is published, the misses are named, and the deterministic core has never varied, because the parts that must never vary do not use a model. Costs and timings wobble between runs. The judgments, across the last four recorded passes, did not.
Where does the data live and who can see it?
In this demo, nowhere but one laptop; the only thing that leaves it is order text sent to the model provider. On a real engagement the answer is settled in writing before any of your data moves: which provider, what they retain and for how long, where anything stored lives, retention windows enforced by a scheduled job rather than by somebody remembering, and how deletion works when you ask.
Who owns the code at the end?
You do. On full payment you own the deliverables built for you: the code, the prompts, the configuration, the documentation and the data flowing through it. You can take it in house, hand it to another vendor, or throw it away. There is no platform to be locked into and no seat licence.
This is one desk. The discipline is not steel specific.
A steel order desk is a good demonstration because the paperwork is unforgiving and the failure is physical: the wrong heat in a structure is not an accounting error. But the pattern underneath is general: documents that make checkable claims, a workflow that has to act on them, arithmetic that should never be delegated to a model, and judgment calls that should reach a named person as a specific question. If a process in your business fits that description, the assessment is where we find out whether it is worth automating.
Start an assessment ↗A fixed scope, a fixed price, and an honest answer if the answer is no.