Sixteen went straight through. It refused to answer three.
A crew writes a field ticket. Somebody in the office turns it into an invoice the operator will actually pay, or works out who to ask about it. This is a fleet of agents doing that job across four separate systems, recorded in full, including the tickets it would not touch.
What is real here, and what is not
Real
The agents, every model call, every tool call, every network hop between systems, the token counts, the dollar costs, the time each ticket took, and the failures. This is one run that actually happened on 4 August 2026. It cost $1.85. Nothing on this page is a mock-up and nothing is scripted.
Stand-ins
The ERP, the field ticketing system, the billing ledger and the operator's accounts payable portal are four small programs, each running separately and each speaking its own dialect. You already own all four of those in some form and no demo can have them. Their behaviour is real: rate sheets, cost centre coding, per-operator submission windows and the specific ways a portal rejects an invoice are how this billing actually works.
Invented, and worth saying first
Every company name, and every number on the rate sheet. Not the wireline unit day rate, not the operator hourly rate, not the price of a perforating gun, not the submission windows. They are the right order of magnitude and the right shape for an East Texas wireline and pump down company. They are nobody's real rate sheet, and if you compare them to yours you will find differences and you will be right.
That matters less than it looks. The machinery does not change when the numbers do. Loading your own contract is a data change, not a rebuild.
What the run did
Every figure in this block was counted from the recorded run. None of it is adjustable and none of it is a projection.
Measured. Counted from the run, not estimated.
$137,313 invoiced across 15 tickets · 11 minutes of wall clock, one ticket at a time · 271k tokens served from cache
Watch three of them
Three of the nineteen, chosen to include one the fleet refused to answer rather than three it got right. Press play and watch it read the ticket, price it against the contract, resolve the cost centre, check it, and send it. The timings are the recorded ones.
The ordinary case. Signed, in window, one open AFE on the lease. Read, priced against the contract, coded, checked and submitted without a person touching it.
PINEYWOODS PRESSURE SERVICES FIELD TICKET #FT-5001 Customer: SABINE RIDGE OPERATING Date of Service: 07/06/26 Well: BURKHALTER 3H Lease: BURKHALTER UNIT AFE#: 26-0431 API: 42-183-31204 ------------------------------------------------------------ WIRELINE UNIT 1 DAY WIRELINE OPERATOR 9 HRS HELPER 9 HRS MILEAGE 38 MI PERF GUNS 3-1/8 4 EA ------------------------------------------------------------ Cust Rep: J. Ainsworth Signed: YES 07/06/26
Press play.
The one that would have been billed twice. A re-keyed copy of a ticket already invoiced. Dispatch was not sure the original made it in. The operator's portal rejected it as a duplicate and the fleet closed it rather than arguing with the rejection.
PINEYWOODS PRESSURE SERVICES FIELD TICKET #FT-5016 Customer: SABINE RIDGE OPERATING Date of Service: 07/17/26 Well: BURKHALTER 3H Lease: BURKHALTER UNIT AFE#: 26-0431 API: 42-183-31204 ------------------------------------------------------------ WIRELINE UNIT 1 DAY WIRELINE OPERATOR 8.5 HRS HELPER 8.5 HRS MILEAGE 38 MI PERF GUNS 3-1/8 5 EA ------------------------------------------------------------ Cust Rep: J. Ainsworth Signed: YES 07/17/26 Second copy - dispatch was not sure the original made it in.
Press play.
The one it refused to answer. One ticket covering perforating and produced-water hauling, same crew, same day, with the AFE line left blank. The work genuinely spans two cost centres and the ticket does not say how to split it. Guessing has a fifty percent chance of being wrong and a hundred percent chance of sounding certain.
PINEYWOODS PRESSURE SERVICES FIELD TICKET #FT-5019 Customer: CADDO LAKE RESOURCES Date of Service: 07/30/26 Well: NOONDAY 1H Lease: NOONDAY AFE#: API: 42-423-27004 ------------------------------------------------------------ WIRELINE UNIT 1 DAY WIRELINE OPERATOR 6 HRS PERF GUNS 3-1/8 2 EA FLUID TRANSFER 210 BBL MILEAGE 61 MI ------------------------------------------------------------ Cust Rep: T. Boudreaux Signed: YES 07/30/26 Crew perforated in the morning and hauled produced water in the afternoon. Same crew, same day, one ticket.
Press play.
Recorded from a real run, not live. The timings are the recorded ones, compressed for viewing. 3 of 19 tickets, chosen to include one it refused to answer.
The screen it is driven from
This is the console itself, photographed mid-run rather than mocked up: the full nineteen ticket queue on the left, one ticket streaming through the seven stages in the middle, and the economics panel keeping the measured and assumed columns apart on the right. It runs locally, next to the four sandbox systems. It is not on the internet, because its run button spends real money.
The three it would not answer
An unsigned ticket cannot be signed by software. A ticket past its contractual window cannot be brought back. A ticket covering two kinds of work with the cost centre left blank cannot be coded without somebody deciding how to split it. For these, the correct outcome is a specific question for a specific person, and the scorecard treats submitting one as a failure.
A fleet that submitted all nineteen would score worse, not better. That is the whole argument.
FT-5017 (Sabine Ridge Operating, $6792.00) stopped before submission.
Sabine Ridge Operating requires an approved signature and this ticket is unsigned.
Get the company man to sign the ticket, or obtain written approval by email.
FT-5018 (Neches Basin Energy, $6317.40) stopped before submission.
Service date is 68 days old; Neches Basin Energy allows 60.
This is outside the contractual window. It needs an operator exception before it can be billed at all.
FT-5019 (Caddo Lake Resources, $9010.50) stopped before submission.
Single ticket contains two distinct scopes of work (morning perforating, afternoon produced-water hauling) that map to two different open AFEs on the same well. Cannot be coded to one cost center without either a split or an owner decision.
Should FT-5019 be split — perforating/wireline lines to CLR-770 and the 210 bbl fluid transfer to CLR-781 (and if so, which AFE takes the 61 mi mileage) — or coded whole to one AFE?
Notice what each one hands over: not "review this", but the actual next action. That is the difference between an exception queue and a pile.
The money that left quietly
After the invoices were accepted, the operators ran their pay cycles. One of them remitted less than was billed and sent no notice about it, which is a thing every controller in this business has lived through. The fleet caught it on the ledger.
Read the amount carefully. The mechanism is real. The specific rule that produces this particular gap was written for the demo, in this case an operator that pays a maximum of eight labour hours per person per day and treats the rest as needing prior written approval. That is a real kind of contract clause applied strictly, so the dollar figure is worth what an invented rate sheet times an invented cap is worth. What the demo is entitled to claim is that the money left silently and the fleet noticed. What it is worth on your invoices is a question about your own receivables.
Notice also what is not in this list. One ticket billed the operator above the contract rate, and the pricing stage repriced it to contract before it ever went out. That short pay never happened, so it never appears anywhere. Prevention beats detection and it is invisible when it works.
Three of the seven stages use no AI at all
This is the design decision worth defending, and it is the opposite of what a demo would do if the goal were to look impressive.
"Is this inside a 90 day window" is a subtraction. "Is it signed" is yes or no. "How many cents is $4,200" is a multiplication. Putting a language model on any of those costs money on every ticket, adds delay, and introduces the one failure a compliance check must never have: the ability to be talked out of its answer. The gate cannot invent a signature and cannot be persuaded that 68 days is inside 60.
The four stages that do use a model do not all use the same one. The two that carry the most text, reading the ticket and matching it to the rate sheet, run on a cheaper model, because extraction and table lookup do not need a frontier model. The expensive one is kept for the two stages where the right answer is sometimes "I will not decide this". That choice was measured rather than assumed: dropping the coding stage to the cheaper model cost 42% less and got three of thirteen decisions wrong, in both directions.
- intake sonnet-5
- pricing sonnet-5
- coding opus-5
- exceptions opus-5
- triage · gate · submit no model call
What this page will not tell you
The cost of the machine is measured. The value of it depends entirely on numbers nobody has counted yet, and they are yours.
How many tickets you write a week, how long one takes your office today, what an hour of that time costs you loaded, and what share of your tickets genuinely need a person. Those four numbers decide whether this is worth building for you, they differ by an order of magnitude between a five truck shop and a regional company, and inventing them here would produce exactly the kind of confident fake number this practice exists to refuse. Establishing them against one real month of your own tickets is what the paid assessment does.
One more thing that is an assumption rather than a measurement: the mix of difficulty in these nineteen tickets. Fifteen clean or resolvable, one duplicate, three unresolvable was reasoned about, not counted. There is no client data behind it and no industry dataset behind it. It is a plausible ordinary week and it is not evidence about yours.
- Measured over 19 tickets. Treat per-ticket cost as an order of magnitude, not a rate. A production sample of a few hundred tickets would tighten it.
- The representative queue escalates 16%, from a modelled ordinary mix: 3 of 19 tickets have no valid automated answer. That distribution was REASONED, not measured, there is no client data behind it. The projection deliberately does NOT use that rate. It uses 20%, which is an assumption, and the real figure comes from counting a month of the client's own tickets.
- Ticket volume, minutes per ticket and the loaded hourly rate are INPUTS, not measurements. Replace them with the client's own numbers before quoting anything.
- Escalations are costed as real human time, not as free.
- Recovered revenue, short pays caught, tickets rescued before their window closes, is excluded from the net saving and reported separately.
- Model prices change. Cost is computed at read time from the published table, so a re-run reprices automatically.
- Every rate, submission window and operator rule in the sandbox ERP was written for this demo. They are the right SHAPE of an MSA and they are not anybody's real rate sheet, so any dollar figure derived from an invoice here sizes a mechanism rather than a client's revenue.
Is it the same every time?
No, and anybody who tells you their system is has not run theirs twice.
Every ticket has a written correct answer, decided before the run and kept in a file the fleet cannot see. That is what makes accuracy a measurement rather than a claim, and it is what makes changing a model a measurement rather than an argument.
Four recorded runs of this exact configuration against this exact queue. The single miss was a ticket that should have been accepted and was not, so nothing wrong reached an operator. It failed toward a person, which is the direction that matters.
The questions this page does not answer
How does it connect to my ERP?
This does not connect to a real one, and no demo can. What it shows is where the connection lives: every request an agent makes goes through one small translation layer, and swapping a stand-in for your real system is a change to that layer. The agents do not know or care what is behind it. If a system of yours has no interface at all, which is common, the options are the same as they have always been: a nightly file, a database view, or a person still doing that one step. Working out which applies to you is the first half of the assessment and the thing most likely to change the price.
What happens when it is wrong and a bad invoice reaches an operator?
You carry it, the same as you do today when a person keys it wrong, and the contract says so. What changes is that every decision is recorded with what the agent saw and why it decided, so a wrong invoice is a five minute investigation rather than an argument. The design decision that matters more is that anything genuinely ambiguous is escalated rather than guessed, and the scorecard penalises guessing.
What does it do at ten thousand tickets?
This runs one ticket at a time on purpose, so a person can follow it on a screen. Tickets do not depend on each other, so real volume runs many at once and the limit becomes how fast the model provider answers, not the design. What genuinely does not scale is the exception queue: at ten thousand tickets a month even a small escalation rate is a full time job, and that is a staffing conversation to have before signing anything, not after.
Who owns the code at the end?
You do. On full payment you own the deliverables built for you: the code, the prompts, the configuration, the documentation and the data flowing through it. You can take it in house, hand it to another vendor, or throw it away. There is no platform to be locked into and no seat licence.
What happens when an operator changes their portal or their contract?
It breaks, and you want it to break loudly rather than quietly submit wrong invoices. The rate sheet and the operator rules are data, so a rate change or a new window is an edit and not a rebuild. A portal that changes its interface is real work, and it is what a monthly arrangement is for. A portal that changes without warning shows up as rejections within a day, which is faster than most offices notice today.
How long before it works on my tickets?
Unknown until somebody looks at your tickets and your systems, which is what the assessment is for. Anybody who quotes you a timeline before seeing a real ticket and a real rate sheet is guessing. What the assessment produces is a fixed scope and a fixed price for the build, or an honest answer that there is not enough here to be worth building.
This is one sector. The machinery is not sector specific.
Field tickets are a good demonstration because the failure modes are sharp and the money is visible, but nothing in the design is about oilfield billing. It is a paperwork chain with four systems that disagree, a contract that governs what is billable, and a counterparty that rejects things without explaining itself. If that describes a process in your business, the assessment is where we find out whether it is worth automating.
Start an assessment ↗A fixed scope, a fixed price, and an honest answer if the answer is no.