Built and running

Recorded run, not a simulation

Sixteen went straight through. It refused to answer three.

A crew writes a field ticket. Somebody in the office turns it into an invoice the operator will actually pay, or works out who to ask about it. This is a fleet of agents doing that job across four separate systems, recorded in full, including the tickets it would not touch.

  • 19 tickets
  • 4 systems
  • 7 stages
  • 3 of them use no AI at all

What is real here, and what is not

Real

The agents, every model call, every tool call, every network hop between systems, the token counts, the dollar costs, the time each ticket took, and the failures. This is one run that actually happened on 4 August 2026. It cost $1.85. Nothing on this page is a mock-up and nothing is scripted.

Stand-ins

The ERP, the field ticketing system, the billing ledger and the operator's accounts payable portal are four small programs, each running separately and each speaking its own dialect. You already own all four of those in some form and no demo can have them. Their behaviour is real: rate sheets, cost centre coding, per-operator submission windows and the specific ways a portal rejects an invoice are how this billing actually works.

Invented, and worth saying first

Every company name, and every number on the rate sheet. Not the wireline unit day rate, not the operator hourly rate, not the price of a perforating gun, not the submission windows. They are the right order of magnitude and the right shape for an East Texas wireline and pump down company. They are nobody's real rate sheet, and if you compare them to yours you will find differences and you will be right.

That matters less than it looks. The machinery does not change when the numbers do. Loading your own contract is a data change, not a rebuild.

What the run did

Every figure in this block was counted from the recorded run. None of it is adjustable and none of it is a projection.

Measured. Counted from the run, not estimated.

84%
Straight through
16 of 19 never reached a person
19/19
Decision accuracy
against an answer key written before the run
3
Escalated
each with a specific question attached
1
Duplicate closed
not billed to the operator twice
0
Failed
the fleet broke
$1.85
Model cost, whole run
$0.10 per ticket

$137,313 invoiced across 15 tickets · 11 minutes of wall clock, one ticket at a time · 271k tokens served from cache

Watch three of them

Three of the nineteen, chosen to include one the fleet refused to answer rather than three it got right. Press play and watch it read the ticket, price it against the contract, resolve the cost centre, check it, and send it. The timings are the recorded ones.

The ordinary case. Signed, in window, one open AFE on the lease. Read, priced against the contract, coded, checked and submitted without a person touching it.

What the crew wrote
PINEYWOODS PRESSURE SERVICES        FIELD TICKET  #FT-5001
Customer: SABINE RIDGE OPERATING         Date of Service: 07/06/26
Well: BURKHALTER 3H      Lease: BURKHALTER UNIT
AFE#: 26-0431
API: 42-183-31204
------------------------------------------------------------
WIRELINE UNIT                    1 DAY
WIRELINE OPERATOR                9 HRS
HELPER                           9 HRS
MILEAGE                          38 MI
PERF GUNS 3-1/8                  4 EA
------------------------------------------------------------
Cust Rep: J. Ainsworth      Signed: YES     07/06/26
What the fleet did

Press play.

Recorded from a real run, not live. The timings are the recorded ones, compressed for viewing. 3 of 19 tickets, chosen to include one it refused to answer.

The screen it is driven from

This is the console itself, photographed mid-run rather than mocked up: the full nineteen ticket queue on the left, one ticket streaming through the seven stages in the middle, and the economics panel keeping the measured and assumed columns apart on the right. It runs locally, next to the four sandbox systems. It is not on the internet, because its run button spends real money.

Captured 4 August 2026, working a ticket end to end. The board reads 84% straight through and 19 of 19 against the answer key; the three amber tickets in the queue are the escalations, and the panel on the right is the same measured against assumed split this page uses. Click the image to view it full size.
The Field Ticket to Cash console: a scoreboard reading 84% straight through and 19 of 19 decision accuracy, a ticket queue with three escalations, a live agent trace working ticket FT-4471 through triage, intake, pricing and coding, and an economics panel splitting measured figures from assumptions.

The three it would not answer

An unsigned ticket cannot be signed by software. A ticket past its contractual window cannot be brought back. A ticket covering two kinds of work with the cost centre left blank cannot be coded without somebody deciding how to split it. For these, the correct outcome is a specific question for a specific person, and the scorecard treats submitting one as a failure.

A fleet that submitted all nineteen would score worse, not better. That is the whole argument.
FT-5017

FT-5017 (Sabine Ridge Operating, $6792.00) stopped before submission.

Sabine Ridge Operating requires an approved signature and this ticket is unsigned.

Get the company man to sign the ticket, or obtain written approval by email.

FT-5018

FT-5018 (Neches Basin Energy, $6317.40) stopped before submission.

Service date is 68 days old; Neches Basin Energy allows 60.

This is outside the contractual window. It needs an operator exception before it can be billed at all.

FT-5019

FT-5019 (Caddo Lake Resources, $9010.50) stopped before submission.

Single ticket contains two distinct scopes of work (morning perforating, afternoon produced-water hauling) that map to two different open AFEs on the same well. Cannot be coded to one cost center without either a split or an owner decision.

Should FT-5019 be split — perforating/wireline lines to CLR-770 and the 210 bbl fluid transfer to CLR-781 (and if so, which AFE takes the 61 mi mileage) — or coded whole to one AFE?

Notice what each one hands over: not "review this", but the actual next action. That is the difference between an exception queue and a pile.

The money that left quietly

After the invoices were accepted, the operators ran their pay cycles. One of them remitted less than was billed and sent no notice about it, which is a thing every controller in this business has lived through. The fleet caught it on the ledger.

INV-8815
Billed$6,666
Remitted$6,400
Short, with no notice$266

Read the amount carefully. The mechanism is real. The specific rule that produces this particular gap was written for the demo, in this case an operator that pays a maximum of eight labour hours per person per day and treats the rest as needing prior written approval. That is a real kind of contract clause applied strictly, so the dollar figure is worth what an invented rate sheet times an invented cap is worth. What the demo is entitled to claim is that the money left silently and the fleet noticed. What it is worth on your invoices is a question about your own receivables.

Notice also what is not in this list. One ticket billed the operator above the contract rate, and the pricing stage repriced it to contract before it ever went out. That short pay never happened, so it never appears anywhere. Prevention beats detection and it is invisible when it works.

Three of the seven stages use no AI at all

This is the design decision worth defending, and it is the opposite of what a demo would do if the goal were to look impressive.

"Is this inside a 90 day window" is a subtraction. "Is it signed" is yes or no. "How many cents is $4,200" is a multiplication. Putting a language model on any of those costs money on every ticket, adds delay, and introduces the one failure a compliance check must never have: the ability to be talked out of its answer. The gate cannot invent a signature and cannot be persuaded that 68 days is inside 60.

The four stages that do use a model do not all use the same one. The two that carry the most text, reading the ticket and matching it to the rate sheet, run on a cheaper model, because extraction and table lookup do not need a frontier model. The expensive one is kept for the two stages where the right answer is sometimes "I will not decide this". That choice was measured rather than assumed: dropping the coding stage to the cheaper model cost 42% less and got three of thirteen decisions wrong, in both directions.

  • intake sonnet-5
  • pricing sonnet-5
  • coding opus-5
  • exceptions opus-5
  • triage · gate · submit no model call

What this page will not tell you

The cost of the machine is measured. The value of it depends entirely on numbers nobody has counted yet, and they are yours.

How many tickets you write a week, how long one takes your office today, what an hour of that time costs you loaded, and what share of your tickets genuinely need a person. Those four numbers decide whether this is worth building for you, they differ by an order of magnitude between a five truck shop and a regional company, and inventing them here would produce exactly the kind of confident fake number this practice exists to refuse. Establishing them against one real month of your own tickets is what the paid assessment does.

One more thing that is an assumption rather than a measurement: the mix of difficulty in these nineteen tickets. Fifteen clean or resolvable, one duplicate, three unresolvable was reasoned about, not counted. There is no client data behind it and no industry dataset behind it. It is a plausible ordinary week and it is not evidence about yours.

  • Measured over 19 tickets. Treat per-ticket cost as an order of magnitude, not a rate. A production sample of a few hundred tickets would tighten it.
  • The representative queue escalates 16%, from a modelled ordinary mix: 3 of 19 tickets have no valid automated answer. That distribution was REASONED, not measured, there is no client data behind it. The projection deliberately does NOT use that rate. It uses 20%, which is an assumption, and the real figure comes from counting a month of the client's own tickets.
  • Ticket volume, minutes per ticket and the loaded hourly rate are INPUTS, not measurements. Replace them with the client's own numbers before quoting anything.
  • Escalations are costed as real human time, not as free.
  • Recovered revenue, short pays caught, tickets rescued before their window closes, is excluded from the net saving and reported separately.
  • Model prices change. Cost is computed at read time from the published table, so a re-run reprices automatically.
  • Every rate, submission window and operator rule in the sandbox ERP was written for this demo. They are the right SHAPE of an MSA and they are not anybody's real rate sheet, so any dollar figure derived from an invoice here sizes a mechanism rather than a client's revenue.

Is it the same every time?

No, and anybody who tells you their system is has not run theirs twice.

Every ticket has a written correct answer, decided before the run and kept in a file the fleet cannot see. That is what makes accuracy a measurement rather than a claim, and it is what makes changing a model a measurement rather than an argument.

19/1919/1919/1918/19

Four recorded runs of this exact configuration against this exact queue. The single miss was a ticket that should have been accepted and was not, so nothing wrong reached an operator. It failed toward a person, which is the direction that matters.

The questions this page does not answer

How does it connect to my ERP?

This does not connect to a real one, and no demo can. What it shows is where the connection lives: every request an agent makes goes through one small translation layer, and swapping a stand-in for your real system is a change to that layer. The agents do not know or care what is behind it. If a system of yours has no interface at all, which is common, the options are the same as they have always been: a nightly file, a database view, or a person still doing that one step. Working out which applies to you is the first half of the assessment and the thing most likely to change the price.

What happens when it is wrong and a bad invoice reaches an operator?

You carry it, the same as you do today when a person keys it wrong, and the contract says so. What changes is that every decision is recorded with what the agent saw and why it decided, so a wrong invoice is a five minute investigation rather than an argument. The design decision that matters more is that anything genuinely ambiguous is escalated rather than guessed, and the scorecard penalises guessing.

What does it do at ten thousand tickets?

This runs one ticket at a time on purpose, so a person can follow it on a screen. Tickets do not depend on each other, so real volume runs many at once and the limit becomes how fast the model provider answers, not the design. What genuinely does not scale is the exception queue: at ten thousand tickets a month even a small escalation rate is a full time job, and that is a staffing conversation to have before signing anything, not after.

Who owns the code at the end?

You do. On full payment you own the deliverables built for you: the code, the prompts, the configuration, the documentation and the data flowing through it. You can take it in house, hand it to another vendor, or throw it away. There is no platform to be locked into and no seat licence.

What happens when an operator changes their portal or their contract?

It breaks, and you want it to break loudly rather than quietly submit wrong invoices. The rate sheet and the operator rules are data, so a rate change or a new window is an edit and not a rebuild. A portal that changes its interface is real work, and it is what a monthly arrangement is for. A portal that changes without warning shows up as rejections within a day, which is faster than most offices notice today.

How long before it works on my tickets?

Unknown until somebody looks at your tickets and your systems, which is what the assessment is for. Anybody who quotes you a timeline before seeing a real ticket and a real rate sheet is guessing. What the assessment produces is a fixed scope and a fixed price for the build, or an honest answer that there is not enough here to be worth building.

This is one sector. The machinery is not sector specific.

Field tickets are a good demonstration because the failure modes are sharp and the money is visible, but nothing in the design is about oilfield billing. It is a paperwork chain with four systems that disagree, a contract that governs what is billable, and a counterparty that rejects things without explaining itself. If that describes a process in your business, the assessment is where we find out whether it is worth automating.

Start an assessment

A fixed scope, a fixed price, and an honest answer if the answer is no.