The Decision Audit for AI Steps
A worksheet for your GTM stack: find the AI steps that only say yes or no, work out what they cost, and set a confidence threshold for handing off.
What this guide is for
In almost every GTM stack a large language model answers questions that only need a yes, a no or a category. Does the company fit the ICP? Is this reply an out-of-office? Is this person a decision maker? The model writes no text for any of that. It still gets paid like an author.
This guide gives you an audit that runs in one afternoon: classify every AI step, work out the cost per million decisions, and set a threshold below which a human takes over. At the end you have a list of what moves, what stays and what becomes a formula.
The starting point, in four numbers
We measured this in our own workspace: 170 real decisions from 13 table columns, each with about 1,000 tokens of input. Same input, four models, priced at list rates:
| Model | Cost per 1 million decisions |
|---|---|
| Claude Sonnet | about $2,300 |
| Claude Haiku | about $1,150 |
| Gemini Flash | about $380 |
| Jev (Typesafe), a typed decision model | about $41 |
Cheap only counts if it is right. Wherever Jev and the stored verdict disagreed, a stronger model judged blind which side followed the column’s rule. Jev was right in 126 of 129 checked rows. At confidence 0.8 or higher, which covered 81 per cent of rows, it was right in 104 of 104.
That matches the research. Bucher and Martini show that small, specialised models still beat large language models at text classification (arXiv 2406.08660). FrugalGPT by Chen, Zaharia and Zou matches the best single model with a cascade of cheap and expensive models at up to 98 per cent lower cost (arXiv 2305.05176).
What does not move
Before you calculate: the biggest cost block in most AI stacks is not the decision, it is the reading. An agent that reads a website, a conversation history and three data sources pays for context. For us that is three quarters of model spend; table columns are only about an eighth. A decision model replaces the choice at the end, not the reading before it.
Counted honestly, this audit is not a lever for half your bill. It is the cheapest lever you have, because it costs nothing in quality.
Step 1: Classify every AI step
Write down every place a model is called: table columns, workflow steps, agent tasks. Then give each one of four classes:
- Only decides. The output is only yes/no, a category from a fixed list or a named level. Candidate for moving.
- Decides and writes. A verdict plus a justification, a hook, a summary. Only the decision part can move; the text stays on the language model.
- Only writes. Copy, salutation, reply draft. Stays on the language model.
- Actually calculates. A field that combines other fields (“fits = all four criteria met”) or just reads a value back. That is a formula, not a question, and then costs nothing.
The fourth class is the one most often missed. In our own test a combined field reached only 48 per cent agreement, not because of the model, but because you compute an AND, you do not ask for it.
Step 2: Work out the cost per million decisions
The formula is enough for a rough estimate:
Cost per decision = input tokens × input price + output tokens × output price
Cost per million = cost per decision × 1,000,000
Two notes:
Measure the input, don’t guess it. The input is the prompt including rules and the row’s data, usually 500 to 4,000 tokens. A decision model bills by input. If you calculate with 500 tokens and send 3,800, you are off by a factor of seven.
Scale to your monthly volume. Decisions per month times cost per decision, once with today’s model, once with the decision model. The difference is the ceiling of the saving, not the saving.
Step 3: Set a confidence threshold
A language model writes “no” with the same face whether it was sure or not. A decision model returns a confidence. That, not the price, is the real gain.
The rule we use:
- Confidence 0.8 or higher: the verdict stands.
- Confidence below 0.8: the verdict is marked uncertain and goes to a human or a bigger model.
Research calls this selective classification: a classifier may abstain on uncertain cases and so hold a chosen error rate (Geifman and El-Yaniv 2017, arXiv 1705.08500). That models can estimate their own accuracy usefully is shown by Kadavath and colleagues (arXiv 2207.05221). And that part of the traffic can be routed to a cheaper model without lowering answer quality is shown by RouteLLM (arXiv 2406.18665).
The cost of the hand-off belongs in the calculation. For us, 19 per cent fell below the threshold. With Sonnet as the fallback, a million decisions then cost about $478 instead of $2,300; with Gemini Flash about $113 instead of $380.
Step 4: Compare on a sample before the old one goes
Switch no column without running both sides next to each other:
- Pull 50 to 100 rows, stratified: as many yes as no verdicts from today’s model.
- Run the decision model on the same rows with the same rules and data.
- Count agreement. Below 90 per cent, read the disagreements by hand.
- For each disagreement, ask: who follows the rule? In our test that was the decision model in 35 of 40 disputed rows. The stored verdict is a reference point, not the truth.
Two patterns we found along the way:
- Default-yes questions work badly. “False only for a clear new founding” reached 26 per cent. Ask positively (“is it a new founding?”) and invert the result.
- Truncated rules look like a bad model. If all disagreements run in one direction, first check whether the model got the whole rule.
The worksheet
One line per AI step. One page, not a deck.
Step / column: _____________________________________________
Class: only decides / decides + writes / only writes / calculates
Output: yes-no / category (list: ______) / level (______)
Input tokens (measured): ______
Decisions per month: ______
Cost today per million: $______
Cost decision model per million: $______
Saving ceiling per month: $______
Confidence threshold: 0.8 / other: ______
Below the threshold goes to: human / model: ______
Sample: ___ rows, agreement ___ %, rule winner: ______
Decision: move / decision part only / stays / becomes a formula
Two rules for the sheet:
An empty token line is not an estimate. No measured input, no cost figure.
“Move” needs a sample. Without a comparison it is a guess with a savings target.
Four principles that stay
Writing costs, deciding does not. Pay author prices only for text.
Confidence is the gain. A verdict that reports its uncertainty can be handed off. One that does not can only be hoped for.
Combined fields get calculated. An AND is not a question.
The reading stays. If you want lower agent costs, shorten the context; don’t just make the decision cheaper.
The short version
Classify every AI step, measure the input, work out the cost per million, trust the decision model from 0.8 confidence and hand off below it. Compare on a sample first.
The effort is one afternoon. The result is a list of where you pay author prices for a yes or a no today.
Sources
- Bucher, M. J. J., Martini, M. (2024): Fine-Tuned ‘Small’ LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification. arXiv:2406.08660
- Chen, L., Zaharia, M., Zou, J. (2023): FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176
- Ong, I. et al. (2024): RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665
- Geifman, Y., El-Yaniv, R. (2017): Selective Classification for Deep Neural Networks. arXiv:1705.08500
- Kadavath, S. et al. (2022): Language Models (Mostly) Know What They Know. arXiv:2207.05221
- Our own measurement in the CegTec workspace, September 2026: 170 table decisions from 13 columns, blind adjudication of 129 rows, costs at provider list prices.
Next step
GTM Goat runs exactly what you just read.