
AI is working its way through organizations, but not at an even pace. The promise of efficiency comes with caveats: check the work, watch for hallucinations. That is the nature of large language models, which generate the most probable answer rather than compute a verified one.
Those caveats weigh most on the functions that depend on accuracy: FP&A, forecasting, budgeting, staffing and pricing among others. In those departments trust comes at a premium. If the output cannot be trusted, every answer has to be checked, and the efficiency gain shrinks accordingly. The question is not only the potential return on AI. It is the potential cost of a mistake, which is why the cost of AI should be measured per correct decision rather than per token.
Input prices for capable AI models now span a 240-fold range, from about four cents to ten dollars per million tokens. Ten major models shipped from eight labs in eleven weeks this summer, and one frontier model announced this fall doubles its price after its introductory period. The cost of a decision is set less by which vendor a company chooses than by how the work is routed, where the numbers come from, and how many wrong answers get through.
In The Tokenomics Trap I made the case against single-vendor lock-in with SignalFlare production data: in a 90-day sample, routing pushed 99.93% of token volume to low-cost models at roughly 1/25th the cost of running everything through a single frontier vendor. That piece was about the bill. This one is about accuracy and what wrong answers cost.
The charts are from a session I am presenting to business and IT leaders. Prices and benchmarks are as of early October 2026 and will change within months.
The price range
Ten dollars buys about 1 million input tokens of GPT-6 Astra or about 238 million of Jev. Output tokens typically cost three to five times the input price, so the spread on a real workload is wider than the input prices show.

The bottom of the range is a different kind of model. Jev, from TypeSafe, reads language the way a generative model does but returns only one of the answers you define, with a calibrated confidence. It generates no text, and output is free. SignalFlare uses it for routing. It was released in early access on September 15, so the track record is short.
When the same task can cost a fraction of a cent or several dollars depending on where it is sent, routing is the largest cost lever in the system. A single-vendor contract does not use it.
Capability and effort
On the Artificial Analysis Intelligence Index, GLM-5.3 Flash scores 42 at $0.25 per index task. Claude Sonnet 5.5 at maximum effort scores 56 at $7.62. The line traces the best score available at each price, and it flattens as cost rises.

Effort is a setting on reasoning models that controls how much internal reasoning the model does before it answers. Those reasoning tokens are billed as output. Sonnet 5.5 scores 47 at high effort for $1.08 and 56 at maximum effort for $7.62: 9 points for about seven times the cost. For most business tasks, effort is a cost setting.
Rankings also change with the task. Open-weight models trail the leader by 13 to 19 points on general intelligence but by only 0.7 to 4.8 points on AIME 2026 competition math, in a vendor-run table. Math has checkable answers, so labs can train on it cheaply with automated grading, and smaller labs have concentrated there. A general benchmark score says little about whether a model is good enough for a specific job.
Generated numbers versus computed numbers
Generative AI applied probability to language and made machines fluent. A language model produces a number the same way it produces a word, by predicting what is most likely to come next. Business decisions carry a deterministic requirement: the same answer for the same input, traceable to its source.
Left to themselves, language models reason first, generate a figure along the way, and keep reasoning from it. The figure cannot be verified without recomputing it. The alternative is to route the calculation first: the model chooses the method, an engine computes the number from governed data, and the model explains the result.

The research points in one direction, though the studies differ in data, models and years. On financial documents, chain-of-thought reasoning got answers with decimals right 61.2% of the time with Llama 3 and 78.6% with GPT-4 Turbo. Having the model write the calculation as a program raised those to 79.9% and 91.7%. On FinSheet-Bench, published this year, accuracy on complex spreadsheet calculations fell to just above 30% for the top three models.
Database queries show the same pattern. In a 2023 study, GPT-4 writing SQL against raw enterprise tables answered 16.7% of business questions correctly. Querying through a knowledge graph raised that to 54.2%, and to 72% with an ontology-based check. A 2026 dbt Labs benchmark found 98.2% to 100% through a semantic layer on well-modeled data, against 84.1% to 90.0% for text-to-SQL. In SignalFlare’s work, queries through a defined schema with clear semantics and ontology have returned accurate results more than 99% of the time, against about 75% for a language model querying the same data.
The failure modes differ as well. In the dbt benchmark, the semantic layer failed with an error message and text-to-SQL failed with a plausible wrong number. A refusal is cheap to catch. A plausible wrong number is not. At the top of the ladder, pre-run machine learning results are validated by holdouts and backtests, and a model run once a night serves thousands of decisions for a few hundred tokens each.
Six layers of infrastructure
Computing numbers instead of generating them is one part of a larger structure. An AI system that makes or informs business decisions needs six layers around the models. Engineers call this structure a harness; applied to business decisions, it is a decision intelligence harness. The name matters less than what it does.

Know is governed data and context: metrics, entities and rules defined once in an ontology, so models are not left to guess what a number means. Route sends each task to the right model or engine. Compute produces the numbers with deterministic engines, with ranges and seeds so they can be reproduced. Verify and deliver recomputes and reconciles every number, keeps an audit trail, and writes results back only with approval. Learn keeps validated findings in memory the company owns, so outcomes improve the next answer and the routing. Govern is the frame around the other five: decision rights, guardrails, security, which models are allowed, and the tests that show a model is good enough.
Most of the work in this structure is done by deterministic code, statistics and people, not by a language model. Each layer either reduces the tokens a decision needs or reduces the chance that a wrong answer gets through, which is how the structure moves the cost per decision.
Cost per decision
The cost of a decision has more terms than the token bill:
Cost per decision = routing + inference + compute + verification + escalation + (probability of an undetected error × damage)
The first five terms are measured in cents. The last can be measured in millions. Take a 3% price increase to offset 3% inflation. Managed well, 90% or more of the price flows through to sales; the restaurant industry average is about 50%; and in recent years, poorly implemented price increases have cost chains as much as 3% of their guests.
In our pricing model for a 50-store chain with $100 million in sales, with food and other costs up 3% and labor up 4%, profit after the increase is about $3.6 million when it is managed well, $2.9 million at the industry average, and $1.9 million with a 3% guest loss. The gap between a well-run and a poorly run price decision is about $1.7 million, close to seven years of the AI bill for a routed design at full automation, estimated below.

On tokens alone, a five-cent model that fails 30% of the time costs about seven cents per success, against a dollar for a model that never fails. Sending 85% of tasks to the cheap model and the rest to a frontier model cuts cost from $7.62 to $1.36 per task. A checker that costs two cents and catches 90% of failures brings it to $1.95, still 74% cheaper. About 15 of every 1,000 tasks remain wrong and unnoticed, which is what verification and review are for.
The order of operations changes the cost as well. I estimated two designs for the same decision: a frontier model that reasons first and generates its numbers along the way, and a design that routes the calculation to an engine first and has a model explain the computed result. With assumptions set in favor of the reasoning-first design, routing math first came out about four times cheaper in tokens and 12 to 34 times cheaper in expected cost per decision, depending on what an error costs. Most of the difference came from fewer reruns and fewer undetected errors, not from tokens.
At $100 million in sales
While managers are asking questions, the AI bill is small under any design. The cost structure changes when agents run on their own, checking every store and item every day.

For an illustrative 50-store chain with $4 million in net profit, 10,000 analytical questions a month cost between $7,000 and $66,000 a year, depending on the design. At 250,000 automated runs a month, sending everything to a frontier model costs $1.65 million a year with Claude Opus 5.5, 41% of profit, or $870,000 with Sonnet 5.5. A routed design costs about $255,000, and about $180,000 once specialized models take over routine work. These are modeled estimates at October 2026 list prices.
The architecture decision has to come before the automation decision. A design that is affordable for questions becomes one of the largest variable costs in the business once agents run continuously.
Optionality
Between mid-July and the end of September, ten major models shipped from eight labs, open and closed, American and Chinese. No single lab led the period. The benchmark moved too: Artificial Analysis re-based its index in the same window, and GLM-5.3 Flash went from 57 at launch to 42.

A system wired to one model inherits that vendor’s price changes and cannot adopt the next model that does a job better or cheaper. Committing to a single model this early is a bet on a winner in a race that is still being run.

For the same illustrative chain, a single frontier model costs about 29 cents per automated run for as long as the vendor holds the price. A routed framework starts around eight and a half cents and falls toward six as routine work moves to specialized models trained on verified answers. Over three years, with automation growing fivefold, the estimates are about $1.6 million against $0.4 million.
Security narrows the choice before price does. Highly sensitive data, such as contracts or compensation, may only be allowed on a small model a company hosts itself. Masked operating data can go to open-weight models, or to frontier models under verified zero-retention terms. A single-vendor design applies one policy to everything. A routed design applies it per decision.
What is strategic
Models will keep changing, and their prices will keep moving. What a company can own is everything around them: its data, the definitions of what its numbers mean, its record of past decisions and outcomes, and the tests that show whether a model is good enough for a task. Those assets hold their value whichever model leads next quarter.
The order of work matters. Build tests for each type of task first, because every later choice depends on measured results. Trace cost, model and outcome for every workflow so cost per decision is visible. Route work to cheaper models only where a checker has been shown to catch most errors. Then train a specialized model on the highest-volume, most stable task. Business leaders, not IT alone, should set what an error costs for each type of decision, because that number determines how much checking each decision is worth.
Frontier models remain the right choice for novel and ambiguous problems. The question is where in each process they are invoked, and whether the system around them computes the numbers, catches the errors, and keeps the freedom to change models as the field moves.
· · ·
Sources: Vendor list prices, October 2026; Artificial Analysis Intelligence Index v4.3.2, October 2026; Thinking Machines AIME 2026 table, July 2026 (vendor-run); Improved LLM Agents for Financial Document Question Answering, arXiv 2506.08726 (2025); FinSheet-Bench, arXiv 2603.07316 (2026); Sequeda, Allemang and Jacob, arXiv 2311.07509 (2023); Allemang and Sequeda, arXiv 2405.11706 (2024); Ganz and Perigaud, dbt Labs, April 7, 2026; SignalFlare production routing data, 90-day sample, previously published in The Tokenomics Trap; cost scenarios and the pricing example are SignalFlare models of an illustrative 50-store, $100 million chain.
Mike Lukianoff is the founder and CEO of SignalFlare.ai.
Read the original post and subscribe for updates here.
Share



