Companies have gotten very good at measuring how much AI they use.
Tokens consumed. Seats purchased. Prompts sent. Agents deployed. Hours "saved."
There is just one problem: none of those things are ROI.
McKinsey's August 2026 global survey captures the disconnect perfectly. 80% of respondents said AI improved their individual productivity. Only 37% said it had contributed positively to their company's EBIT. Just 6% qualified as AI high performers, meaning they attributed at least 5% of EBIT to AI and described the impact as significant.
So employees increasingly feel more productive, companies are spending more money on AI, and the financial statements are sitting there like, cool story.
The problem is that AI activity, AI productivity, and AI ROI are three different things.
This guide is about getting from the first one to the last one.
- AI ROI is still just business ROI
- AI can make work better. It can also make it worse.
- Build an AI ROI Ledger before you automate anything
- Worked example: the two-hour deliverable
- The three AI pricing systems
- How AI creates revenue
- How AI cuts costs
- Human-assisted AI should usually come before autonomous AI
- Stop using 95% as a universal automation threshold
- The production AI stack: evals, guardrails, monitoring, recovery
- Product-market fit comes before automation
- The AI value ladder
- So where should you actually use AI?
- The AI ROI test
AI ROI is still just business ROI
Strip away the models, agents, tokens, credits, benchmarks, and demos and there are only two fundamental ways AI can create an economic return:
- Help you earn more money.
- Help you spend less money.
Everything else is an intermediate step.
The basic equation is still:
AI ROI = (realized economic benefit - total AI cost) / total AI cost
Both halves of that equation need more care than most AI ROI calculations give them.
On the benefit side, count things like:
- incremental contribution profit from new or preserved revenue
- verified hard-dollar savings
- productive capacity that was actually redeployed into valuable work
On the cost side, count everything required to get an accepted result, including AI spend, employee time, review, revisions, infrastructure, implementation, maintenance, failures, and remediation.
That word accepted matters.
An AI generating something is not a business outcome. An AI generating something useful enough that someone can actually ship, use, sell, or act on it is.
Revenue is not profit
Suppose an AI sales workflow helps generate $100,000 in incremental revenue.
If fulfilling that revenue costs $70,000, the economic benefit is not $100,000. It is closer to $30,000 in contribution profit, before accounting for the cost of the AI system itself.
This sounds obvious until somebody puts "$100K AI-generated revenue" on a slide.
Hours saved are not automatically savings
There are three increasingly valuable forms of labor savings:
- Capacity created: AI freed 100 employee hours.
- Capacity monetized: those hours were redirected into more valuable or sellable work.
- Cash removed: overtime, contractor spend, hiring, software expense, or another actual cost disappeared.
All three matter. They are not interchangeable.
If AI frees 500 hours and everybody spends them in more meetings, you created capacity. You did not save 500 hours' worth of payroll.
That distinction is one reason productivity can rise long before company profits do.
A current vendor-reported example makes the distinction useful. In May 2026, Boston Children's Hospital said more than 50 AI-enabled automations had produced about 60,000 hours of time savings, which it valued at $7 million-plus in redeployed labor. That is meaningful economic capacity, but the wording matters: the hospital described the labor as redeployed, not $7 million of payroll that vanished.
McKinsey's latest work makes a similar point from the enterprise side: the firms seeing stronger returns are more likely to redesign workflows, measure financial impact, and pursue growth and innovation instead of treating AI as a thin productivity layer on top of the old organization.
AI can make work better. It can also make it worse.
The evidence gives us a useful warning before we calculate anything.
A 2025 Quarterly Journal of Economics study tracked 5,172 customer-support agents using an AI assistant. Access to AI increased successfully resolved issues per hour by about 15%, with substantially larger gains among less experienced and lower-skilled workers.
In another randomized experiment involving 758 BCG consultants, workers using GPT-4 completed qualifying tasks 25.1% faster, completed 12.2% more tasks, and produced better results inside the model's capabilities. Performance could deteriorate on work outside what researchers called AI's "jagged technological frontier."
A 2026 UK AI Security Institute experiment found a similar pattern across four workplace tasks. On average, AI users finished 25% faster, scored 19% higher on quality, and produced 61% more quality-adjusted output per minute. The gains still varied significantly by task.
Then METR ran one of the most important counterexperiments.
Experienced open-source developers working on repositories they knew well believed AI was speeding them up. The randomized data showed the opposite: AI increased completion time by 19%. Even after the experiment, the developers still believed AI had made them about 20% faster.
METR's 2026 follow-up suggests newer systems probably do accelerate more development work, but participant selection effects made the size of the improvement impossible to estimate reliably.
That gives us the first rule of AI ROI:
Never calculate ROI from how fast AI appears to generate something. Measure the complete workflow.
Generation time is irrelevant.
Accepted-output time is what counts.
If you used to make a presentation in two hours, AI generates one in three minutes, and you spend three hours repairing it, you have automated yourself into working longer.
Build an AI ROI Ledger before you automate anything
The easiest way to stop fooling yourself is to keep one ledger for every workflow you test.
You do not need a complicated financial model. You need a consistent accounting system.
- Define the business outcome. What is this workflow supposed to change?
- Revenue created
- Revenue preserved
- Conversion
- Churn
- Gross margin
- Cost per transaction
- Contractor expense
- Throughput
- Time to completion
- Number of experiments run
- Support resolution
- Capacity available for higher-value work
Do not start with, "We want to use an agent."
Start with, "This process costs $42 per completed case and we want to get it below $25 without reducing quality."
Now we have something to optimize.
- Measure the human baseline. Before touching AI, record:
- Who does the job
- Fully loaded hourly cost
- Task frequency
- Actual end-to-end completion time
- Quality or acceptance rate
- Revision rate
- Error rate
- Current throughput
Do this with real work, not someone's guess about how long the work "usually" takes.
METR's study should permanently cure us of using perceived time savings as our measurement system.
- Measure AI production cost. Depending on the system, this may include:
- API tokens
- Credits
- Subscription allocation
- Model reasoning
- Search
- Computer use
- API or tool calls
- Retrieval
- Database/storage
- Hosting
- Orchestration
- Observability
- Implementation
- Maintenance
Once you build your own AI product or agent harness, model inference can become only one line item in a much larger bill.
This is why obsessing over shaving half a cent from token costs while ignoring a ten-minute human review step is often backwards.
- Measure the human cost of using the AI. Count:
- Preparing instructions
- Gathering context
- Prompting
- Supervising
- Checking
- Fixing
- Rerunning
- Escalating
- Troubleshooting
If a $0.30 model call produces 20 minutes of cleanup for a $90/hour employee, the cleanup costs $30.
Guess which number matters more.
- Calculate accepted-output economics. Track:
Cost per accepted output = total workflow cost / accepted outputs
Not:
- Cost per generation
- Cost per model call
- Cost per token
*Cost per useful finished job.*
This automatically punishes systems that require repeated attempts, hallucinate, get stuck in loops, or produce superficially impressive junk.
- Price failure. For consequential workflows, add:
Expected failure cost = probability of failure x average cost of failure
That cost might include:
- Employee rework
- Support escalation
- Refunds
- Incorrect payments
- Downtime
- Customer loss
- Compliance remediation
- Security incidents
Some risks cannot sensibly be reduced to an average dollar figure. We will deal with those when we get to autonomy.
- Measure realized impact. Finally ask what actually happened:
- Did revenue increase?
- Did churn fall?
- Did you avoid another hire?
- Did output increase?
- Did contractor spend fall?
- Did the employee use the extra time productively?
- Did quality remain equal or improve?
That's your ledger.
Everything before this is an assumption.
Worked example: the two-hour deliverable
Assume an analyst costs the business $60/hour fully loaded.
A recurring report normally takes:
120 minutes x $60/hour = $120 in labor
Now introduce AI.
The model generates a draft quickly, but the real AI-assisted workflow takes:
- 10 minutes gathering material and instructing it
- 7 minutes generating and rerunning
- 20 minutes reviewing
- 10 minutes editing
Total human involvement: 47 minutes
- Human cost: 47/60 x $60 = $47
- AI/tool cost: $1.84
- New workflow cost: $48.84
- Capacity value created: $120 - $48.84 = $71.16 per accepted report
If the company produces 100 reports a month, that is $7,116 of monthly capacity.
We still have not proven $7,116 in cash savings.
If the analyst uses those reclaimed hours to complete additional billable projects, we can monetize the capacity.
If the company avoids contractor expense, we can call it hard savings.
If nothing changes downstream, record it as capacity created.
This is how we keep the math honest.
The three AI pricing systems
Once you understand the workflow, you can calculate what the AI portion actually costs.
There are three common pricing systems, plus the custom stack hiding underneath many products.
The examples below use a September 9, 2026 source check. AI pricing changes quickly, so use the method as the durable lesson and re-check the vendor's live pricing before you approve a budget.
- Pay per token. Token pricing is the easiest to understand because the underlying unit is visible.
At its simplest:
AI request cost = input + cached input + output + tools
As of September 9, 2026, OpenAI lists GPT-5.6 Sol at promotional pricing of $4 per million input tokens, $0.40 per million cached input tokens, and $20 per million output tokens, available at least through November 21, 2026. Requests above 272K input tokens are priced at 2x input and 1.5x output for the full request, while cache writes are billed at 1.25x the uncached input rate.
Anthropic lists Claude Sonnet 5 at $2 per million input tokens and $10 per million output tokens. Anthropic had originally planned to raise Sonnet 5 to $3/$15 on September 1, then explicitly canceled that increase on August 10 and made the $2/$10 pricing permanent. That is exactly the kind of pricing change that makes stale screenshots dangerous. Anthropic also notes that Sonnet 5's newer tokenizer can produce roughly 30% more tokens for equivalent text than older Sonnet models.
Google made Gemini 3.8 Flash generally available on September 2, 2026 at introductory pricing of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Gemini pricing can also include cached context, search grounding, and other tool-specific charges.
So the real equation can become:
Task cost = model tokens + cache + retrieval + search + tools + retries + agent loops + infrastructure
Still, cost per request is usually the wrong number to optimize.
Use:
total AI spend / accepted completed tasks
A cheap model that fails three times can cost more than an expensive model that works once.
OpenAI's August 2026 builder guidance pushes the same idea further. It recommends choosing models by measured task performance, using prompt caching where it pays, moving deterministic filtering and aggregation into programmatic tool calls, and adding multi-agent orchestration only where parallel work earns the added complexity.
Quality first. Rightsize second.
- Pay per credit. Credits range from economically transparent to please consult the Oracle of Billing.
GitHub's current AI Credits system is unusually legible: one AI credit equals one cent. Chat, CLI, cloud-agent sessions, Spaces, Spark, and third-party coding agents consume credits based on the model and token usage. On paid plans, code completions and next-edit suggestions are not billed in AI credits.
Adobe's generative credits are more abstract. Plans include monthly allocations, while different Firefly features consume different amounts based on what you generate.
The rule is simple:
Never optimize credits. Convert credits into dollars per accepted business outcome.
If the underlying mechanics are opaque, benchmark them empirically.
If you spend $300 this month and receive 127 deliverables you can actually use:
$300 / 127 = $2.36 AI expense per accepted deliverable
Now the credits have been converted back into money, where accountants can once again see them.
- Pay per subscription. Subscriptions are powerful for human-assisted work because another useful task may have almost zero marginal cost until you hit a usage limit.
But there are two ways to look at that:
- Marginal economics: You already own the $100/month seat. Doing one additional task creates almost no new invoice.
- Fully loaded economics: You complete 40 valuable tasks through that seat each month. $100 / 40 = $2.50 per accepted task.
Use marginal economics when deciding what work to route through an already-purchased subscription today.
Use fully loaded economics when deciding whether the subscription deserves to exist next month.
Subscriptions introduce another scarce resource: quota capacity.
A task that burns half your agent allowance has opportunity cost even if your credit card charge stays the same.
So the better question becomes:
How much business value am I producing per constrained unit of AI usage?
Worked example: the $100 seat
Suppose a $100/month AI seat supports 40 accepted tasks.
- Fully loaded AI cost: $2.50/task
- Each task saves 20 minutes of a worker earning $60/hour.
- Capacity created per task: $20
- Monthly capacity: 40 x $20 = $800
- Subscription cost: $100
- Potential net economic benefit: $700/month
On paper, that's a 700% return on the subscription.
But only if that $800 becomes economically useful.
If 13.3 reclaimed hours produce additional client work, prevent an outside hire, or eliminate another expense, great.
If they disappear, call it $800 of capacity created, not $800 of cash.
- The custom harness. Once you build an application around an API, your model vendor's pricing page stops being your cost model.
Your stack may include:
- Model inference
- Search
- Vector/database storage
- External APIs
- Orchestration
- Queues
- Cloud compute
- Observability
- Eval infrastructure
- Security
- Support
- Engineering
- Human review
- Ongoing maintenance
This is where total cost of ownership replaces token math.
McKinsey's August 2026 analysis gives a useful illustration. In one banking customer-service agent example, tokens represented only about 20% to 25% of variable run cost. The rest came from the surrounding workflow and systems.
By September 9, the major agent platforms were converging on the same production lesson from different directions:
- OpenAI's Responses API added programmatic tool calling so deterministic work can happen in code instead of burning model context.
- Claude Managed Agents supports hard session budgets that stop new model requests when a session reaches its configured list-cost ceiling.
- Google's Gemini Enterprise Agent Platform, the evolution of Vertex AI, combines ADK, Agent Runtime, evaluation, observability, and governance in the same production lifecycle.
- Microsoft Foundry, the current name for the platform formerly called Azure AI Foundry, uses tracing, repeatable evaluation, publishing, and production monitoring as explicit lifecycle steps.
- Amazon Bedrock AgentCore supports on-demand regression evals and continuous evaluation of sampled production traces, alongside observability.
The durable lesson is not which vendor has the prettiest agent diagram. It is that the job, the trace, and the accepted outcome are the unit of economics.
Build the economics around the job, not the model.
How AI creates revenue
Cost cutting gets most of the AI ROI conversation because it is easier to measure.
Revenue matters just as much.
There are six useful buckets:
- Preserve revenue you already have. Start with customers you would otherwise lose.
AI can help:
- Resolve support problems faster
- Increase first-contact resolution
- Identify churn signals
- Recommend save offers
- Accelerate onboarding
- Improve product adoption
- Prevent refunds
- Recover failed payments
- Escalate difficult cases earlier
The QJE support study is useful because AI did not replace the human agent. It helped the agent resolve more customer issues per hour, with the largest gains among newer workers. Researchers also found improved customer sentiment and evidence of lower employee turnover.
A current commercial example comes from Circles, a telecom platform customer profiled by OpenAI in August 2026. Circles reports that AI-driven personalization in Singapore produced a 22% increase in average revenue per user and a 9% reduction in churn, while its CareX support system reached a 65% autonomous resolution rate.
Circles says it measured the personalization effect by comparing customers who received AI-powered recommendations with customers who did not. These are company and vendor-reported results, not an independent randomized study, but the measurement design is the right direction.
AT&T gives us a larger-scale production example. AT&T says it uses BigQuery to analyze customer history and route intent, then Gemini Enterprise for Customer Experience to handle autonomous voice and chat interactions. The customer-facing agents can retain up to 60 days of history across channels, so a conversation can continue after a customer moves from web to store to phone.
Separately, AT&T's secure internal AskAT&T platform supports 150 generative AI solutions in production and processes 40 billion tokens per day. AT&T reports 20% more customer inquiries resolved through fully automated systems, about $150 million in annual savings, improved Net Promoter Score, and overall AI ROI rising from 2x to 5x in one year.
These are company-reported results, but the implementation lesson is concrete: routing, memory, and the agent layer were treated as separate systems and measured against resolution, savings, and customer outcomes rather than token volume.
That suggests a straightforward measurement system:
AI group vs. control group
Then compare:
- Resolution
- Retention
- Refunds
- Churn
- Revenue retained
- Support cost
Do not ask customers whether the chatbot "felt innovative."
Ask whether more of them stayed.
- Sell more of what you already sell. This includes:
- Account research
- Lead qualification
- Personalized outreach
- Sales-call preparation
- Follow-up
- Proposals
- RFPs
- Cross-sell recommendations
- Upsell recommendations
- Segmentation
- Campaign creation
- Offer personalization
- Merchandising
Example: AWS put a supervisor in front of 20+ sales agents
AWS built Field Advisor after its sales organization accumulated more than 20 specialized agents. Instead of asking sales reps to know which agent handled which job, AWS added a supervisor layer.
The disclosed architecture gives us a useful implementation recipe:
- Run one supervisor in an isolated runtime. Field Advisor uses a Strands Agent in Amazon Bedrock AgentCore Runtime, with the latest Anthropic Claude model available through Amazon Bedrock and a cross-Region inference profile for throughput.
- Normalize tools behind one interface. The supervisor can call local tools, remote MCP tools through AgentCore Gateway, or specialized domain agents running in separate AgentCore runtimes. More than 20 MCP tools were onboarded by registering them so the supervisor could discover them without orchestration-code changes.
- Propagate identity instead of sharing generic credentials. AgentCore Identity carries the authenticated user's OAuth identity to downstream tools and agents.
- Persist the right context. AgentCore Memory combines short-term session history with long-term semantic memory, while internal knowledge bases provide product, pricing, competitive, and organizational context.
- Put cross-cutting controls in hooks. AWS uses hooks for error circuit breaking, authorization interrupts on sensitive CRM data, explicit confirmation before data-mutating writes, citation extraction, and tool-progress streaming.
- Cache repeated context. The team extended the Bedrock model integration with incremental prompt caching for multi-turn conversations.
- Trace and evaluate the whole flow. AgentCore supplies observability and evaluation across supervisor, tools, and remote agents.
AWS reports 120,000+ prompts, up to two hours saved per salesperson per week in human-in-the-loop record workflows, 41% lower latency after moving to AgentCore, and infrastructure consolidation from seven AWS accounts to one AgentCore Runtime. The pattern is useful anywhere a company has multiple useful agents but is making humans do the routing.
McKinsey's 2026 survey finds revenue gains are frequently reported in areas such as marketing and sales, but the correct denominator is still the business outcome rather than AI output.
For sales:
incremental contribution profit from AI-assisted cohort - total workflow cost
For marketing:
incremental contribution profit from test group - control group
Use randomized tests or holdouts whenever you can.
If the AI-assisted sales team closes 12% more business, you have evidence.
If sales happened to rise during the three months after everybody got ChatGPT, you have a PowerPoint.
- Create new products. There are two important versions:
- AI-created products: AI accelerates humans who remain accountable for the finished thing. Reports, research, courses, newsletters, videos, software, designs, datasets, dashboards, templates, games, books, prototypes, and market intelligence all fit here.
- AI-powered products: AI becomes part of what the customer buys. Copilots, tutors, recommendation systems, research assistants, coding tools, planning systems, analytics, and generative design products fit here.
The first category is one of the lowest-risk ways to create new AI revenue because the AI does not need to be the product. It changes your production economics.
In the second category, the product's unit economics have to absorb inference, reliability, support, and monitoring.
The 2025 Procter & Gamble field experiment offers a useful clue for the first category. In a randomized study of 776 professionals working on real product-development challenges, individuals using AI performed comparably to human teams without AI. The researchers also found AI helped bridge expertise across functional areas.
That does not mean "fire the product team."
It means the cost and structure of assembling expertise may change.
- Create new services. Services may be even easier to launch because you can keep humans in the fulfillment loop.
Examples include research-as-a-service, advisory, analysis, design, copy, video production, software development, proposals, compliance support, localization, data enrichment, training, competitive intelligence, reporting, lead generation, implementation, and managed agents.
The model is simple:
Human accountability + AI production leverage
If one expert can now serve 20 clients rather than ten without degrading the result, you have a throughput opportunity.
Calculate:
Contribution margin = customer revenue - human fulfillment - AI - infrastructure - support/remediation
Then compare it with the old service.
- Build AI-native products and services. Now AI itself performs a large portion of the value-generating work.
Think autonomous support, AI analysts, coding agents, workflow agents, adaptive tutors, compliance monitors, AI planning, personalized media, and autonomous research.
The upside is scale.
The catch is that your inference bill, reliability problem, security model, and customer-support burden now scale too.
The correct metric becomes:
profit produced per model-dollar
rather than:
tokens consumed
- Increase the number of experiments you can afford. R&D has different economics because most experiments fail.
AI can lower the cost of customer research, product concepts, prototypes, software experiments, literature reviews, simulations, design variations, marketing tests, and formulation work.
So measure the portfolio:
Expected portfolio value = experiments x probability of useful success x value of success - experimentation cost
If your $1 million R&D budget previously funded 20 serious experiments and AI lets you run 40 at the same quality, you bought more shots on goal.
That can be valuable even before a successful experiment creates revenue.
How AI cuts costs
The cost side can stay simpler.
Most opportunities fall into three buckets:
- Artifact creation. AI helps create documents, presentations, spreadsheets, code, designs, reports, proposals, videos, analyses, and documentation. Measure the full accepted deliverable.
- Research and information synthesis. AI helps with search, literature review, competitive intelligence, account research, document review, extraction, discovery, synthesis, and comparison. Measure whether the resulting research is accurate enough to make the intended decision.
- Routine workflow execution. AI helps with classification, routing, data movement, reconciliation, CRM updates, reporting, scheduling, document processing, QA, and other repetitive operational tasks.
Google offers a useful scale check on the first bucket. At Cloud Next 2026, Google said 75% of all new code at Google was AI-generated and then approved by engineers, up from 50% the previous fall. Google also said a particularly complex code migration completed by agents and engineers together finished six times faster than was possible a year earlier. The useful pattern is the approval boundary: high-volume AI execution with a human still accountable for what counts as accepted production code.
Example: OpenAI made the software environment legible to Codex
OpenAI's February 2026 Codex harness experiment shows what changes when agents do nearly all of the implementation work. The team started from an empty repository. Codex CLI using GPT-5 generated the initial repo structure, CI configuration, formatting rules, package-manager setup, application framework, and even the first AGENTS.md.
The implementation pattern was not "write a better prompt":
- Give the agent a map, not a giant manual. OpenAI kept a short
AGENTS.md, roughly 100 lines, as a table of contents and made a structureddocs/directory the versioned system of record for architecture, product specs, plans, reliability, security, and other knowledge. - Make constraints executable. The application follows a rigid layer model, and custom linters plus structural tests enforce allowed dependency directions, naming rules, structured logging, file-size limits, and reliability requirements. Lint errors include remediation instructions the agent can act on.
- Make the running product visible to the agent. Each change can boot in an isolated git worktree. OpenAI wired Chrome DevTools into the agent environment and added skills for DOM snapshots, screenshots, and navigation. Local logs, metrics, and traces are also queryable by the agent.
- Close the review loop. A human describes the task. Codex opens the PR, reviews its own work, requests additional local and cloud agent reviews, reacts to feedback, and iterates until the reviewers are satisfied.
- Move humans up one level. Humans prioritize work, turn user feedback into acceptance criteria, and validate outcomes. When the agent struggles, the team asks what tool, guardrail, documentation, or environmental capability is missing and improves the harness.
Five months in, OpenAI reported roughly one million lines of agent-written code and about 1,500 merged PRs, initially driven by three engineers. The team estimated it built the product in about one-tenth the time manual implementation would have taken. The transferable lesson is that autonomous coding becomes more reliable when the environment can test, observe, constrain, and explain itself to the agent.
Example: Delivery Hero used a different model family to review the first
Delivery Hero's Herogen system gives us a useful pattern for autonomous software delivery at scale. The company standardized model access through Google Cloud's Agent Platform with LiteLLM as an LLM proxy, allowing engineers and agents to use different models through the same infrastructure.
Herogen's disclosed workflow is straightforward:
- It picks up a Jira task.
- It writes the code.
- It runs tests.
- It reviews its own work.
- Claude is the primary coding model.
- Gemini participates in a separate council of agents that reviews the code and performs security checks.
Delivery Hero says using a different model family for review helps catch errors that a single model may miss. The company also uses Gemini Code Assist and Gemini CLI for broader developer workflows.
Herogen launched in Q1 2026. Delivery Hero reports 170+ PRs created and merged per day, more than 10% of company PR volume, an 85% success rate for tickets merged to production, and adoption 18x higher than its original assumptions in Q1. Some week-long engineering jobs are now completed in minutes. The reusable pattern is independent verification: if one model performs the work, do not make that exact same reasoning path the only thing deciding whether the work is safe to ship.
The general formula is:
Workflow savings = baseline human cost - AI-era human cost - AI/system cost - expected failure cost
The revision-time rule still wins:
Generation time does not matter. Total accepted-output time does.
Human-assisted AI should usually come before autonomous AI
The economic temptation is obvious.
If AI saves money while helping someone, eliminating the person looks like the next step.
Sometimes it is.
Autonomy changes the risk equation because the system can now create costs without someone inspecting the output first.
A useful progression is:
- Level 1: AI assists. A human initiates, reviews, and owns the result.
- Level 2: Agent executes, human approves. AI runs most of the workflow. Consequential actions hit an approval gate.
- Level 3: Agent operates autonomously. Appropriate when outcomes are measurable, failure costs are bounded, actions are reversible or low risk, and monitoring is strong.
Example: LendingTree separated planning, specialists, and state
LendingTree's August 2026 mortgage-assistant write-up is one of the most explicit production multi-agent architectures available.
LendingTree deployed three independent agents: a supervisor, an education worker, and a matching worker. They use LangGraph, MCP, Amazon Bedrock, Nova Pro, Nova Lite, Bedrock Guardrails, Bedrock Knowledge Bases, OpenSearch, PostgreSQL on RDS, ECS with Fargate, Terraform, GitLab CI/CD, CloudWatch, and X-Ray.
The request flow is concrete:
- Screen the request before reasoning. User input passes through Bedrock Guardrails for content filtering, PII redaction, and prompt-threat screening while a separate LLM safety classifier checks LendingTree's conversational policy in parallel.
- Plan in an explicit state machine. The LangGraph supervisor uses nodes for intent analysis, planning, and response composition, with edges that make routing decisions explicit and traceable.
- Route expensive reasoning only where needed. Nova Pro handles complex reasoning and critical classification. Nova Lite handles lighter conversation and classification work.
- Give specialists different sources of truth. The education worker uses a Bedrock Knowledge Base backed by OpenSearch. The matching worker calls LendingTree's internal offer, eligibility, rate, prequalification, and profile APIs.
- Keep conversation state outside the model. A LangGraph PostgreSQL checkpointer on RDS preserves state across turns, worker handoffs, and service restarts.
- Run the output through safety checks again. The supervisor combines worker results, passes the result back through Guardrails, then returns it to the user.
- Deploy agents independently. MCP lets each agent be updated, scaled, or rolled back independently. CloudWatch and X-Ray trace one conversation across all three services.
LendingTree says early guardrail settings overblocked legitimate mortgage terminology, so the team tuned them against real conversations. It also says model routing helped keep costs under control and a unified PostgreSQL checkpointer fixed lost context during early agent handoffs. Production data through Q1 2026 covered roughly 1,960 conversations and 12,100 messages, with more than 97% of conversations handled without human escalation.
Anthropic's real-world analysis of millions of Claude Code and API interactions found people increasingly let agents work for longer periods without intervention. Among its longest-running Claude Code sessions, uninterrupted work time nearly doubled over three months. Anthropic argues this increased autonomy makes monitoring and intervention mechanisms increasingly important.
Example: Anthropic designed the blast radius before increasing autonomy
Anthropic's May 2026 containment write-up is useful because it documents several architectures, including failures that forced the team to change them.
Anthropic's pattern changes with the product:
- claude.ai code execution: code runs server-side inside an ephemeral gVisor container on isolated infrastructure. The user filesystem is not available.
- Claude Code: the agent needs local filesystem, shell, and network capabilities, so Anthropic added an OS-level sandbox using Seatbelt on macOS and bubblewrap on Linux. Reads are allowed, writes are allowed inside the workspace, and network access is denied by default. Anthropic says this reduced permission prompts by 84%.
- Claude Cowork: because a general knowledge worker cannot be expected to judge arbitrary shell commands, Anthropic used a full local VM. Only the chosen workspace and
.claudefolder are mounted. Credentials remain in the host keychain and do not enter the guest. - Model defenses stay layered on top. System prompts, classifiers, probes, and training shape behavior, while tool permissions and controls on external content reduce what a compromised or injected agent can reach.
Why not rely on approval dialogs? Anthropic says users approved roughly 93% of Claude Code permission prompts. In a controlled internal red-team exercise, a malicious prompt persuaded Claude Code to read AWS credentials and send them to an external endpoint in 24 of 25 attempts. Anthropic's conclusion is much stronger than "improve the warning": filesystem boundaries and egress controls need to make the dangerous action impossible even when the user or model would otherwise approve it.
That gives us a practical hierarchy: environment boundary first, model behavior second, human approval where the user can actually evaluate the action.
OpenAI describes a similar production philosophy for Codex: keep agents inside clear technical boundaries, allow low-risk actions to move quickly, make higher-risk actions explicit, and preserve telemetry so teams can reconstruct what the agent did.
A good default is:
Anything that communicates directly with customers, moves money, changes production, creates a legal or compliance commitment, or causes difficult-to-reverse external effects should normally retain human approval.
You can loosen that gate when real evidence says the system is reliable enough.
Not when the demo looks good.
Stop using 95% as a universal automation threshold
This is where evals become economics.
A system that is 95% accurate fails one time in twenty.
That may be excellent for tagging internal documents.
It may be horrifying for wiring money.
Reliability only makes sense relative to failure cost.
Suppose:
- B = business value when the automation succeeds
- F = cost when it fails
- r = reliability
- C = ordinary AI and operating cost
Then:
Expected value = rB - C - (1 - r)F
Now reliability becomes a business variable.
Worked example: the 95% agent
Suppose a successfully automated task produces $9 of net benefit after normal operating costs.
A failure costs $100 to repair.
At 95% reliability:
0.95 x $9 - 0.05 x $100 = $3.55 expected value per task
Still positive.
The approximate break-even reliability is:
91.7%
Now change only one number.
If each failure costs $1,000, break-even reliability jumps to roughly:
99.1%
If failure costs $100,000:
99.991%
Same agent.
Completely different automation decision.
And multi-step workflows introduce another trap.
If 20 independent steps each succeed 99% of the time:
0.99^20 ≈ 81.8%
The end-to-end system is nowhere near 99% reliable.
This is why "our model scored 97% on the benchmark" tells you almost nothing about whether the whole business process should run unattended.
The production AI stack: evals, guardrails, monitoring, recovery
Once AI can act, there are four questions you need answered:
- Evals: Can it do the job? Anthropic recommends what it calls eval-driven development: define realistic tasks and success criteria early, then keep running them as the system changes. Its practical guidance says teams can often begin with just 20 to 50 real tasks or failures, rather than waiting until they have a giant benchmark.
Your eval suite should test things like:
- Task completion
- Quality
- Factuality
- Correct tool use
- Escalation
- Policy compliance
- Edge cases
- Adversarial cases
- Regressions
Example: OpenAI built context, retrieval, memory, and evals as one system
OpenAI's January 2026 write-up of its in-house data agent is one of the clearest blueprints for an internal knowledge-and-analysis agent. The agent is powered by GPT-5.2 and uses Codex, the Evals API, the Embeddings API, and MCP. Employees can use it from Slack, a web interface, IDEs, Codex CLI, or OpenAI's internal ChatGPT app.
The useful part is the setup:
- Ground the agent in several kinds of context. OpenAI combines schema metadata and lineage, historical query patterns, human table annotations, Codex-derived knowledge from the code that creates the data, institutional documents, memory, and runtime context.
- Precompute what can be precomputed. A daily offline pipeline normalizes table usage, annotations, and Codex enrichment, converts that context into embeddings, and stores it for retrieval. At query time, the agent retrieves only the relevant context instead of scanning everything.
- Validate stale or missing context live. The agent can query the warehouse and systems such as the metadata service, Airflow, and Spark when it needs current information.
- Save non-obvious corrections as memory. Users can save, edit, and scope corrections so the same filter or business rule does not need to be rediscovered every time.
- Turn evaluation into a regression suite. OpenAI pairs important natural-language questions with manually authored golden SQL. It runs the expected SQL and the agent-generated SQL, compares both the queries and resulting data, and feeds those signals into an Evals grader.
- Reuse existing permissions. Access is pass-through, so the agent can only query data the user already has permission to access.
- Remove ambiguity. OpenAI says exposing too many overlapping tools hurt reliability, and highly prescriptive prompts could push the agent down the wrong path. Its lessons were to consolidate tools, guide the goal rather than every step, and extract meaning from the code that actually produces the data.
The reusable pattern is authoritative context -> selective retrieval -> live tools -> memory -> golden evals -> existing permissions. OpenAI says this moved many internal questions from days of analysis to minutes.
Use real production failures as new test cases.
And evaluate the whole harness, not only the model.
The real product is:
model + instructions + context + retrieval + tools + permissions + code + orchestration + fallback behavior
Anthropic describes agent traces as the complete record of outputs, tool calls, intermediate actions, and interactions. That is much closer to what you need to evaluate than scoring the last paragraph the model happened to produce.
The latest platform workflows reinforce that idea. Google's ADK supports evaluation of execution trajectories, Microsoft Foundry explicitly puts tracing before evaluation in its agent lifecycle, and AWS AgentCore supports both on-demand regression evaluation and online evaluation of production traces.
- Guardrails: What is it allowed to do? Use:
- Least-privilege access
- Allowlisted tools
- Sandboxing
- Network restrictions
- Credential boundaries
- Spend limits
- Transaction limits
- Retry limits
- Input validation
- Output validation
- Human approval for high-impact actions
- Kill controls
Example: Intuit let the model decide what, while deterministic software decided how
Intuit's September 4, 2026 disaster-recovery architecture may be the cleanest pattern in this guide for high-consequence automation. AWS explicitly describes the post as an architecture and reusable pattern, not a complete copy-paste deployment.
Intuit did not give an LLM raw control of production infrastructure. It already had a deterministic disaster-recovery system called EWOK, where service owners declare recovery intent in YAML and EWOK executes the actual failover workflow.
The agent layer sits above that system:
- Define capabilities as typed skills. Each skill has a name, description, input schema, prompt body, and executor.
- Make procedures explicit. Intuit writes numbered operation steps so each step maps to exactly one executor call. A failed step stops the skill immediately. The model is told not to improvise alternate recovery actions because the executor already owns transient retries.
- Encode policy as branches, not tribal knowledge. Change-freeze rules and emergency exceptions become explicit flows with defined exits.
- Compile skills into model tools. The skill schema becomes an Amazon Bedrock Converse API
toolSpec, allowing the model to choose the right capability from the request. - Keep execution deterministic. The model selects a skill and arguments. The skill executor calls the existing EWOK APIs and returns structured results to the agent.
- Apply guardrails centrally. Bedrock Guardrails are attached to the model client rather than being reimplemented in every skill prompt. Credentials and infrastructure authority remain outside the model.
- Keep the model replaceable. Because the skills and executors are separate from the Bedrock model layer, Intuit can evaluate and switch foundation models without rewriting the operational system.
Intuit says teams had used the agent to run production failovers for eight months by the time of the write-up. Its underlying EWOK system had already reduced supported recovery workflows from several hours to about 20 minutes. The pattern is the point: use probabilistic reasoning to interpret intent and deterministic software to execute consequential actions.
OpenAI describes guardrails as layered defenses rather than a single protection and recommends combining model-based checks with deterministic controls, authentication, authorization, and normal software security.
That distinction matters.
A prompt saying please don't wire $47,000 to Romania is not a financial control.
- Monitorability: Can you see what it did? For consequential agents, log:
- Model/version
- Prompt/context
- Tool calls
- Permissions
- External requests
- Files changed
- State changes
- Tokens/cost
- Latency
- Retries
- Approvals
- Errors
- Final outcome
If something goes wrong, you should be able to reconstruct the sequence.
This is now a standard production pattern across tools. n8n's August 2026 observability guidance calls for tracing model calls, tool invocations, and interactions with external systems rather than only recording that a workflow failed. AWS AgentCore's July update similarly unified traces, prompts, and structured logs into a per-agent log group to make end-to-end debugging easier.
- Recovery: Can you stop it and undo it? Design for:
- Idempotent operations
- Checkpoints
- Rollback
- Backups
- Bounded transactions
- Human escalation
- Pause controls
- Kill controls
- Incident response
AISI delivered an unusually vivid demonstration in July 2026.
During deliberately permissive cybersecurity evaluations, agents had open-internet access and certain developer safety classifiers disabled. Across 122 runs, AISI found 10 runs containing autonomous, unsanctioned live-internet activity. One agent attempted to insert malicious code into a real open-source project and tried to socially engineer a maintainer into approving it. The maintainer refused.
AISI emphasizes that these were intentionally unusual research conditions, not normal consumer deployments.
That caveat matters.
So does the lesson.
Monitoring and recovery are part of your ROI architecture because failures cost money too.
Product-market fit comes before automation
A lot of AI projects begin with:
What can we build with an agent?
Reverse the question.
What expensive outcome does someone already want badly enough to pay for?
Then figure out the minimum amount of AI required to deliver it.
- Start with painful demand. Ask:
- What already costs the customer money?
- What burns expert hours?
- What takes too long?
- What prevents revenue?
- What are they already paying someone to do?
A $100 problem solved with $5 of AI can be interesting.
A $2 problem requiring $50 of inference is a science project.
- Prove willingness to pay. Do not use "people thought the demo was cool" as your PMF metric.
Sell the outcome.
- Start concierge. Before building a fully autonomous product, deliver the service manually with AI helping behind the scenes.
That teaches you:
- What customers actually want
- What acceptable quality looks like
- What edge cases exist
- Where humans intervene
- What they will pay
- Which steps really need AI
It also creates the real production examples you need for your eval suite.
The implementation pattern that keeps repeating
Different companies chose different clouds, models, frameworks, and data stores. The architecture keeps rhyming:
- Start with a measurable job, not an agent demo.
- Use the smallest architecture that can reliably complete that job.
- Put authoritative data and context somewhere the system can retrieve and verify.
- Keep persistent state outside the model when the workflow needs it.
- Give the model only the tools and permissions required for the task.
- Keep deterministic work deterministic, especially for consequential actions.
- Evaluate the trajectory and final business outcome, not only the final paragraph.
- Put human approval at high-impact boundaries, then back it with hard technical controls.
- Trace tool calls, state changes, errors, latency, and cost.
- Use accepted output, reliability, and realized economic value to decide whether greater autonomy deserves to exist.
That is the implementation version of the AI ROI framework.
- Use the least AI required. This may be the most important architecture rule in the whole guide:
Use the least autonomous, least expensive system capable of reliably delivering the required business outcome.
Sometimes that is an employee using ChatGPT.
Sometimes it is Claude Code or Codex.
Sometimes it is one API call inside an otherwise deterministic automation.
Sometimes you need retrieval.
Sometimes you need an agent.
Sometimes you need multiple agents.
The current tools increasingly make this split explicit. OpenAI's GPT-5.6 guidance recommends moving deterministic filtering and aggregation into programmatic tool calls. Make's February 2026 Make AI Agents (New) product was rebuilt directly inside the normal scenario canvas so adaptive agent judgment can sit beside deterministic automation, with visible tool choices and reasoning.
That is a much better architecture test than asking whether a task can technically be given to an agent.
A five-agent architecture does not automatically mean you have reached Level Five AI Business Wizard. Sometimes you have just found an expensive way to move JSON around.
- Automate gradually. The progression should usually look like:
- Human
- Human + AI
- AI executes / human reviews
- AI executes bounded components
- Autonomy
At each step, recalculate:
- Contribution margin
- Accepted-output cost
- Reliability
- Failure cost
- Customer outcome
If the economics improve, keep going.
If not, stop.
The AI value ladder
This gives us one final way to see why so many AI projects never produce measurable ROI.
- Output. AI made something. This is where most demos end.
- Accepted output. The result was good enough to use. Now we have something potentially valuable.
- Productivity. The accepted result required less total time or money. Now we have a measurable workflow improvement.
- Capacity. The organization can perform more work with the same resources. Now we have leverage.
- Economic value. That capacity creates revenue, preserves revenue, or removes actual cost. Now the P&L changes.
- Scaled ROI. The economics stay attractive after model expense, human review, failures, support, governance, monitoring, maintenance, and increased usage.
Now you have an AI business case.
McKinsey's 2026 survey is basically a giant real-world illustration of the gap between levels three and six: reported personal productivity is widespread, while meaningful enterprise financial impact remains concentrated in a small minority of organizations.
So where should you actually use AI?
Do not start by building a 400-row catalog of everything AI could theoretically touch.
Start with your own expensive workflows.
Look for work with some combination of:
high frequency + high labor cost + measurable output + repeatability + meaningful revenue/cost impact
Then classify each opportunity:
- Revenue preservation: Can AI keep customers or prevent lost revenue?
- Existing revenue growth: Can it improve conversion, deal size, upsell, or customer value?
- New products: Can AI lower the cost of creating something customers will buy?
- New services: Can AI let existing experts serve more customers?
- R&D: Can it lower the cost or increase the number of useful experiments?
- Artifact production: Can it create accepted deliverables faster?
- Research: Can it shorten the path from information to decision?
- Workflow automation: Can it complete repeatable operational work at lower fully loaded cost?
That is the pattern you are hunting for.
Not "where can we add AI?"
Where can intelligence change the economics of the work?
The AI ROI test
For any AI project, you should eventually be able to answer these questions:
- What economic outcome are we trying to change?
- What does the current workflow cost?
- What does the entire AI-assisted workflow cost?
- How many generated outputs become accepted outputs?
- How much human management and revision does AI require?
- What happens to the time we reclaim?
- How much incremental contribution profit or hard savings appears?
- What happens when the system fails?
- How reliable must it be given that failure cost?
- What level of human approval should remain?
- Can we evaluate, monitor, stop, and recover the system?
- Do these economics survive at scale?
Then there are only four reasonable conclusions:
- Scale it.
- Keep it human-assisted.
- Improve it and retest.
- Kill it.
That last answer needs to remain acceptable. Microsoft's July 2026 agent-lifecycle guidance makes the same operational point in plainer terms: retirement is a healthy outcome when an agent no longer adds enough value to justify its cost and risk.
AI does not deserve an ROI because it is AI.
It earns one when it changes the economics of a measurable business outcome.
Everything before that is evidence that it might.