Operations

The Agent P&L: Measure Resolved Work, Not Tokens

An AI agent can use fewer tokens this month and still become a worse business investment. It can also cost more per run while creating far more value. Tokens tell you what the model consumed. They do not tell you whether the work got done. The unit that matters is cost per resolved workflow.

Start with the work before the agent

Before automating a workflow, measure how it operates today. How many requests arrive? How many labor minutes does each one consume? How often does someone correct an error, chase missing information, or reopen the work? How long does a request sit between handoffs?

That baseline prevents a common accounting trick: comparing the agent's visible model bill with no equivalent measure for the existing process. Manual work has inference costs too. They show up as payroll, interruptions, delayed responses, rework, and management time. Without the baseline, a cheap agent looks impressive even when it simply moves the cleanup to someone else.

Define “resolved” in business terms

A run is not a resolution. A drafted response is not a resolution. A green API response is not a resolution. The workflow is resolved when the intended business outcome is complete: the invoice is reconciled, the intake is accepted, the appointment is booked, or the exception reaches the right person with enough context to decide.

Write that definition down before launch. Include the quality bar and the time window. An invoice match that silently uses the wrong purchase order is not resolved. A customer reply that arrives three days after the service-level target is not fully resolved either. This is how the scorecard stays connected to operations instead of drifting toward convenient technical metrics.

The agent's job is not to produce output. It is to move a unit of work to a verified business outcome at an acceptable cost and risk.

Count the whole cost of the run

Model usage belongs in the calculation, but it is only one line. Add integration hosting, retrieval, observability, storage, human review, maintenance, and the cost of failed attempts. If a workflow needs three retries and ten minutes of employee cleanup, all four attempts belong to the cost of the one resolution.

The same rule applies to exceptions. Human review is not evidence that an agent failed. Review can be the correct design for a high-consequence decision. But it is still a cost, and hiding it makes expansion decisions worse. Track review minutes, escalation rate, and correction rate alongside infrastructure spend.

A useful monthly view is straightforward: total operating cost divided by verified resolutions. Put cycle time, first-pass resolution, exception rate, and error correction beside it. Now a change that saves tokens but creates more rework is visible for what it is.

Separate healthy exceptions from unresolved work

Some cases should leave the automated path. A missing contract, an ambiguous customer identity, or a payment above an approval limit should trigger an escalation. That is a controlled resolution path, not a defect, when the agent packages the evidence and routes it correctly.

The dangerous category is unresolved work disguised as success: abandoned retries, drafts nobody reviews, records written to the wrong account, or exceptions dumped into a queue without ownership. Give every terminal state a name, an owner, and a measurable outcome. If the system cannot explain where the work ended, it cannot claim the work was resolved.

Use the P&L to decide what happens next

Once the unit economics are visible, the operating choices get clearer. Expand an agent when resolution cost and cycle time beat the baseline without weakening quality. Tune it when a small group of recurring exceptions drives most of the review burden. Narrow its scope when low-frequency edge cases consume disproportionate maintenance. Retire it when the workflow never had enough volume or value to justify ongoing ownership.

This also keeps teams from defending sunk cost. An agent is not successful because it was difficult to build, uses an impressive model, or completed thousands of runs. It earns more scope by resolving valuable work reliably. The portfolio should reward that evidence.

A practical scorecard

For each workflow, track volume received, verified resolutions, automated resolution rate, human review minutes, correction rate, median cycle time, total operating cost, and cost per resolution. Add one risk measure specific to the work, such as incorrect payment rate or messages sent without approval. Review the trend, not a single week.

Tokens still matter. They are useful for diagnosing a cost change and comparing technical approaches inside the same workflow. They just belong below the business metric, not above it. Optimize the model bill after you know the workflow is producing the right outcome.

Foundation builds managed agents with the operating scorecard attached: outcomes, review, exceptions, cost, and the evidence needed to earn more autonomy. If you want to find the real unit economics of a workflow in your business, talk to us.