Pattern 04 · AI cost and token waste

The bill is visible. What caused it is not.

AI consumption grows quietly and arrives as a single number. When nobody can decompose that number into the agents, applications and teams behind it, cost stops being something anyone can manage and becomes something someone has to explain.

What happens

Spend without attribution is spend without a lever.

Provider bills aggregate by account or key. Several agents commonly share one credential, so the total cannot be decomposed. The question that gets asked at the end of a quarter — why did this grow — has no answer that anyone can act on.

Nobody owns the number

The bill lands with finance, the consumption belongs to a dozen teams, and the mapping between them does not exist.

Waste is silent

An agent stuck in a loop and an agent doing valuable work look identical on an invoice. Both are just tokens.

The response is a freeze

Without attribution the only available control is blunt — slow everything down — which costs more in stalled adoption than it saves.

The control that is not there

Cost is measured where it is billed, not where it is caused.

Consumption is recorded by the provider against a credential. The thing you actually want to manage — which agent, doing what, on whose behalf — is not in that record at all.

Cloud cost management solved this years ago with tagging and showback. AI consumption has no equivalent by default, because the unit of work is a call made by a non-human caller that nothing was recording.

Where FaburAI helps

Measure it where it happens.

Governed calls already carry identity, so the same record that answers "was this allowed" also answers "what did it consume."

Attributed, not aggregated

Token consumption is recorded per call and broken down by AI actor, application, model and server. Not a total to be reconciled — a breakdown that was never combined in the first place.

Consumption becomes cost

Operator-configured rate cards turn tokens into currency at rates you set, so the breakdown is in the units the conversation is actually held in.

Trend, not snapshot

Consumption over time with the model behind it, so a step change can be traced to the week it started and the agent that started it.

Budgets are policy

A token budget or call-rate limit is a policy condition like any other. A rule can deny, warn or elevate auditing once an agent crosses a threshold in a window — through the same approval as any access rule.

That last one is the part most cost tooling cannot do. Because spend and access run through the same policy mechanism, a limit is enforced rather than reported — the agent stops, instead of appearing in next month's review.

Waste detection

Three patterns, found in your own activity.

Analysis runs on a schedule against recorded calls and surfaces what it finds as recommendations. Each can be acted on or dismissed, and a dismissal sticks.

Redundant calls

An agent making the same request over and over with near-identical arguments. Usually a missing cache or a loop that re-asks rather than remembering — and every repeat is paid for.

Retry loops

An agent repeatedly retrying a call that keeps being refused. It consumes tokens without ever succeeding, and it is invisible on a bill because a denied call still costs.

Unusual call density

One agent running at several times the median rate of its peers. Sometimes that is the job; often it is a scheduling mistake nobody has noticed.

These are rule-based patterns over real activity, not predictions. Each one points at a specific agent and a specific behavior, which is what makes it actionable by the team that owns that agent.

What FaburAI alone does not solve

Where this control ends.

  • We do not fix your agents.A redundant-call pattern tells you an agent is re-asking rather than caching. Changing that belongs to the team that owns it — the platform surfaces the behavior, it does not rewrite the application.
  • We are not your billing system of record.Owned by your finance and FinOps tooling. Rate cards produce an operational estimate for managing behavior, not an invoice to reconcile against.
  • We do not choose models or route between them.Owned by your AI platform team. We can show that one agent's model choice is expensive; deciding what to run is theirs.
  • We see governed calls.Consumption that never passes through a governed call is not in the record, so coverage of spend follows coverage of governance.

Every line above is a reason we integrate rather than replace. We are glad to sit alongside the finance and platform tooling already in place and be the layer that attributes consumption to the agent that caused it.

Talk to us

If your AI spend grew and nobody can say why.

That is the usual starting position, and attribution is normally the first thing that makes the conversation tractable.