Summary

  • The 28 September AAuth Budgets Internet-Draft proposes a resource-enforced spending ceiling on each authorization token, not an automatic ceiling across all of a person's agents or missions.
  • When tokens overlap, the person server must account for the full amount it has allocated and wait for credible usage readings before releasing uncertain headroom.

Imagine a person authorizes two agent missions at the same inference service. Each receives a token with a hypothetical $10 ceiling. The service can enforce both perfectly and still be entitled to charge as much as $20. If the person server meant to allow only $10 in total, the failure was not at either token's stop gate. It happened when the issuer promised the same headroom twice.

Dick Hardt's 28 September 2026 initial draft, AAuth Budgets, makes that distinction unusually explicit. It is an individual, exploratory Internet-Draft, with “Standards Track” as its intended status. The IETF Datatracker currently says only that the draft exists and defines no RFC stream; this is not a ratified IETF rule or evidence of broad deployment. The proposal extends a separate AAuth protocol draft. An agent asks for a budget, a resource offers a unit and possible amount, the person server and any access server can narrow it, and a signed authorization token carries the grant to the resource. The resource counts the consumption. The person server is not in the request path and keeps the person's overarching ceiling to itself.

The limit is deliberately specific. The draft's resource-side invariant applies to the token presented on a request: committed consumption plus outstanding reservations against that token must not exceed its granted amount. For a response whose eventual output determines the charge, the resource first needs a finite maximum. It can use a declared output bound or a documented default, reserve that possible cost, then commit the actual cost and release the difference. A request may be refused because its maximum exceeds the remaining allowance even if the eventual charge would have fitted. A narrower request can follow. This is how a purported hard cap differs from a dashboard warning issued after the bill.

The same resource also maintains a person-level usage ledger, but that ledger is not a second limit of its own invention. It does not know the person's private ceiling. Multiple live tokens may therefore be individually valid at one resource. The draft tells the issuer to size their sum against its ceiling. Its example risk is arithmetic rather than a reported incident: n concurrent tokens of amount X authorize exposure of up to nX while they remain live. In the $10 thought experiment, stopping one conduit does not retroactively erase what flowed through it or constrain the other.

Accounting becomes harder when a token expires, is refreshed without being presented again, or is stranded by a crashed agent. The issuer may know what it granted but not what was consumed. Under the draft, it must temporarily assume the whole unknown allocation was spent. A consumption record attached to another challenge is only a snapshot; it does not close a still-valid token. The usage endpoint's meter reading, complete through a stated as_of time, can settle allocations that had expired by then. A recorded revocation can advance that point; unsupported or unavailable revocation cannot safely release the reservation early, and a request already in flight may finish. Conservative accounting can delay legitimate work, but a cheerful zero would permit the issuer to promise money already spent.

This is not merely a finance feature. A budget does not replace scope: a cheap irreversible action is still irreversible. Nor does the token itself identify which of several billing accounts should be charged; that binding remains a separate resource-account decision. Delegating to another agent does not create an automatic sub-budget drawn from the parent. Each downstream token needs its own grant through the person server, and those grants again compete for the same private ceiling.

The draft offers an AAuth-Budget response header so an agent can see a request's cost and remaining token balance. It is unsigned by default and is a pacing signal, not the source of authorization. It recommends, but does not require, signatures on usage responses. A signature records what a resource reported; it does not prove its meter honest or settle a billing dispute. The draft's implementation section includes author-supplied reports, not independently established adoption. Its common case of provider-paid, provider-bundled agent inference is outside this extension; the sharper use case is agent software spending on the person's own metered account.

The immediate editorial test is therefore not whether a token has a reassuring round-number cap. It is whether the issuer can show, for the same person and resource, which grants are still live, which amounts remain reserved, which meter reading is current, and what a stop request actually stopped. That is the difference between a limit one component obeys and a total the person can rely on.

Sources