fix(agent): count Anthropic cache tokens and stop message_delta erasing usage - #2127
Conversation
…ng usage
Two counting bugs on the Anthropic path, both silent.
Anthropic reports `input_tokens` net of the prompt cache and bills
`cache_read_input_tokens` and `cache_creation_input_tokens` separately, so
`input_tokens + output_tokens` is not what the call is charged on. It never sends
`total_tokens`, so the computed branch is always the one that runs. A call billed on
21,623 tokens reported 123. `cache_creation_input_tokens` was not read anywhere in the
repository, and Anthropic bills cache writes above the base rate.
OpenAI is the opposite: its `cached_tokens` sits inside `prompt_tokens`, so adding it
would double count. Only the Anthropic-shaped keys are summed, and a test pins that.
Second, `tokenUsageFromRecord` returns every key and sets the ones it did not find to
`undefined`. Anthropic's streaming `message_delta` carries `output_tokens` only, so
`usage = { ...usage, ...deltaUsage }` overwrote `inputTokens` and `cachedInputTokens`
with `undefined`, discarding what `message_start` had already reported. After any
streamed call the whole input side of the accounting was gone, with nothing raised.
`mergeTokenUsage` merges only defined values. It does not touch `totalTokens`, because
whether a cache counter is additive is a provider question and that information is gone
once usage has been normalised, so the caller owns the total.
Three of the added tests fail on main.
There was a problem hiding this comment.
2 issues found across 4 files
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="libraries/typescript/packages/agent/src/llm/types.ts">
<violation number="1" location="libraries/typescript/packages/agent/src/llm/types.ts:118">
P2: The inspector will drop Anthropic cache-creation token details from recorded `tokenUsage` and trace spans even though this field is now emitted by the agent. Propagating `cacheCreationInputTokens` through `InspectorTokenUsage`, parsing, and aggregation would keep the new public usage field observable.</violation>
</file>
<file name="libraries/typescript/packages/agent/src/llm/usage.ts">
<violation number="1" location="libraries/typescript/packages/agent/src/llm/usage.ts:40">
P2: `totalTokens` can still undercount cache-read tokens for camel-case usage payloads. `cacheReadOutsideInput` only checks `cache_read_input_tokens`, so records that provide `cachedInputTokens` are parsed but not included in the computed total fallback.</violation>
</file>
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
| /** Input tokens served from a provider cache. */ | ||
| cachedInputTokens?: number; | ||
| /** Input tokens written to a provider cache. Billed, and on Anthropic above base rate. */ | ||
| cacheCreationInputTokens?: number; |
There was a problem hiding this comment.
P2: The inspector will drop Anthropic cache-creation token details from recorded tokenUsage and trace spans even though this field is now emitted by the agent. Propagating cacheCreationInputTokens through InspectorTokenUsage, parsing, and aggregation would keep the new public usage field observable.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At libraries/typescript/packages/agent/src/llm/types.ts, line 118:
<comment>The inspector will drop Anthropic cache-creation token details from recorded `tokenUsage` and trace spans even though this field is now emitted by the agent. Propagating `cacheCreationInputTokens` through `InspectorTokenUsage`, parsing, and aggregation would keep the new public usage field observable.</comment>
<file context>
@@ -114,6 +114,8 @@ export interface TokenUsage {
/** Input tokens served from a provider cache. */
cachedInputTokens?: number;
+ /** Input tokens written to a provider cache. Billed, and on Anthropic above base rate. */
+ cacheCreationInputTokens?: number;
/** Tokens used for model reasoning. */
reasoningTokens?: number;
</file context>
| // input_tokens alone is not what the call is charged on. OpenAI reports cached_tokens | ||
| // INSIDE prompt_tokens, so adding that one would double count. Only the Anthropic-shaped | ||
| // keys are summed here. | ||
| const cacheReadOutsideInput = numberAt(usage, "cache_read_input_tokens"); |
There was a problem hiding this comment.
P2: totalTokens can still undercount cache-read tokens for camel-case usage payloads. cacheReadOutsideInput only checks cache_read_input_tokens, so records that provide cachedInputTokens are parsed but not included in the computed total fallback.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At libraries/typescript/packages/agent/src/llm/usage.ts, line 40:
<comment>`totalTokens` can still undercount cache-read tokens for camel-case usage payloads. `cacheReadOutsideInput` only checks `cache_read_input_tokens`, so records that provide `cachedInputTokens` are parsed but not included in the computed total fallback.</comment>
<file context>
@@ -33,10 +33,23 @@ export function tokenUsageFromRecord(raw: unknown): TokenUsage | undefined {
+ // input_tokens alone is not what the call is charged on. OpenAI reports cached_tokens
+ // INSIDE prompt_tokens, so adding that one would double count. Only the Anthropic-shaped
+ // keys are summed here.
+ const cacheReadOutsideInput = numberAt(usage, "cache_read_input_tokens");
+ const cacheCreationInputTokens = numberAt(
+ usage,
</file context>
| const cacheReadOutsideInput = numberAt(usage, "cache_read_input_tokens"); | |
| const cacheReadOutsideInput = numberAt( | |
| usage, | |
| "cachedInputTokens", | |
| "cache_read_input_tokens" | |
| ); |
… inspector dropping cache Follow-up on the review. Four points were raised; three were right and are fixed, one would have introduced a bug and the reasoning is below. The same wrong total existed in a third place. `message_stop` recomputed it as `inputTokens + outputTokens`, so the fix at `message_delta` survived only because `??` short-circuited on the total already set. Two places knowing how Anthropic bills is how this drifts, so `message_stop` now owns it alone and `message_delta` just merges. Added an end-to-end test that drives real `message_start` then `message_delta` SSE through `streamChat`, rather than exercising `mergeTokenUsage` in isolation. On the base commit it reports `totalTokens: undefined`, not merely a wrong number, because the input counters were gone by the time `message_stop` ran. The inspector was dropping more than the new field. `InspectorTokenUsage` declares `cachedInputTokens` and `reasoningTokens`, the trace view reads them, and `inspectorTokenUsageFromUnknown` returned only `inputTokens`, `outputTokens` and `totalTokens`. Both have always rendered as absent whatever the provider sent. It also repeated the same input-plus-output total. Both fixed, and `cacheCreationInputTokens` is carried through. The merge test now seeds `cache_creation_input_tokens` too, since cache writes are billed above base rate and need the same protection as cache reads. Not done, deliberately: adding `cachedInputTokens` to the computed total inside `tokenUsageFromRecord`. On Anthropic that counter sits outside `input_tokens`; on OpenAI `cached_tokens` sits inside `prompt_tokens`. Once usage is normalised that distinction is gone, so summing it would double count OpenAI. Only the Anthropic-shaped keys are added, and a test pins the OpenAI case at 14.
|
Thanks for the review. Three of the four were right and are fixed in The end-to-end test found a third instance of the same defectTesting
The new test in Not a wrong number, an absent one, because the input counters were gone by the time The inspector was dropping more than the new fieldWorth flagging as a separate pre-existing bug. Cache creation in the merge testAdded. Cache writes are billed above the base rate, so they need the same protection across the merge that cache reads do. The one not doneAdding Happy to split the inspector change into its own PR if you would rather keep this one to the agent package. |
|
Thanks for these. Going through them in order. Three of the four landed in P2, the inspector dropping cache-creation detail. Fixed. P3, the merge fixture only seeding P3, the tests exercising P2, The first is Anthropic-shaped and does undercount, exactly as you describe. The second is OpenAI-shaped and is already correct, because OpenAI reports It is also unreachable from the current call sites, which all pass a raw provider payload: Anthropic and OpenAI send snake case, Google sends Where a provider genuinely needs its cache counted, the provider owns its own total rather than leaning on the generic parser, which is what Happy to add a test pinning that camel-case OpenAI record at 14 if it is better to have that reasoning enforced than written in a comment. |
@mcp-use/agent
@mcp-use/cli
@mcp-use/client
create-mcp-use-app
@mcp-use/inspector
mcp-use
@mcp-use/tunnel
commit: |
|
A note while this is waiting for review, in case it helps size it. The same shape is in four other codebases that share no code with this one: Pipecat, LiveKit Agents, Haystack and LlamaIndex. The first three have merged the fix. Anthropic reporting Since it repeats, it is checkable, and I packaged the check: https://github.com/arthi-arumugam-git/cachecheck. Correcting myself on one point: run against this branch's base it flags |
format:check fails the typescript-lint job on this one file. Prettier wraps the cachedInputTokens call across three lines, the way the two calls below it already are. Formatting only, no behaviour change.
khandrew1
left a comment
There was a problem hiding this comment.
Thanks for the fix—the Anthropic accounting and streaming changes look good. Before merging, please address the two inspector follow-ups below and add changesets covering both @mcp-use/agent and @mcp-use/inspector.
| outputTokens?: number; | ||
| totalTokens?: number; | ||
| cachedInputTokens?: number; | ||
| cacheCreationInputTokens?: number; |
There was a problem hiding this comment.
cacheCreationInputTokens is parsed and placed on individual usage events, but addUsage() does not aggregate the new field. As a result, it disappears from state.usage, the LLM span’s usage, and the aggregate tokenUsage written to the raw-chat payload.
Please add the following field to the object returned by addUsage():
cacheCreationInputTokens: sum(
current?.cacheCreationInputTokens,
next.cacheCreationInputTokens
),It would also be useful to extend trace.test.ts to assert that this counter survives in both aggregate state and LLM span usage.
| "cache_creation_input_tokens" | ||
| ); | ||
| const reasoningTokens = number("reasoningTokens", "thoughtsTokenCount"); | ||
| const uncountedCache = |
There was a problem hiding this comment.
This fallback can double-count OpenAI cache reads because cachedInputTokens is provider-neutral. For OpenAI, cached tokens are already included in inputTokens, so a normalized record such as { inputTokens: 10, outputTokens: 4, cachedInputTokens: 3 } currently produces 17 instead of 14 when totalTokens is absent.
Please keep parsing cachedInputTokens for observability, but calculate uncountedCache from the Anthropic-specific cache_read_input_tokens key instead. This would mirror the provider-safe behavior in agent/src/llm/usage.ts.
A regression test for the normalized OpenAI example above would pin the distinction.
…unting OpenAI cache reads Addresses @khandrew1's review on mcp-use#2127. addUsage dropped cacheCreationInputTokens: the field was parsed onto individual usage events but never summed, so it vanished from state.usage, the LLM span's usage, and the aggregate tokenUsage in the raw-chat payload. It is now summed alongside the other counters. The fallback total could also over-count. uncountedCache was built from the provider-neutral cachedInputTokens, but on OpenAI those tokens are already inside inputTokens, so { inputTokens: 10, outputTokens: 4, cachedInputTokens: 3 } totalled 17 instead of 14. The total now adds only the Anthropic-shaped cache_read_input_tokens, mirroring the provider-safe path in agent/src/llm/usage.ts. cachedInputTokens is still parsed and returned for observability. Two tests in trace.test.ts pin both: the OpenAI-vs-Anthropic total distinction, and that cacheCreationInputTokens survives aggregation and the LLM span. Changesets added for @mcp-use/agent and @mcp-use/inspector.
|
Thanks, both inspector follow-ups are in as of
The OpenAI double-count. Good catch. Two tests in |
|
@khandrew1 gentle ping on this one. Both of your review points went in on 11 August in db1daa1: Happy to rebase if it has drifted. If you would rather I split the inspector fix from the accounting fix into two PRs, say the word and I will do that instead. |
* fix(agent): count Anthropic cache tokens and stop message_delta erasing usage (#2127) * fix(agent): count Anthropic cache tokens and stop message_delta erasing usage Two counting bugs on the Anthropic path, both silent. Anthropic reports `input_tokens` net of the prompt cache and bills `cache_read_input_tokens` and `cache_creation_input_tokens` separately, so `input_tokens + output_tokens` is not what the call is charged on. It never sends `total_tokens`, so the computed branch is always the one that runs. A call billed on 21,623 tokens reported 123. `cache_creation_input_tokens` was not read anywhere in the repository, and Anthropic bills cache writes above the base rate. OpenAI is the opposite: its `cached_tokens` sits inside `prompt_tokens`, so adding it would double count. Only the Anthropic-shaped keys are summed, and a test pins that. Second, `tokenUsageFromRecord` returns every key and sets the ones it did not find to `undefined`. Anthropic's streaming `message_delta` carries `output_tokens` only, so `usage = { ...usage, ...deltaUsage }` overwrote `inputTokens` and `cachedInputTokens` with `undefined`, discarding what `message_start` had already reported. After any streamed call the whole input side of the accounting was gone, with nothing raised. `mergeTokenUsage` merges only defined values. It does not touch `totalTokens`, because whether a cache counter is additive is a provider question and that information is gone once usage has been normalised, so the caller owns the total. Three of the added tests fail on main. * fix(agent,inspector): one owner for the Anthropic total, and stop the inspector dropping cache Follow-up on the review. Four points were raised; three were right and are fixed, one would have introduced a bug and the reasoning is below. The same wrong total existed in a third place. `message_stop` recomputed it as `inputTokens + outputTokens`, so the fix at `message_delta` survived only because `??` short-circuited on the total already set. Two places knowing how Anthropic bills is how this drifts, so `message_stop` now owns it alone and `message_delta` just merges. Added an end-to-end test that drives real `message_start` then `message_delta` SSE through `streamChat`, rather than exercising `mergeTokenUsage` in isolation. On the base commit it reports `totalTokens: undefined`, not merely a wrong number, because the input counters were gone by the time `message_stop` ran. The inspector was dropping more than the new field. `InspectorTokenUsage` declares `cachedInputTokens` and `reasoningTokens`, the trace view reads them, and `inspectorTokenUsageFromUnknown` returned only `inputTokens`, `outputTokens` and `totalTokens`. Both have always rendered as absent whatever the provider sent. It also repeated the same input-plus-output total. Both fixed, and `cacheCreationInputTokens` is carried through. The merge test now seeds `cache_creation_input_tokens` too, since cache writes are billed above base rate and need the same protection as cache reads. Not done, deliberately: adding `cachedInputTokens` to the computed total inside `tokenUsageFromRecord`. On Anthropic that counter sits outside `input_tokens`; on OpenAI `cached_tokens` sits inside `prompt_tokens`. Once usage is normalised that distinction is gone, so summing it would double count OpenAI. Only the Anthropic-shaped keys are added, and a test pins the OpenAI case at 14. * style(inspector): run prettier on trace.ts format:check fails the typescript-lint job on this one file. Prettier wraps the cachedInputTokens call across three lines, the way the two calls below it already are. Formatting only, no behaviour change. * fix(inspector): aggregate cacheCreationInputTokens and stop double-counting OpenAI cache reads Addresses @khandrew1's review on #2127. addUsage dropped cacheCreationInputTokens: the field was parsed onto individual usage events but never summed, so it vanished from state.usage, the LLM span's usage, and the aggregate tokenUsage in the raw-chat payload. It is now summed alongside the other counters. The fallback total could also over-count. uncountedCache was built from the provider-neutral cachedInputTokens, but on OpenAI those tokens are already inside inputTokens, so { inputTokens: 10, outputTokens: 4, cachedInputTokens: 3 } totalled 17 instead of 14. The total now adds only the Anthropic-shaped cache_read_input_tokens, mirroring the provider-safe path in agent/src/llm/usage.ts. cachedInputTokens is still parsed and returned for observability. Two tests in trace.test.ts pin both: the OpenAI-vs-Anthropic total distinction, and that cacheCreationInputTokens survives aggregation and the LLM span. Changesets added for @mcp-use/agent and @mcp-use/inspector. * chore: enter canary prerelease mode * chore(typescript): version packages (canary) * fix(inspector): scope loading state by chat session (#2232) * fix(inspector): give each chat session one id and one state record Chat state lived in hook-wide React state, so starting a new chat cleared the values an in-flight stream was still writing to: the fresh chat inherited the previous session's loading state and could not send. Rework it around a single identity. A session is created with one id and keeps it; that id is handed to ChatStorageProvider.createChat, so runtime state, chat history and OAuth retry all name the same chat. All of a session's state and turn machinery live in one record inside ChatSessionStore, which notifies only the subscribers of the session that changed, so a background turn keeps streaming into its own record. useChatSessions owns the lifecycle: minting ids, activating a session, reattaching to persisted chats, and reopening the session a full-page OAuth redirect interrupted. New Chat now only switches the active id. Closes #2217 * refactor(inspector): simplify chat session state * fix(inspector): harden chat session recovery --------- Co-authored-by: Andrew Khadder <andrew@manufact.com> * chore(typescript): version packages (canary) * fix(ci): give inspector tests the same timeout as every other package (canary) (#2290) * fix(ci): give inspector tests the same timeout as every other package Same one-line change as #2268, which targets main. canary carries the same 10s cap and is failing on it independently: run 32053629834 compressed-assets.test.ts 10230ms Test timed out in 10000ms run 32051250836 same test main is red on it too (run 32054211669). Recent margins are 10029ms, 10230ms and 10837ms, all just over the cap. The inspector is the only package at 10s; agent, client, cli and server all use 60000, and these cases mount a server and decompress the bundled app. 45 files, 217 tests still pass locally. * chore(ci): remove timeout rationale comment --------- Co-authored-by: Andrew Khadder <andrew@manufact.com> * feat(oauth): add Scalekit provider (#2272) * feat(oauth): add Scalekit provider Add oauthScalekitProvider next to Auth0 and Clerk. The verifier accepts both environment-root and resource-scoped issuers, binds audience to the res_… resource id, and advertises DCR plus CIMD. No Scalekit SDK. * fix(oauth): keep Scalekit resource-id audience with extra aud Always verify aud includes the resource id. Optional audience is an extra AND check, not a replacement. Drop the new v1 provider page. * fix(oauth): drop Scalekit issuerBoundAccessTokens option Scalekit binds tokens to the resource id. The dashboard will own that opt-in. Keep the static oauthMetadata copy because mcp-use requires it. Custom claims stay on ctx.auth.payload. * docs: drop v1 and legacy skill changes * docs: fix Scalekit environment variables --------- Co-authored-by: Saif Shines <saifshine7@gmail.com> Co-authored-by: Andrew Khadder <andrew@manufact.com> * chore(typescript): version packages (canary) * fix(inspector): add a test script so pnpm test covers this package (#2289) * fix(inspector): add a test script so pnpm test covers this package The root test script is `pnpm run -r test`, and the inspector only had test:unit. pnpm skips a package with no matching script and reports nothing, so 217 unit tests were absent from the workspace-wide run: $ pnpm --filter @mcp-use/inspector test (no output) $ pnpm --filter @mcp-use/inspector test:unit Test Files 45 passed (45) Tests 217 passed (217) CI is unaffected; its step calls test:unit directly. This is the documented `pnpm test` contributors run. Uses run mode rather than watch, matching tunnel, so a recursive run cannot sit waiting on a watcher. * chore(inspector): remove test-only changeset --------- Co-authored-by: Andrew Khadder <andrew@manufact.com> * fix(types): align pnpm workspace tooling (#2301) * docs: update TypeScript release changelogs * fix(inspector): use secure chat session ids --------- Co-authored-by: Arthi Arumugam <62213284+arthi-arumugam-git@users.noreply.github.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Mert Çetin <97505365+merttcetn@users.noreply.github.com> Co-authored-by: Ayaan Gazali <ayaangazali.work@gmail.com> Co-authored-by: saif-at-scalekit <saif.shaik@scalekit.com> Co-authored-by: Saif Shines <saifshine7@gmail.com>
Two token counting bugs on the Anthropic path. Both are silent: the number comes out wrong and nothing raises.
1.
totalTokensleaves out the cache countersAnthropic reports
input_tokensnet of the prompt cache, and billscache_read_input_tokensandcache_creation_input_tokenson top of it. It never sendstotal_tokens, so the computed branch intokenUsageFromRecordis always the one that runs.cache_creation_input_tokenswas not read anywhere in the repository, so cache writes were invisible, and Anthropic bills those above the base rate.OpenAI is the opposite shape: its
cached_tokenssits insideprompt_tokens, so adding it would double count. Only the Anthropic-shaped keys are summed, and there is a test pinning the OpenAI case at 14 so this cannot regress into a double count.2.
message_deltaerased the counters frommessage_starttokenUsageFromRecordreturns every key and sets the ones it did not find toundefined. Anthropic's streamingmessage_deltacarriesoutput_tokensonly, so:overwrote
inputTokensandcachedInputTokenswithundefined.After any streamed call the entire input side of the accounting was gone.
mergeTokenUsagemerges only defined values. It deliberately does not recomputetotalTokens: whether a cache counter is additive is a provider question, and that information is gone once usage has been normalised, so the caller owns the total.streamChatin the Anthropic provider does that sum itself, where the semantics are known.Tests
Added to
src/llm/__tests__/usage.test.ts. Three of them fail onmain, the fourth is the OpenAI regression guard and passes on both. The two existing assertions in that file still pass unchanged.Happy to split this into two PRs if you would rather review them separately.
Summary by cubic
Fixes Anthropic token accounting across
agentandinspector. Totals now include cache read/write tokens, streaming keeps prior counters, and the inspector aggregates cache writes while avoiding OpenAI double-counts.Bug Fixes
cache_read_input_tokensandcache_creation_input_tokensin totals (and trace); addcacheCreationInputTokenstoTokenUsage; OpenAI unaffected.mergeTokenUsagesomessage_deltadoesn’t erase inputs; recompute the billed total once atmessage_stopto avoid drift.inspector: aggregatecacheCreationInputTokensacross events and the LLM span; compute the fallback total by adding only Anthropic-shaped cache counters, not OpenAI’scached_tokens; surfacecachedInputTokens,cacheCreationInputTokens, andreasoningTokens.Refactors
inspectortrace to satisfy lint; no behavior change.Written for commit db1daa1. Summary will update on new commits.