Token counts on Reason and ask-about-user responses cover the
whole internal reasoning pass — several LLM calls of retrieval and synthesis — not just
the answer text you receive. A short answer with a few hundred
output_tokens is
normal, and those totals are what credits are charged on.
The organization endpoints back the dashboard and need a signed-in session — an API
key gets
401 session_required there. GET /v1/usage/me is the one to call from
inside your application.
Key usage
A lightweight balance check to run from inside your application.Response headers
Live limits are also returned on every API response, so you can track them without an extra call.Organization usage
GET /v1/usage and GET /v1/usage/breakdown both accept a period window and back the
usage views in the dashboard.
Memory context
Both also report what memory injection cost your prompts on the drop-in proxy, and how much of it your provider had already seen.GET /v1/usage/breakdown carries the same two totals per space.
Both figures are len/4 estimates, not tokenizer counts. And cache_prefix_tokens is
eligibility rather than a saving: an unchanged prefix is what makes a cache hit possible,
but whether your provider charged the discounted rate also depends on their minimum
cacheable length and their own cache TTL. Raise it with
block_order: "stable".
Next steps
Billing
Plans, credit costs, and subscription management.
Errors
What happens when a quota runs out.