Skip to main content
Most memory APIs stop at retrieval: they hand back a ranked list and leave the reasoning to you. Reason goes a step further. It runs an agentic pass across the memories in a space, connecting related facts, weighing recency, and resolving what matters, then returns one synthesized answer.

Reason over a space

Depth

Reason keeps a layer of standing answers — memory models — above the notes and the raw memories. On the default depth, a memory model that is current can answer on its own, which is what makes a typical call fast. When that model is out of date, so is the answer, and asking the same question again does not help. depth: "thorough" is how you say look harder. It makes the pass read the notes and the raw memories underneath before it answers, whatever the models say.
Credits are charged on the tokens the call actually uses, so thorough costs more because it does more, not because it is priced differently. A very deep question over a very large space can still run out of time; that is 503 reason_timeout, and dropping back to fast is the first thing to try.

A default for the space

A space that is habitually asked hard questions can make thorough its default, so callers do not have to remember the field:
Both clients take model and depth together for the reason below: they are one resource, and an argument you leave out is an argument that clears the other field. A depth on the request still wins, so one caller can drop back to fast for a simple question without changing the space.
This PUT replaces the whole resource. A body naming only model clears depth, and the other way round — send both fields, or send DELETE to clear everything.

Reason about one user

Pass any of the scope keys and the answer is synthesized from that scope alone, the same way Retrieve narrows a search. Without them, reason covers every memory in the space.
A scoped question never falls back to the wider space: memories written without a scope are not visible to a scoped call, so a space that adopts scoping part-way through will not see its earlier history here until those memories are rewritten with a scope. Response 200 OK

What the answer was built from

sources says which layer the answer came from, so a result that looks wrong can be diagnosed instead of guessed at. These are the memories the reasoning pass declared it used, not everything it looked at — which is what makes “the memory is missing” separable from “the memory is there and was not used”. Those need opposite fixes. A non-empty models list is the first thing to check when an answer looks out of date: it means a standing answer spoke. Refresh that model, or ask again with depth: "thorough" to read past it. "memories": 0, "notes": 0 with no models is not an error — it means the pass found nothing to build on, which is usually an empty space or a scope that matches no memories.
sources carries counts, not the memories themselves, because it rides every answer. To see which memories are behind a named model, read GET /v1/spaces/{space_id}/models/{model_id} — that returns the list with their text.
client.reason() returns the synthesized string directly, or None when nothing was found. The TypeScript client returns the whole InsightsResult, so read .insights.

Which rules shaped the answer

A space can be given rules — standing instructions it must follow whatever the memories say. rules_applied names the ones that actually applied to this answer:
The rule’s text is deliberately not echoed. It is your own configuration, readable once from GET /v1/spaces/{space_id}/rules, while this rides every answer — and a space may hold 25 active rules of 2,000 characters each. The name is what identifies which rule spoke. rules_applied is always present and never null. An empty list is the finding that no rule applied, which is worth being able to establish: a rule overrides what the memories say, so the memories are wrong and a rule said otherwise look identical in the prose and need opposite fixes. When an answer is surprising, this is the first field to read after sources.
The TypeScript client returns the whole result, so rules_applied is on it. The Python client.reason() returns the synthesized text alone — read the field from the API response if you need it.

How it works

Reason is read-only and never writes new memories. Internally it runs a multi-step pass: retrieving relevant facts, expanding context through linked entities, and cross-referencing related memories. Answer quality improves as more relevant memories accumulate in the space.
usage covers that whole internal pass, not just the answer you see. Reason may make several LLM calls before it settles on a synthesis, so output_tokens is routinely several times larger than the visible answer — a one-sentence conclusion can honestly report a few hundred output tokens. Credits are charged on these totals. The same applies to asking about one user.

Choose the model that answers

Reason is the call where the model matters most, because it is an agentic pass, not a single completion, so a stronger model changes the quality of the synthesis rather than just its wording. Name one per request, or set a default for the space:
The response echoes the model that actually answered, and credits are charged at that model’s rate. Full detail in model selection.

Error responses

See the full error reference.

Next steps

Retrieve API

Query memories for specific information.

Chat

OpenAI-compatible chat with memory built in.