Skip to main content
Retrieve understands meaning rather than keywords. “Where does Alice work?”, “Alice’s employer”, and “What company is Alice at?” all surface the same memory. Your query is embedded and compared against stored memory embeddings, then results are ranked by relevance_score.

Parameters

Response

Both SDKs return the results list directly, and an empty list when nothing matches. Response fields keep the API’s own snake_case names in TypeScript too, so a result is r.relevance_score rather than r.relevanceScore.

Let Anona choose the space

Send route: "auto" in place of space_id when you know the question but not which space holds the answer. Anona picks one space, searches it, and reports which one in searched. Exactly one of the two fields must be present: sending both, or neither, is 422 route_conflict.
One space is searched, never several.
On a retrieve that named a space_id, searched is absent from the response body, not null. The key does not appear at all, so a call that does not use routing is byte-for-byte what it always was. Check whether the key is present rather than comparing against null.
Read searched. A query answered from the wrong space looks exactly like a space that holds nothing, and this is the only field that tells the two apart.

When no space fits

An abstention is a 200 with an empty results array and one searched entry whose space_id is null. The entry is present, rather than searched being empty, because the reason has to be carried somewhere.
reason is no_space_fits when the model reported that none of your spaces holds the answer, and low_confidence when it picked one but not confidently enough to act on.
Your default space is deliberately not searched here. On a write the default space is the sink and filing there is correct; on a read it is the bag of everything that fitted nowhere, which makes it the least topically coherent space you have and so the worst place to look for a specific answer. Irrelevant memories are worse than none. Widen deliberately instead: name a space_id, or pass fallback_space_id.
An abstained query is charged. A retrieve that names a space and matches nothing is charged today, and a routed retrieve that matches nothing is charged the same way. The usage row groups under your organization’s default space, or under unrouted if you have none set.
Four other reasons are not answers about your memories, and are refused rather than reported as an empty result set. Without a fallback_space_id each is 422 no_route_target: no_candidates (no space of yours is eligible for routing), router_disabled (routing is not switched on for this deployment), answer_outside_candidates (the model named a space never offered to it) and model_unavailable (the decision provider could not be reached). Sending a fallback_space_id searches that space instead, with stage: "fallback" and the reason carried along.

Routing applies to retrieve only

route is accepted on POST /v1/retrieve and the MCP retrieve tool. POST /v1/reason requires a space_id, and so do the SDKs’ get_context and retrieve_receipt, which are renderings of this endpoint that do not take route. Those calls are addressed only. Which spaces are eligible, and how to describe a space so it gets chosen, are in Automatic space routing. One difference from the write side is worth knowing here: a shared space you hold as a viewer is eligible for a routed read, because a viewer can read it, while a routed write needs developer.
retrieve still returns a list, so nothing about existing code changes. searched rides along as an attribute on it, and is None on an addressed read.

Memory types

Every result carries a memory_type. The same content can exist at more than one level, as raw evidence and as a distilled version of it, so filtering by type asks for exactly the layer you want. By default (prefer_observations: true) a note is returned in place of the facts it was built from, so you see one distilled result instead of the same content several times. Set prefer_observations: false to receive the raw facts alongside their note. A note names its evidence either way. Every result carries source_ids, the memory_id of each memory it was synthesized from, so a distilled answer can still be attributed, cited or audited back to the raw writes behind it. source_ids is empty on a fact, which is evidence rather than a synthesis. A note’s own metadata is empty and its document_id is null, because both belong to the facts it was built from rather than to the synthesis. To read them, set prefer_observations: false: the raw facts come back alongside their note, each carrying the metadata you wrote it with, and source_ids tells you which of them belong to which note.

Entities and time

Each result carries structured signal beyond the raw text. The same entities are exposed as a browsable graph. See Graph.

Searching over time

Every memory carries two independent times, and knowing which is which is most of the work: They differ whenever you import history: a memory about last June that you upload today has an event time of last June and a record time of today.
The timestamp you send is returned as timestamp. It is not copied into occurred_start. Those two are filled only from a date found in the memory’s own text, so this write:
comes back with timestamp: "2024-06-15T10:00:00Z" and both occurred_* fields null. That is not a gap in the data. occurred_start: null is the honest answer to “what event window did the text describe”, and the window filter below falls back to timestamp, so this memory still matches a June 2024 window. Read timestamp when you want to know when something happened; read occurred_start only when you specifically want a span the text spelled out.
as_of and query_timestamp are not variants of each other: one filters, the other ranks, and both work on record time. as_of — point-in-time recall. What did the space know on June 1st? Anything recorded after that instant is dropped, however relevant.
query_timestamp — move “now”. Rank as though it were June 1st, and read “last quarter” in the query relative to that date. Nothing is removed.
Use as_of to reproduce a past answer or audit a decision: “why did the agent say that on June 1st”. Use query_timestamp when the query itself contains a relative date, or when you want older-but-more-time-relevant memories to rank higher without excluding anything. Both take an ISO 8601 instant. A malformed one is rejected as 422 before the search runs, so a typo costs you nothing.
as_of filters on when a memory was recorded, not on when the event happened. A bulk import is recorded today no matter what timestamp each item carries, so as_of a year ago returns nothing from it. To filter on when the event happened, use occurred_after / occurred_before below.
Under as_of, retrieval skips its graph-expansion pass and answers from semantic, keyword and temporal retrieval only. Expansion follows entity links without a time bound, so it could pull in a memory recorded after the cutoff: a wrong answer for point-in-time recall rather than a ranking artifact. The practical effect is that an as_of search is slightly narrower than the same search without it.

Filtering on event time

occurred_after and occurred_before bound when the thing happened, which is what as_of cannot address: history you import today all shares one record time and spans years of event time.
Each bound is optional, so either one alone is an open-ended window. Both take an ISO 8601 instant, and a malformed one is rejected as 422 before the search runs. occurred_after later than occurred_before is rejected the same way rather than answered with an empty 200, which would send you looking for missing memories instead of at your own arguments. A memory is matched on the event window its text described, falling back to the timestamp you recorded it with. That fallback is the important half: most memories carry no occurred_start at all (see the note above), and they are filtered correctly regardless. An event that spans either edge of your window counts as inside it.

Prompt-ready context

Most callers take the results array and join it into a system prompt. Ask for format: "block" and that is done for you, including the token budget, which a hand-written loop usually skips.
get_context / getContext return the rendered block as a plain string. Over the raw API:
  • results stays populated, so wanting both costs one round trip, not two.
  • The block is numbered plain text with no header, so you compose your own prompt around it.
  • context_max_tokens drops whole memories, lowest-ranked first. Text is never cut mid-sentence: half a fact still reads as a fact.
  • token_estimate is an estimate, derived from a characters-per-token approximation rather than a tokenizer. Treat it as a guide when budgeting a context window, not an exact count.
Omitting format behaves exactly as before.

Keeping the block prompt-cache friendly

Providers discount a prompt prefix they have already seen, and that discount ends at the first byte that differs from last time. A block ordered by relevance reshuffles between turns (relevance is a property of the query, not of the memory), so its first line usually changes and the discount ends there. block_order: "stable" orders by record time and bullets the lines instead of numbering them:
A newly recorded memory appends at the tail, so everything above it is unchanged, and a memory dropping out of recall no longer renumbers the lines below it.
  • It changes the rendered context only. results stays relevance-ordered, and no memory is added or removed.
  • It protects the prefix, which means memories older than any that changed. If the recalled set changes completely, no ordering can keep the block the same.
  • The proxy endpoints take the same knob, where it also moves the block after your system prompt and reports how much of it your provider had already seen. See Prompt caching.

Compound tag filters

tags is a flat list under a single match mode, so it can say “any of these” or “all of these” and nothing else. tag_groups takes a boolean expression instead, which is what you need the moment a filter has two clauses that combine differently, or a clause that excludes. Groups in the list are AND-ed. Each group is either a leaf:
or one of and, or, not, nested however deep you need:
That reads as: tagged project:alpha or project:beta, and not tagged status:archived. match on a leaf takes the same five values as tags_match and defaults to any_strict. The two non-strict modes (any, all) treat an untagged memory as matching, so a leaf written with any widens an otherwise narrow expression rather than narrowing it.
For that reason any and all are rejected inside a not, with 422 validation_error. Negating “matches, including untagged” excludes every untagged memory in the space, so {"not": {"tags": ["status:archived"], "match": "any"}} reads as “exclude archived” and would actually mean “exclude archived and everything that carries no tags at all”. Use any_strict or all_strict, or leave match unset.
The filter is compiled into the search itself, not applied to its output. An excluded memory never competes for a slot in your results, so what you get back is the best n matches that passed the filter rather than whatever survived it.
tag_groups composes with everything else rather than replacing it. tags, user_id, agent_id and session_id are all AND-ed onto your expression, so a scoped query stays scoped when you add a filter, and you can keep a flat tags list and add a single not beside it.
Limits: at most 25 leaves in total and 5 levels of nesting per request; at most 500 tags in any one leaf, and 1,000 across the whole expression. Every tag list must be non-empty, as must every and / or. A leaf naming hundreds of tags is a normal shape, not abuse. One leaf compiles to a single set-containment test, so naming every candidate you might mean and letting the filter decide which of them the space actually has is cheaper than spreading the same tags over separate leaves. The flat tags field keeps its own, much smaller limit of 50, because that one bounds what a single memory may carry, which is a different question.

Matching a name you are not sure of

Tags match as exact strings, which is right when your own code wrote them and wrong when it did not. An agent building a filter out of something a user typed has to guess the exact string a memory was stamped with, and one typo means the memory is unreachable. Set resolve to "fuzzy" on a leaf and its tags become candidates instead of requirements. Each one is compared against the tags your space actually holds, and the close ones are added to the filter:
name:typsecript matches nothing on its own. Resolved, it also matches name:typescript, because the two are 0.47 similar by trigram overlap and the threshold is 0.45. Five things are worth knowing before you rely on it:
A fuzzy leaf always keeps the tags you named and adds to them, so it matches at least everything the same leaf would have matched exactly. Turning it on cannot take a result away from you.
The part before the first colon has to match exactly, and only what follows it is compared. name:chekcout can reach name:checkout and can never reach status:checkout. Without that rule the shared name: prefix alone would make every name 0.38 similar to every other name, and the filter would stop narrowing anything.
A candidate of fewer than four characters after its namespace is never resolved fuzzily, because trigram similarity on two or three characters is noise: api and apis score 0.5, and they are different things. Short names still match exactly, so name:ts finds name:ts.
A candidate has to be about as long as what it matches, within 30%. Without that, name:pipeline would resolve to name:functional pipeline: almost all of the shorter string is inside the longer one, so it scores as well as a real typo does. Resolution is for names typed imperfectly, not for partial name search.
kubernetes and k8s are 0.07 similar, so no threshold reaches one from the other. Abbreviations and alternative spellings come from the write side instead: a multi-text label group has extraction record every name a thing is known by, and then the exact filter finds them.
resolve may only be combined with match of any or any_strict, or left unset. A resolved leaf is a set of alternatives, so all_strict would ask one memory to carry every alternative at once, which matches nothing; that combination is 422 validation_error rather than a filter that silently returns empty. Each candidate contributes at most 10 resolved tags, and the vocabulary it is resolved against is the 5,000 most used tags in the space. Resolution costs one extra lookup on a request that uses it and nothing at all on a request that does not. The same filter is accepted by /v1/reason, where it narrows what the reasoning agent may look at, and by the retrieve tool on the MCP server. The rules above apply identically on all three.
A filter is only as good as the tags it has to work with. If you are tagging by hand on every write, a label taxonomy can have extraction do it for you: name a dimension once, and every memory is classified along it at write time and filterable here.

Debugging what was cut

Every call already computes which results got dropped (by dedup, by min_score, by limit, or by the token budget above), whether or not you ask to see it. Pass receipt: true to get a fetchable id for that record. See Context Receipts.

Scoping within a space

A space is the coarse boundary. user_id, agent_id and session_id partition it, so one space can serve many end users without their memories mixing, and you do not need a space per user.
Scoping is strict, and that is the point:
  • A scoped search returns only memories written under the same scope. Bob’s memories can never appear in Alice’s results.
  • Memories stored without a scope are not returned to a scoped search either. If you turn scoping on for a space that already has history, that history stays visible to unscoped searches but not to scoped ones, so backfill the scope on those memories if you need them.
  • Passing several keys ANDs them: user_id + session_id returns only that user’s memories from that session.
  • Consolidated memories are built per user, so a synthesis never mixes two users’ facts. Sessions roll up into the user, which is what lets the memory improve across conversations.
Tags beginning with anona: are reserved for this and rejected with 422 reserved_tag, since otherwise a hand-written tag could impersonate another user’s scope. Use the scope fields instead.
Scoping costs nothing when unused. A request without these fields behaves exactly as it did before they existed.

Latency modes

mode trades relevance quality against speed.

Error responses

See the full error reference.

Next steps

Reason API

Get a synthesized answer instead of a ranked list.

Graph API

Inspect the entities behind your results.

Context Receipts

See exactly what was cut from a call, and why.