relevance_score.
Parameters
Response
results list directly, and an empty list when nothing matches.
Response fields keep the API’s own snake_case names in TypeScript too, so a result is
r.relevance_score rather than r.relevanceScore.
Let Anona choose the space
Sendroute: "auto" in place of space_id when you know the question but not
which space holds the answer. Anona picks one space, searches it, and reports
which one in searched. Exactly one of the two fields must be present: sending
both, or neither, is 422 route_conflict.
Read
searched. A query answered from the wrong space looks exactly like a space
that holds nothing, and this is the only field that tells the two apart.
When no space fits
An abstention is a200 with an empty results array and one searched entry
whose space_id is null. The entry is present, rather than searched being
empty, because the reason has to be carried somewhere.
reason is no_space_fits when the model reported that none of your spaces holds
the answer, and low_confidence when it picked one but not confidently enough to
act on.
An abstained query is charged. A retrieve that names a space and matches
nothing is charged today, and a routed retrieve that matches nothing is charged
the same way. The usage row groups under your organization’s default space, or
under
unrouted if you have none set.fallback_space_id each is
422 no_route_target: no_candidates (no space of yours is eligible for
routing), router_disabled (routing is not switched on for this deployment),
answer_outside_candidates (the model named a space never offered to it) and
model_unavailable (the decision provider could not be reached). Sending a
fallback_space_id searches that space instead, with stage: "fallback" and the
reason carried along.
Routing applies to retrieve only
route is accepted on POST /v1/retrieve and the MCP retrieve tool.
POST /v1/reason requires a space_id, and so do the
SDKs’ get_context and retrieve_receipt, which are renderings of this endpoint
that do not take route. Those calls are addressed only.
Which spaces are eligible, and how to describe a space so it gets chosen, are in
Automatic space routing. One difference from the
write side is worth knowing here: a shared space you hold as a viewer is
eligible for a routed read, because a viewer can read it, while a routed write
needs developer.
retrieve still returns a list, so nothing about existing code changes.
searched rides along as an attribute on it, and is None on an addressed read.
Memory types
Every result carries amemory_type. The same content can exist at more than one level,
as raw evidence and as a distilled version of it, so filtering by type asks for exactly
the layer you want.
By default (
prefer_observations: true) a note is returned in place of the facts it
was built from, so you see one distilled result instead of the same content several
times. Set prefer_observations: false to receive the raw facts alongside their note.
A note names its evidence either way. Every result carries source_ids, the
memory_id of each memory it was synthesized from, so a distilled answer can still be
attributed, cited or audited back to the raw writes behind it. source_ids is empty on
a fact, which is evidence rather than a synthesis.
A note’s own metadata is empty and its document_id is null, because both belong to
the facts it was built from rather than to the synthesis. To read them, set
prefer_observations: false: the raw facts come back alongside their note, each
carrying the metadata you wrote it with, and source_ids tells you which of them
belong to which note.
Entities and time
Each result carries structured signal beyond the raw text.
The same entities are exposed as a browsable graph. See Graph.
Searching over time
Every memory carries two independent times, and knowing which is which is most of the work:
They differ whenever you import history: a memory about last June that you
upload today has an event time of last June and a record time of today.
The comes back with
timestamp you send is returned as timestamp. It is not copied into
occurred_start. Those two are filled only from a date found in the memory’s
own text, so this write:timestamp: "2024-06-15T10:00:00Z" and both occurred_*
fields null. That is not a gap in the data. occurred_start: null is the
honest answer to “what event window did the text describe”, and the window
filter below falls back to timestamp, so this memory still matches a June 2024
window. Read timestamp when you want to know when something happened; read
occurred_start only when you specifically want a span the text spelled out.as_of and query_timestamp are not variants of each other: one filters, the
other ranks, and both work on record time.
as_of — point-in-time recall. What did the space know on June 1st? Anything
recorded after that instant is dropped, however relevant.
query_timestamp — move “now”. Rank as though it were June 1st, and read “last
quarter” in the query relative to that date. Nothing is removed.
as_of to reproduce a past answer or audit a decision: “why did the agent
say that on June 1st”. Use query_timestamp when the query itself contains a
relative date, or when you want older-but-more-time-relevant memories to rank
higher without excluding anything.
Both take an ISO 8601 instant. A malformed one is rejected as 422 before the
search runs, so a typo costs you nothing.
Under
as_of, retrieval skips its graph-expansion pass and answers from
semantic, keyword and temporal retrieval only. Expansion follows entity links
without a time bound, so it could pull in a memory recorded after the cutoff:
a wrong answer for point-in-time recall rather than a ranking artifact. The
practical effect is that an as_of search is slightly narrower than the same
search without it.Filtering on event time
occurred_after and occurred_before bound when the thing happened, which
is what as_of cannot address: history you import today all shares one record
time and spans years of event time.
422 before the search
runs. occurred_after later than occurred_before is rejected the same way
rather than answered with an empty 200, which would send you looking for
missing memories instead of at your own arguments.
A memory is matched on the event window its text described, falling back to the
timestamp you recorded it with. That fallback is the important half: most
memories carry no occurred_start at all (see the note above), and they are
filtered correctly regardless. An event that spans either edge of your window
counts as inside it.
Prompt-ready context
Most callers take the results array and join it into a system prompt. Ask forformat: "block" and that is done for you, including the token budget, which
a hand-written loop usually skips.
get_context / getContext return the rendered block as a plain string. Over the raw
API:
resultsstays populated, so wanting both costs one round trip, not two.- The block is numbered plain text with no header, so you compose your own prompt around it.
context_max_tokensdrops whole memories, lowest-ranked first. Text is never cut mid-sentence: half a fact still reads as a fact.token_estimateis an estimate, derived from a characters-per-token approximation rather than a tokenizer. Treat it as a guide when budgeting a context window, not an exact count.
format behaves exactly as before.
Keeping the block prompt-cache friendly
Providers discount a prompt prefix they have already seen, and that discount ends at the first byte that differs from last time. A block ordered by relevance reshuffles between turns (relevance is a property of the query, not of the memory), so its first line usually changes and the discount ends there.block_order: "stable" orders by record time and bullets the lines instead of numbering
them:
- It changes the rendered
contextonly.resultsstays relevance-ordered, and no memory is added or removed. - It protects the prefix, which means memories older than any that changed. If the recalled set changes completely, no ordering can keep the block the same.
- The proxy endpoints take the same knob, where it also moves the block after your system prompt and reports how much of it your provider had already seen. See Prompt caching.
Compound tag filters
tags is a flat list under a single match mode, so it can say “any of these”
or “all of these” and nothing else. tag_groups takes a boolean expression
instead, which is what you need the moment a filter has two clauses that
combine differently, or a clause that excludes.
Groups in the list are AND-ed. Each group is either a leaf:
and, or, not, nested however deep you need:
project:alpha or project:beta, and not
tagged status:archived.
match on a leaf takes the same five values as tags_match and defaults to
any_strict. The two non-strict modes (any, all) treat an untagged
memory as matching, so a leaf written with any widens an otherwise narrow
expression rather than narrowing it.
The filter is compiled into the search itself, not applied to its output. An
excluded memory never competes for a slot in your results, so what you get back
is the best n matches that passed the filter rather than whatever survived it.
tag_groups composes with everything else rather than replacing it. tags,
user_id, agent_id and session_id are all AND-ed onto your expression, so
a scoped query stays scoped when you add a filter, and you can keep a flat
tags list and add a single not beside it.and / or.
A leaf naming hundreds of tags is a normal shape, not abuse. One leaf compiles
to a single set-containment test, so naming every candidate you might mean and
letting the filter decide which of them the space actually has is cheaper than
spreading the same tags over separate leaves. The flat tags field keeps its
own, much smaller limit of 50, because that one bounds what a single memory may
carry, which is a different question.
Matching a name you are not sure of
Tags match as exact strings, which is right when your own code wrote them and wrong when it did not. An agent building a filter out of something a user typed has to guess the exact string a memory was stamped with, and one typo means the memory is unreachable. Setresolve to "fuzzy" on a leaf and its tags become candidates instead of
requirements. Each one is compared against the tags your space actually holds,
and the close ones are added to the filter:
name:typsecript matches nothing on its own. Resolved, it also matches
name:typescript, because the two are 0.47 similar by trigram overlap and the
threshold is 0.45.
Five things are worth knowing before you rely on it:
It widens, it never narrows
It widens, it never narrows
A fuzzy leaf always keeps the tags you named and adds to them, so it matches at
least everything the same leaf would have matched exactly. Turning it on cannot
take a result away from you.
Short candidates are matched literally
Short candidates are matched literally
A candidate of fewer than four characters after its namespace is never resolved
fuzzily, because trigram similarity on two or three characters is noise:
api and apis score 0.5, and they are different things. Short names still
match exactly, so name:ts finds name:ts.A word of a longer tag is not a misspelling of it
A word of a longer tag is not a misspelling of it
A candidate has to be about as long as what it matches, within 30%. Without
that,
name:pipeline would resolve to name:functional pipeline: almost all of
the shorter string is inside the longer one, so it scores as well as a real typo
does. Resolution is for names typed imperfectly, not for partial name search.It fixes misspellings, not abbreviations
It fixes misspellings, not abbreviations
kubernetes and k8s are 0.07 similar, so no threshold reaches one from the
other. Abbreviations and alternative spellings come from the write side
instead: a multi-text
label group has extraction
record every name a thing is known by, and then the exact filter finds them.resolve may only be combined with match of any or any_strict, or left
unset. A resolved leaf is a set of alternatives, so all_strict would ask one
memory to carry every alternative at once, which matches nothing; that
combination is 422 validation_error rather than a filter that silently
returns empty.
Each candidate contributes at most 10 resolved tags, and the vocabulary it is
resolved against is the 5,000 most used tags in the space. Resolution costs one
extra lookup on a request that uses it and nothing at all on a request that
does not.
The same filter is accepted by /v1/reason, where it
narrows what the reasoning agent may look at, and by the retrieve tool on the
MCP server. The rules above apply identically on all three.
Debugging what was cut
Every call already computes which results got dropped (by dedup, bymin_score, by limit, or by the token budget above), whether or not you ask
to see it. Pass receipt: true to get a fetchable id for that record. See
Context Receipts.
Scoping within a space
A space is the coarse boundary.user_id, agent_id and session_id partition
it, so one space can serve many end users without their memories mixing, and you do
not need a space per user.
- A scoped search returns only memories written under the same scope. Bob’s memories can never appear in Alice’s results.
- Memories stored without a scope are not returned to a scoped search either. If you turn scoping on for a space that already has history, that history stays visible to unscoped searches but not to scoped ones, so backfill the scope on those memories if you need them.
- Passing several keys ANDs them:
user_id+session_idreturns only that user’s memories from that session. - Consolidated memories are built per user, so a synthesis never mixes two users’ facts. Sessions roll up into the user, which is what lets the memory improve across conversations.
anona: are reserved for this and rejected with
422 reserved_tag, since otherwise a hand-written tag could impersonate another
user’s scope. Use the scope fields instead.
Scoping costs nothing when unused. A request without these fields behaves
exactly as it did before they existed.
Latency modes
mode trades relevance quality against speed.
Error responses
See the full error reference.
Next steps
Reason API
Get a synthesized answer instead of a ranked list.
Graph API
Inspect the entities behind your results.
Context Receipts
See exactly what was cut from a call, and why.