Memory
Long-term facts, session buffer, episodic log, and the scratchpad.
In short: four different things get called "memory," and they solve different problems. What the agent knows about a customer, what was just said, what happened during this run, and where it puts things too big to hold in its head. Keeping them separate is what stops long conversations getting slow and expensive.
"Memory" gets used for four different problems. agentFast keeps them separate, because they have different lifetimes, different costs and different failure modes.
The four tiers
| Tier | Scope | Backed by | Answers |
|---|---|---|---|
| Long-term facts | Per user | pgvector (or Chroma) | "What do I know about this person?" |
| Session buffer | Per session | The run store | "What did we just say?" |
| Episodic log | Per run | The run store | "What happened during this run?" |
| Scratchpad | Per run | Files on disk | "Where do I put the 40-page document?" |
The scratchpad exists because not everything belongs in a context window. An agent researching twelve sources writes them to files and keeps paths in context, rather than carrying the text and running out of room on source nine.
Recall is visible
When long-term memory injects a fact, it writes a memory:recall step into the trace — so a
surprising answer can be traced back to what the agent remembered, not just what it retrieved:
| memory | recall | 2 facts · cust_alice · prefers email contact | — |
| llm | claude-sonnet-4-5 | context included recalled facts | $0.0031 |
Semantic vs lexical retrieval
Every retrieval path here — knowledge-base search, long-term fact recall, and grounding scoring — runs in one of two modes, decided by whether an embeddings key is configured.
embeddings:
provider: auto # auto | openai | voyage | none
auto turns semantic retrieval on whenever an OPENAI_API_KEY or VOYAGE_API_KEY is present, and
leaves it off otherwise. That's why the keyless quickstart still works — with no key, retrieval
falls back to Postgres full-text matching and everything keeps running.
The difference is not cosmetic. Lexical matching can only find documents that share words with the question:
| Question | Document | Lexical | Semantic |
|---|---|---|---|
| "when does the money reach my account?" | "Refunds land in 5-10 business days." | ✗ no shared words | ✓ found |
Real users paraphrase. If your agent answers from a knowledge base, you want a key set.
Vector search only sees rows that have an embedding, so anything ingested before you set a key
is invisible to it. agentFast warns about this at startup and counts the affected rows — run
agentfast reindex to backfill. It's safe to re-run and resumes where it stopped.
Changing embedding model later is the same problem plus a dimension change, so agentFast refuses to boot on a mismatch rather than failing later inside pgvector.
What about running out of context?
That's a different question with its own answer. Memory is what the agent knows; the context window is what currently fits. When a long run outgrows the window, agentFast compacts the history rather than letting the provider reject the call — see Context and compaction.
Turning it off
memory_enabled: false
The runtime treats memory as an optional collaborator, so this is a genuine no-op rather than a stubbed-out path — useful when your agent is stateless by design and you'd rather not carry a vector store.