Skip to content
agentFast

Memory

Long-term facts, session buffer, episodic log, and the scratchpad.

In short: four different things get called "memory," and they solve different problems. What the agent knows about a customer, what was just said, what happened during this run, and where it puts things too big to hold in its head. Keeping them separate is what stops long conversations getting slow and expensive.

"Memory" gets used for four different problems. agentFast keeps them separate, because they have different lifetimes, different costs and different failure modes.

The four tiers

TierScopeBacked byAnswers
Long-term factsPer userpgvector (or Chroma)"What do I know about this person?"
Session bufferPer sessionThe run store"What did we just say?"
Episodic logPer runThe run store"What happened during this run?"
ScratchpadPer runFiles on disk"Where do I put the 40-page document?"

The scratchpad exists because not everything belongs in a context window. An agent researching twelve sources writes them to files and keeps paths in context, rather than carrying the text and running out of room on source nine.

Recall is visible

When long-term memory injects a fact, it writes a memory:recall step into the trace — so a surprising answer can be traced back to what the agent remembered, not just what it retrieved:

run trace
memoryrecall2 facts · cust_alice · prefers email contact—
llmclaude-sonnet-4-5context included recalled facts$0.0031
The recall step is a first-class part of the run, visible in the dashboard and on the live stream. If an agent says something you didn't expect, this is where you look first.

Semantic vs lexical retrieval

Every retrieval path here — knowledge-base search, long-term fact recall, and grounding scoring — runs in one of two modes, decided by whether an embeddings key is configured.

embeddings:
  provider: auto        # auto | openai | voyage | none

auto turns semantic retrieval on whenever an OPENAI_API_KEY or VOYAGE_API_KEY is present, and leaves it off otherwise. That's why the keyless quickstart still works — with no key, retrieval falls back to Postgres full-text matching and everything keeps running.

The difference is not cosmetic. Lexical matching can only find documents that share words with the question:

QuestionDocumentLexicalSemantic
"when does the money reach my account?""Refunds land in 5-10 business days."✗ no shared words✓ found

Real users paraphrase. If your agent answers from a knowledge base, you want a key set.

CarefulEnabling it later needs a reindex

Vector search only sees rows that have an embedding, so anything ingested before you set a key is invisible to it. agentFast warns about this at startup and counts the affected rows — run agentfast reindex to backfill. It's safe to re-run and resumes where it stopped.

Changing embedding model later is the same problem plus a dimension change, so agentFast refuses to boot on a mismatch rather than failing later inside pgvector.

What about running out of context?

That's a different question with its own answer. Memory is what the agent knows; the context window is what currently fits. When a long run outgrows the window, agentFast compacts the history rather than letting the provider reject the call — see Context and compaction.

Turning it off

memory_enabled: false

The runtime treats memory as an optional collaborator, so this is a genuine no-op rather than a stubbed-out path — useful when your agent is stateless by design and you'd rather not carry a vector store.