Skip to content
agentFast

Configuration

agentfast.yaml, environment variables, and budgets.

One file: agentfast.yaml, in your project root. Every value supports ${VAR} and ${VAR:-default} interpolation, so the same file works across environments without a templating step.

A complete file

agent: support                      # support | research | code
sdk: langgraph                      # langgraph | vanilla | crewai
                                    # | claude_agent_sdk | openai_agents_sdk
model: ${AGENTFAST_MODEL:-claude-sonnet-4-5}   # or openai:gpt-4o

database_url: ${AGENTFAST_DATABASE_URL:-postgresql+asyncpg://agentfast:agentfast@localhost:5432/agentfast}
redis_url: ${AGENTFAST_REDIS_URL:-}

kb_dir: kb
scratchpad_dir: .agentfast/scratchpad
vector_backend: pgvector            # pgvector | chroma

harness_llm:                        # agentFast's OWN calls, not your agent's
  provider: auto                    # auto | anthropic | openai | none
  model:                            # defaults per provider
  api_key:                          # else ANTHROPIC_API_KEY / OPENAI_API_KEY
  base_url:                         # any OpenAI-compatible endpoint

embeddings:
  provider: auto                    # auto | openai | voyage | none
  model: text-embedding-3-small     # or voyage-3
  api_key: ${OPENAI_API_KEY:-}      # else OPENAI_API_KEY / VOYAGE_API_KEY
  dimensions:                       # only to truncate text-embedding-3-*
  batch_size: 128

api_host: 0.0.0.0
api_port: 8321
api_auth_token: ${AGENTFAST_API_TOKEN:-}
cors_origins: ["http://localhost:3400", "http://localhost:3401"]

budget:
  max_iterations: 50
  max_tokens: 150000
  max_duration_ms: 300000

memory_enabled: true
complexity_classifier: heuristic    # off | heuristic | llm

compaction:                         # keeps long runs inside the context window
  enabled: auto                     # auto | on | off
  mode: reset                       # reset | tiered
  threshold_tokens: 100000
  keep_recent: 6                    # tiered only
  model:                            # defaults to harness_llm.model
  retain_history: true              # keep dropped messages in the episodic log

team:                                # fan-out settings, used when sdk: team
  max_parallel: 5                    # children in flight at once
  max_tasks: 8                       # hard cap on how many angles the supervisor can plan
  worker_max_iterations: 12          # per-child budget
  worker_max_tokens: 60000           # per-child budget — see the Honest limits note below
  on_child_failure: continue         # continue | fail

guardrails:
  pii_enabled: true
  pii_kinds: [email, card, ssn, phone, ip, secret]
  safety_enabled: true
  injection_action: flag            # flag | block | allow

rate_limits:
  global: "100/minute"
  default_per_tool: "20/minute"

tool_overrides:
  refund_request:
    risk: high
    hitl_mode: suspend              # suspend | defer

observability:
  otel_endpoint: ${OTEL_ENDPOINT:-}
  langsmith:
    api_key: ${LANGSMITH_API_KEY:-}
  pricing:
    claude-sonnet-4-5:
      input_per_1m: 3.00
      output_per_1m: 15.00
      cache_read_per_1m: 0.30

evals:
  cases: evals/cases.yaml
  thresholds:
    containment: 0.8
    tool_use: 0.9
    grounding: 0.7

mcp_servers: []
mcp_provider:
  expose_agent: true
  expose_tools: true
  auto_approve: false

The keys that matter most

KeyDefaultWhy you'd change it
sdklanggraphSwitch orchestration. Nothing else changes
modelclaude-sonnet-4-5Your agent's model. openai:gpt-4o, anthropic:claude-haiku-4-5, or a bare id like gpt-4o. scripted-support runs keyless
harness_llmautoThe model agentFast uses for its own calls — eval judge, llm classifier. Independent of model
tool_overrides{}The most important key here — decides what needs a human
api_auth_tokenunsetSet it before exposing the API anywhere
budget50 / 150k / 5minSized for a support ticket. Research runs want more
complexity_classifierheuristicllm for a smarter pre-flight; off for fixed budgets
compaction.moderesettiered for conversations, where recent turns must stay verbatim. See Context and compaction
compaction.threshold_tokens100000Match your model's window — roughly 60–75% of it. Too low and you pay to summarise more often than you save
team.max_parallel5Children in flight at once, on sdk: team. Admission is re-checked as slots free, so this doesn't decide spend by itself — the run budget still does
team.max_tasks8Hard cap on how many angles the supervisor can plan, regardless of how many it proposes
team.worker_max_tokens60000Per-child token ceiling. A projected-size check runs before every model call, including the closing wrap-up turn — see Agent teams for how the estimate works
team.on_child_failurecontinuecontinue synthesises from whatever succeeded; fail fails the whole team on any one child's failure
redis_urlunsetRequired before scaling past one instance

Environment variables

VariableOverrides
ANTHROPIC_API_KEYAnthropic credential — agent models, harness calls
OPENAI_API_KEYOpenAI credential — agent models, harness calls, embeddings
OPENAI_BASE_URLAny OpenAI-compatible endpoint (Azure, Together, Groq, vLLM, Ollama)
VOYAGE_API_KEYVoyage credential — embeddings only
AGENTFAST_MODELmodel
AGENTFAST_DATABASE_URLdatabase_url
AGENTFAST_REDIS_URLredis_url
AGENTFAST_API_TOKENapi_auth_token

These work through interpolation in the YAML, so the file stays the single source of truth rather than the env quietly winning.

Budgets are a safety mechanism

Checked at the top of every iteration. Exceeding one ends the run as error with a readable reason rather than looping until someone notices the bill:

budget:
  max_iterations: 50      # hard cap on model calls
  max_tokens: 150000      # input + output across the run
  max_duration_ms: 300000 # wall clock
CarefulDurable doesn't mean unbounded

A run that survives restarts and also has no budget is a run that can survive restarts forever. Set budgets deliberately — they're the layer that stops a stuck agent becoming an invoice.

Scaling checklist

Before more than one instance:

  • Set redis_url — otherwise rate limits are per-process and effectively N× your intended limit.
  • Set api_auth_token.
  • Narrow cors_origins to your real front-end origins.
  • Know that live streaming re-attach is per-worker until the bus is Redis-backed. Polling GET /api/runs/{id} works across workers today.