On-premise, sandboxed AI agent platform for regulated sectors
Public sector, finance, and healthcare — where data must stay on your infrastructure.
Cat-Agent is a Python framework for building LLM agents that run fully on-premise. You get:
- Agents —
Assistant, multi-agentGroupChat, graph workflows (StateGraph) - Tools —
@tooldecorator, RAG, code interpreter (Docker or WASM), MCP - Serve & deploy — FastAPI HTTP server + Nomad deploy via
cat-agent deploy - Scheduling — recurring collect-and-report jobs (email / webhook)
- Synthesis — Markdown draft → interviewed spec → sandboxed
@tool - Security — air-gap mode, encrypted storage, audit trail, PII redaction
- Tracing — structured JSONL execution traces (schema v1.0, RunLimits, redaction, token/cost totals)
- Context — conversation-window management (observation masking with residue, compaction,
fold()) - Failure analysis — MAST taxonomy over traces (14 modes; Tier-1 detectors + optional LLM judge)
The base install is lightweight (OpenAI-compatible client + native Rust RAG). Heavy backends (Transformers, LlamaCpp, MLX) are optional extras.
- Zero to hero
- Architecture
- Installation
- Configuration
- Core concepts
- CLI reference
- Deploy to Nomad
- Scheduled reports
- Tool synthesis
- Security & compliance
- Advanced topics
- Examples index
- Development
Follow this path in order. Each step links to a runnable example in examples/.
| Step | Goal | Command / example |
|---|---|---|
| 0 | Clone & configure | cp .env.example .env |
| 1 | Install | pip install cat-agent |
| 2 | First agent | 2.1 API · 2.2 Transformers · 2.3 LlamaCpp · 2.4 MLX |
| 3 | HTTP serve | examples/serve_fastapi/ |
| 4 | Nomad deploy | cat-agent-stack + cat-agent deploy |
| 5 | Multi-agent team | examples/multi_agent/ |
| 6 | Scheduled reports | examples/scheduling/ |
| 7 | Tool synthesis | examples/synthesis/from_draft/ |
| 8 | Production hardening | Security & compliance |
git clone https://github.com/kemalcanbora/cat-agent.git
cd cat-agent
cp .env.example .env
# Edit .env — at minimum set your LLM gateway or Ollama credentialsCat-Agent loads .env automatically on import cat_agent and when using the CLI. Shell exports override file values. Point elsewhere with CAT_AGENT_ENV_FILE=/path/to/custom.env.
Requires Python 3.10+. On zsh, quote extras:
pip install cat-agent # base: agents, tools, native RAG
pip install 'cat-agent[serve]' # FastAPI HTTP server
pip install 'cat-agent[platform]' # Nomad deploy
pip install 'cat-agent[scheduler]' # scheduled reports
pip install 'cat-agent[rag]' # doc parsing + ONNX embeddings
pip install 'cat-agent[local]' # transformers + llama + wasm
pip install 'cat-agent[all]' # everythingEvery backend uses the same pattern: define a @tool, attach it to Assistant, call run. Pick the LLM that matches your hardware.
Works with OpenAI, Ollama Cloud, or any on-prem gateway. No extra install beyond cat-agent.
from cat_agent.agents import Assistant
from cat_agent.tools import tool
@tool
def sum_two_number(a: float, b: float) -> float:
"""Adds two numbers."""
return a + b
bot = Assistant(
llm={'model': 'gpt-4o-mini', 'model_type': 'oai', 'model_server': 'https://api.openai.com/v1'},
function_list=[sum_two_number],
)
print(list(bot.run([{'role': 'user', 'content': 'What is 2 + 3?'}]))[-1]['content'])# Ollama Cloud — set OLLAMA_API_KEY + OLLAMA_API_BASE in .env
python examples/tool_decorator/sum_two_number.py
python examples/multi_agent/team_example.py # three models on one gatewayNomad-deployable when paired with agent.yaml (model.type: api). See Step 4.
Local PyTorch models — CUDA or Apple MPS.
pip install 'cat-agent[transformers]'
python examples/transformers_math_guy/math_guy.pybot = Assistant(
llm={
'model': 'Qwen/Qwen3.5-0.8B',
'model_type': 'transformers',
'device': 'cuda:0', # or 'mps' on Mac
},
function_list=['sum_numbers'],
)For models with native HF tool-calling templates (e.g. FunctionGemma), add
'use_chat_template_tools': True to the config — see LLM backends.
Local-only — no agent.yaml / Nomad deploy (model weights live on your machine).
Quantised GGUF models via llama-cpp-python. CPU or GPU offload.
pip install 'cat-agent[llama]'
python examples/llama_cpp_math_guy/llama_cpp_example.py
# HTTP serve (still local-only — no agent.yaml):
cat-agent serve --factory llama_cpp_example:registrybot = Assistant(
llm={
'model_type': 'llama_cpp',
'repo_id': 'Salesforce/xLAM-2-3b-fc-r-gguf',
'filename': 'xLAM-2-3B-fc-r-F16.gguf',
'n_gpu_layers': -1,
},
function_list=['sum_two_number'],
)Multimodal: examples/llama_cpp_vision/. Local-only for Nomad.
Fast local inference on Mac with mlx-lm.
pip install 'cat-agent[mlx]'
python examples/mlx_lm_math_guy/math_guy.pybot = Assistant(
llm={
'model_type': 'mlx_lm',
'model': 'mlx-community/Qwen3.5-0.8B-MLX-8bit',
},
function_list=['sum_numbers'],
)Local-only for Nomad. For deployable HTTP agents use the API path (Step 2.1 + Step 3).
| Backend | Extra | Example | Nomad deploy |
|---|---|---|---|
oai |
base | tool_decorator/ |
yes (with agent.yaml) |
transformers |
[transformers] |
transformers_math_guy/ |
no |
llama_cpp |
[llama] |
llama_cpp_math_guy/ |
no |
mlx_lm |
[mlx] |
mlx_lm_math_guy/ |
no |
Keep agents loaded in-process and call them via REST:
pip install 'cat-agent[serve]'
python examples/serve_fastapi/serve_math_guy.py
curl -s http://127.0.0.1:8080/agents/calculator/run \
-H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"sum 42 and 58"}]}'Every deployable agent exposes a zero-arg registry() factory:
from cat_agent.serve import AgentRegistry
def registry() -> AgentRegistry:
reg = AgentRegistry()
reg.register(my_assistant, name='calculator')
return regNomad deploy needs two repos:
| Repo | Role |
|---|---|
| cat-agent (this repo) | Agent code, CLI, cat-agent deploy |
| cat-agent-stack | Local HashiCorp stack: Consul + Vault + Nomad + LiteLLM + Traefik |
Clone them side by side:
git clone https://github.com/kemalcanbora/cat-agent.git
git clone https://github.com/kemalcanbora/cat-agent-stack.gitpip install 'cat-agent[serve,platform]'
cd cat-agent-stack
cp .env.example .env # VAULT_TOKEN=root + Ollama/OpenAI keys
export CAT_AGENT_STACK_DIR=$PWD
export CAT_AGENT_CONFIG=$PWD/cat-agent.config.toml
cat-agent stack bootstrap # docker compose up + Vault seed + demo team key
cat-agent doctor # must show docker_network: cat-agent-stack_hashicorpStack details, Vault key layers, and LAN DNS: cat-agent-stack README.
From the cat-agent checkout, point at a directory with agent.yaml + registry():
cd ../cat-agent
cat-agent deploy --dir examples/serve_fastapiWhat deploy does:
- Reads
agent.yaml(team, name, model alias, resources, env) - Validates the model exists on the live LiteLLM gateway
- Builds a Docker image (local registry mode — no push by default)
- Renders and submits a Nomad job
- Prints the Traefik URL (default
http://{team}-{name}.localhost:8088)
curl -sS http://demo-calculator.localhost:8088/readyz
cat-agent ls # all deployed agents
cat-agent status demo/calculator # health + URL
cat-agent logs demo/calculator # tail allocation logsDeploy more examples:
cat-agent deploy --dir examples/multi_agent
cat-agent deploy --dir examples/schedulingRemoves the Nomad job and stops the container. Does not stop the stack.
cat-agent rm demo/calculator --yes
cat-agent ls # should be empty (or list remaining agents)To stop the whole infrastructure:
cd ../cat-agent-stack
cat-agent stack downFull operator guide: Deploy to Nomad.
Three agents, three models, one round-robin GroupChat:
python examples/multi_agent/team_example.py
cat-agent deploy --dir examples/multi_agentDetails: examples/multi_agent/README.md.
Two examples — pick the one that matches your question:
| File | What it shows |
|---|---|
scheduled_report_example.py |
Local loop; Job(interval_seconds=60) visible in code |
schedule_agent.py |
Deployable HTTP agent; seeds a job into SQLite/Postgres |
pip install 'cat-agent[scheduler]'
python examples/scheduling/scheduled_report_example.py # runs ticks locally
cat-agent deploy --dir examples/scheduling # HTTP + persisted job
cat-agent schedule run-due # worker executes due jobsDetails: examples/scheduling/README.md.
Business users write a Markdown draft; Cat-Agent interviews, confirms, and synthesises a WASM-validated @tool:
pip install 'cat-agent[wasm,synthesis]'
cat-agent synth init my_tool --lang en
cat-agent synth run my_tool_draft.mdDetails: examples/synthesis/from_draft/README.md.
Enable air-gap, encryption, and audit before going live:
# .env
CAT_AGENT_OFFLINE=1
CAT_AGENT_ENCRYPT_AT_REST=1
CAT_AGENT_AUDIT=1
cat-agent offline-check --strict
cat-agent encrypt-storage --workspace ./workspaceSee Security & compliance and deploy/README.md for the air-gapped Docker package.
flowchart TB
subgraph dev ["Your code"]
A[Assistant / GroupChat / Graph]
T["@tool functions"]
R["registry() factory"]
end
subgraph runtime ["Cat-Agent runtime"]
LLM[LLM backends]
NAT["cat_agent._native\nBM25 · HNSW · PDF · tokenizer"]
SCH[JobStore + scheduling]
SYN[ToolSmith + WASM sandbox]
end
subgraph serve ["HTTP serve"]
API[FastAPI /agents/name/run]
end
subgraph platform ["Platform (optional)"]
NOM[Nomad jobs]
GW[LiteLLM gateway]
V[Vault secrets]
end
A --> T
A --> LLM
T --> NAT
R --> API
API --> A
R --> NOM
NOM --> API
GW --> LLM
V --> GW
SCH --> API
| Word | Means | CLI / path |
|---|---|---|
| deploy | Build + submit an agent to Nomad | cat-agent deploy |
| promote | Point a group's active.json at a synthesised tool |
cat-agent synth promote |
| run (async) | In-process HTTP job for a served agent | POST /agents/{name}/jobs |
| report job | Scheduled collect → LLM report → delivery | cat-agent schedule … |
deploy/ package |
Air-gap library image (+ optional k8s CronJob) | deploy/docker-compose.yml |
cat-agent deploy never promotes WASM tools. cat-agent synth promote never touches Nomad.
| Extra | Installs | Use when |
|---|---|---|
rag |
doc parsers, ONNX runtime | document Q&A, hybrid search |
transformers |
PyTorch, HuggingFace | local GPU models |
llama |
llama-cpp-python | GGUF models (+ vision) |
mlx |
mlx-lm | Apple Silicon local models |
wasm |
wasmtime | WASM code interpreter |
wasm-bundled |
wasmtime in wheel | air-gap (no runtime download) |
mcp |
MCP SDK | Model Context Protocol tools |
scheduler |
SQLAlchemy, APScheduler | scheduled reports |
serve |
FastAPI, uvicorn | HTTP agent server |
platform |
Jinja2, YAML, import-linter | Nomad deploy |
synthesis |
PyYAML | ToolSpec helpers |
email |
Resend | optional email provider |
pii |
Presidio | NER-based PII redaction |
otel |
OpenTelemetry | trace export |
code_interpreter |
Jupyter stack | Docker code interpreter server |
local |
transformers + llama + wasm | all local backends |
all |
everything above | full dev environment |
Before publishing, verify the wheel like an end user:
./scripts/install_consumer.sh rag examples/rag_keyword/rust_keyword_search_demo.pyPublished wheels ship cat_agent._native (PyO3). No Python fallbacks for these paths:
| Module | Used by |
|---|---|
| BM25 index | KeywordSearch, RagIndex |
| HNSW vector index | VectorSearch, VectorIndex |
| Hash embeddings | offline vector recall |
| Tokenizer / truncation | count_tokens, truncate_messages |
| Document chunking | DocParser.split_doc_to_chunk |
| PDF text extraction | .pdf ingestion |
import cat_agent._native as native
print(native.__version__)Source installs build via maturin; published wheels do not require a local Rust toolchain.
| Concern | Where | Examples |
|---|---|---|
| Secrets (API keys) | .env only — never in yaml |
OLLAMA_API_KEY, OPENAI_API_KEY |
| API base URL | .env (local) / gateway on deploy |
OLLAMA_API_BASE, OPENAI_BASE_URL |
| Model id | agent.yaml |
model.alias, env.CAT_AGENT_LLM_MODEL_* |
| Scheduler DSN | agent.yaml or .env |
CAT_AGENT_SCHEDULER_DSN |
| Non-secret tuning | agent.yaml env: |
LOG_LEVEL, per-agent model ids |
# agent.yaml (deploy manifest — no API keys)
team: demo
name: calculator
runtime:
entrypoint: serve_math_guy:registry
model:
type: api
alias: minimax-m3
env:
LOG_LEVEL: INFOOn Nomad deploy, CAT_AGENT_MANAGED=1 prevents a baked .env from redirecting the LLM off the gateway. Use llm_config_from_env() in factories — CAT_AGENT_LLM_* beat legacy OPENAI_* / OLLAMA_*.
Full template: .env.example.
| Class | Role |
|---|---|
Agent |
Base class — run / arun streaming |
Assistant |
Function-calling agent (most common) |
FnCallAgent |
Lower-level tool loop |
ReActChat |
ReAct-style reasoning |
DocQAAgent |
Document Q&A with retrieval (BasicDocQA) |
ParallelDocQA |
Exhaustive per-chunk doc QA (high recall, high cost; see below) |
GroupChat |
Multi-agent round-robin or auto-router |
Router |
Route queries to specialised agents |
GraphAgent |
Compiled DAG from StateGraph |
Assistant with retrieval asks the model to call search and answers from top-k snippets — cheap and usually enough.
ParallelDocQA LLM-scans every document chunk in parallel, keeps passages that claim to answer, then runs GenKeyword + retrieval + a summary pass. That raises recall on long corpora and costs one member LLM call per chunk (default hard cap max_chunks=32; over-budget sets fail with the chunk count and limit). Use estimate_member_calls(messages) before spending.
Example: examples/parallel_doc_qa/.
Register plain functions with @tool — schemas come from type hints and docstrings:
from cat_agent.tools import tool
@tool
def my_tool(query: str) -> str:
"""Search internal docs.
Args:
query: Natural language search query
"""
return "..."Network tools (web_search, image_search, web_extractor) are opt-in — not in the default registry. Enable with enable_optional_tools(...). Blocked when CAT_AGENT_OFFLINE=1.
Tool schemas on the wire are OpenAI JSON Schema objects (parameters.type == "object"). Older list-style parameter declarations (e.g. hub tools) are converted automatically when exported.
| Sync | Async |
|---|---|
run |
arun |
run_nonstream |
arun_nonstream |
The async path does not stream tokens — it yields complete message lists. Multiple tool calls in one turn run concurrently via asyncio.gather. Use arun from FastAPI/Jupyter; calling sync run() inside a running event loop blocks and emits a warning.
Wire and dict yields use tool_calls as the canonical field (not a derived function_call key). In-process code can still read Message.function_call for the first call.
Configured via rag_searchers (default: keyword + front-page):
| Searcher | Backend |
|---|---|
keyword_search |
Rust BM25 (persistent index) |
vector_search |
Rust HNSW (hash or ONNX embeddings) |
front_page_search |
Heuristic first-chunk boost |
hybrid_search |
Fusion when multiple searchers configured |
Indexes persist under workspace/storage/keyword_indexes/ and vector_indexes/.
Cross-session memory with encrypted SQLite + vector recall:
agent = Assistant(
llm=llm_cfg,
memory_cfg={
'scope': 'user:alice',
'top_k': 5,
'auto_record': True,
'auto_summarize': True,
'session_window_tokens': 8000,
},
)Example: examples/long_term_memory/.
Compose agents and tools into branching graphs:
from cat_agent.graph import StateGraph, AgentNode, FunctionNode, END
app = (
StateGraph()
.add_node(FunctionNode("classify", classify_fn))
.add_node(AgentNode("math", math_agent))
.set_entry("classify")
.add_conditional_edges("classify", route_fn)
.add_edge("math", END)
.compile(name="MathGraph")
)Visualise with MermaidExporter or OpenTelemetryHandler. Example: examples/graph/.
cat_agent.multi_agent provides blackboard artifacts, handoff, and ask-agent tools for team workflows. Example: examples/multi_agent/team_example.py.
Handlers are opt-in — attach to agents or compiled graphs:
from cat_agent.observability import PrintHandler, CallbackHandler, MermaidExporter
bot = Assistant(llm=..., handlers=[PrintHandler()])For structured run traces (JSONL, token totals, RunLimits), context-window management, and MAST failure analysis, see Features — tracing, context, failure analysis.
| Event | When |
|---|---|
run.start / run.end |
Agent lifecycle |
node.start / node.end |
Graph node execution |
llm.start / llm.end |
LLM calls |
tool.start / tool.end |
Tool invocations |
Enable trace logging: CAT_AGENT_TRACE=1. Langfuse example: examples/langfuse/.
| Backend | model_type |
Extra | Tool calling |
|---|---|---|---|
| OpenAI-compatible | oai |
base install | Native tools / tool_calls on the wire |
| Transformers | transformers |
[transformers] |
Prompt path (Nous <tool_call> markup) or native HF chat template (use_chat_template_tools: true) |
| LlamaCpp | llama_cpp |
[llama] |
Prompt path |
| LlamaCpp Vision | llama_cpp_vision |
[llama] |
Prompt path |
| MLX-LM | mlx_lm |
[mlx] |
Prompt path |
Native HF chat template tools — Models like Google FunctionGemma ship with a
Jinja chat template that formats tool schemas and parses tool calls natively.
Set use_chat_template_tools: true to use it instead of the prompt-based path:
bot = Assistant(
llm={
'model': 'google/functiongemma-270m-it',
'model_type': 'transformers',
'use_chat_template_tools': True,
'device': 'cuda:0', # or 'mps'
'generate_cfg': {'max_new_tokens': 128, 'do_sample': False},
},
function_list=['sum_numbers'],
)Any HF model whose tokenizer supports apply_chat_template(..., tools=...) can
use this path — it is not FunctionGemma-specific.
Serve local Qwen (or any tools-capable model) through Ollama / vLLM / llama.cpp server with model_type: oai to get the native path. Both paths keep every tool call the model emits in one turn — there is no single-call trim.
parallel_tool_calls is not injected by the oai backend; leave it unset unless your gateway needs an explicit value (OpenAI allows parallel by default when the key is absent; some models reject the parameter entirely).
Live check against any OpenAI-compatible endpoint: examples/native_parallel/.
cat-agent <command> [options]| Command | Purpose |
|---|---|
deploy --dir <path> |
Build image + submit Nomad job from agent.yaml |
ls |
List deployed agents |
status <team>/<name> |
Job health and URL |
logs <team>/<name> |
Tail allocation logs |
rm <team>/<name> --yes |
Tear down deployment |
rollback <team>/<name> |
Revert to previous version |
doctor |
Platform readiness (network, gateway, config) |
build-base |
Build shared runtime base image |
Requires sibling cat-agent-stack repo:
| Command | Purpose |
|---|---|
stack bootstrap |
docker compose up + Vault seed |
stack up / stack down |
Start / stop infrastructure |
stack seed |
Inject LLM credentials into Vault |
stack compose |
Raw docker compose passthrough |
Auto-discovers ../cat-agent-stack or $CAT_AGENT_STACK_DIR.
| Command | Purpose |
|---|---|
serve --factory <mod:fn> [--port 8080] |
HTTP server for named agents |
Requires [scheduler]:
| Command | Purpose |
|---|---|
schedule add --user U --topic T --every H --channel C --target T |
Create report job |
schedule list [--user U] |
List jobs |
schedule rm <job_id> |
Delete job |
schedule run <job_id> [--dry-run] |
Run one job now |
schedule run-due [--limit N] |
Claim and execute due jobs (CronJob path) |
schedule doctor |
Validate DSN, channels, LLM creds |
Also available as cat-agent-scheduler entry point.
Requires [wasm,synthesis]:
| Command | Purpose |
|---|---|
synth init <name> [--lang en] |
Blank Markdown draft template |
synth run <draft.md> |
Interview + synthesise tool |
synth promote / synth demote |
Group active tool pointer |
synth list / synth gc |
Inventory / cleanup |
synth share / synth adopt |
Cross-host artifact transfer |
Promote workflow: docs/synthesis-promote.md.
| Command | Purpose |
|---|---|
offline-check [--strict] |
Air-gap readiness report |
fetch-runtime --output <dir> |
Copy WASM assets for offline transfer |
encrypt-storage [--workspace <dir>] |
Encrypt plaintext caches and indexes |
encrypt-cache --path <dir> |
Encrypt one cache directory |
audit-verify --path <file> |
Verify tamper-evident audit chain |
audit-export --path <file> --output <file> |
Export audit records |
Nomad deploy uses the sibling stack repo: github.com/kemalcanbora/cat-agent-stack (Consul, Vault, Nomad, LiteLLM gateway, Traefik). Agent packaging and the cat-agent deploy CLI live in this repo.
Install platform extra and bootstrap the stack once:
pip install 'cat-agent[serve,platform]'
git clone https://github.com/kemalcanbora/cat-agent-stack.git
cd cat-agent-stack
cp .env.example .env
export CAT_AGENT_STACK_DIR=$PWD
cat-agent stack bootstrap # compose up + Vault seed + demo team virtual key
cat-agent doctor # must show docker_network: cat-agent-stack_hashicorp
cd ../cat-agent
cat-agent deploy --dir examples/serve_fastapi
curl -sS http://demo-calculator.localhost:8088/readyz| Command | What it does |
|---|---|
cat-agent deploy --dir <path> |
Build image + submit Nomad job from agent.yaml |
cat-agent ls |
List deployed agents (team/name) |
cat-agent status <team>/<name> |
Job health, allocation, public URL |
cat-agent logs <team>/<name> |
Stream stdout/stderr from the running allocation |
cat-agent rm <team>/<name> --yes |
Stop and remove the Nomad job (agent gone; stack keeps running) |
cat-agent rollback <team>/<name> |
Revert to the previous deployment version |
cat-agent doctor |
Platform readiness (network, gateway, config file) |
cat-agent stack down |
Stop Consul/Vault/Nomad/LiteLLM (run from cat-agent-stack dir) |
cat-agent ls
cat-agent status demo/calculator
cat-agent logs demo/calculator
cat-agent rm demo/calculator --yesEach deployable folder needs:
my-agent/
├── agent.yaml # manifest (team, name, model, resources, env)
└── my_agent.py # registry() → AgentRegistry
agent.yaml requirements:
runtime.entrypoint: my_agent:registry— zero-arg factorymodel.type: api— API-backed models only (no local GGUF on Nomad)trigger.type: http— HTTP service (orperiodic/dispatchfor workers)
Local-only demos (llama.cpp, MLX) can expose registry() for cat-agent serve but omit agent.yaml.
Platform config lives in cat-agent-stack:
cat-agent-stack/cat-agent.config.toml
Deploy auto-discovers sibling ../cat-agent-stack, or set $CAT_AGENT_STACK_DIR / $CAT_AGENT_CONFIG.
Key settings:
| Setting | Purpose |
|---|---|
platform.docker_network |
Required on Mac Docker Desktop (netns) |
platform.ingress_host_template |
Traefik Host rule ({team}-{name}.localhost) |
platform.public_url_template |
Human-readable URL after deploy |
Model validation on deploy checks the live gateway model list (not a fixed allowlist). Escape hatch: --skip-alias-check.
Traefik URLs for LAN/corp: set a real DNS name in ingress_host_template — see cat-agent-stack README Shared access (LAN / corp).
| Directory | Agent | Notes |
|---|---|---|
examples/serve_fastapi/ |
calculator | Simplest HTTP agent |
examples/multi_agent/ |
earth-spin team | 3 agents, 3 models |
examples/scheduling/ |
report-scheduler | Job seed + schedule tools |
examples/tool_decorator/ |
sum tool | Minimal manifest |
Collect sources on a cadence, generate an LLM Markdown report, deliver by email or webhook.
sequenceDiagram
participant User
participant Agent as HTTP agent / CLI
participant Store as JobStore (SQLite/Postgres)
participant Worker as schedule run-due
participant Channel as SMTP / webhook
User->>Agent: create job (interval_seconds)
Agent->>Store: upsert Job row
Worker->>Store: claim due jobs
Worker->>Worker: collect sources + LLM report
Worker->>Channel: deliver Markdown
Worker->>Store: update next_run_at
from cat_agent.scheduling.models import Job
job = Job(
id='report:alice:ai-news',
user_id='alice',
kind='collect_and_report',
topic='AI news',
interval_seconds=3600, # cadence — this is what you set
channel='webhook', # smtp | resend | webhook
target='https://hooks.example/report',
enabled=True,
next_run_at=...,
)LLM tools (create_schedule, list_schedules, cancel_schedule) in cat_agent/scheduling/tools.py wrap the same store. create_schedule converts every_hours * 3600 → interval_seconds.
| Driver | When | Entry |
|---|---|---|
| APScheduler | Dev / single-node | APSchedulerDriver(store).start() |
| Kubernetes CronJob | Multi-replica | cat-agent schedule run-due |
K8s manifest: deploy/k8s/cronjob.yaml. Set CAT_AGENT_SCHEDULER_DSN to Postgres in production so all replicas share state.
pip install 'cat-agent[scheduler]'
cat-agent schedule add --user alice --topic "AI news" --every 5 \
--channel smtp --target alice@example.com
cat-agent schedule run report:alice:ai-news --dry-run
cat-agent schedule run-dueReports use delivered_at IS NULL watermarking — missed runs do not drop sources.
Turn business requirements into sandboxed, WASM-validated tools:
draft.md → interview → confirmation → ToolSpec → ToolSmith → @tool
pip install 'cat-agent[wasm,synthesis]'
cat-agent synth init vat_calculator --lang en
# edit vat_calculator_draft.md
cat-agent synth run vat_calculator_draft.mdArtifacts land in workspace/generated_tools/<name>/:
| File | Purpose |
|---|---|
<name>.py |
Assistant-ready @tool (logic inlined) |
tool.py |
Sandboxed proxy |
impl.py |
Generated implementation |
spec.json |
Compiled ToolSpec |
Load in agents:
from cat_agent.synthesis import load_generated_tools
from cat_agent.tools import enable_optional_tools
tools = load_generated_tools('vat_calculator')
enable_optional_tools(tools)Group promote/demote for production rollout: docs/synthesis-promote.md. Threat model: docs/synthesis-threat-model.md.
# .env
CAT_AGENT_OFFLINE=1
CAT_AGENT_OFFLINE_ALLOW_HOSTS=llm.internal,10.0.0.0/8- Disables network-dependent tools at registration
- Blocks outbound HTTP/sockets with
OfflineViolationError OPENAI_BASE_URL/CAT_AGENT_LLM_BASE_URLauto-added to allowlist- Self-hosted search:
CAT_AGENT_SEARXNG_URL
cat-agent offline-check --strict
cat-agent fetch-runtime --output ./wasm-runtime # transfer WASM assets offline
pip install 'cat-agent[wasm-bundled]' # runtime baked into wheelEnabled by default (CAT_AGENT_ENCRYPT_AT_REST=1). AES-GCM for:
| Data | Location |
|---|---|
| Doc-parser cache | workspace/tools/doc_parser/ |
| Parsed document cache | workspace/tools/simple_doc_parser/ |
| Agent memory | workspace/tools/storage/ |
| RAG indexes | workspace/storage/keyword_indexes/, vector_indexes/ |
| Scheduler store | CAT_AGENT_SCHEDULER_DSN path |
Key management (first match wins):
CAT_AGENT_ENCRYPTION_KEY— base64 32-byte AES key (recommended for air-gap)- OS keyring (
cat-agent/encryption-key)
cat-agent encrypt-storage --workspace ./workspace
# Strict: refuse startup if plaintext remains
CAT_AGENT_REQUIRE_ENCRYPTED_STORAGE=1Hash-chained JSONL for prompts, outputs, tool calls, and file access:
CAT_AGENT_AUDIT=1
CAT_AGENT_AUDIT_PATH=./workspace/storage/audit/audit.jsonl
cat-agent audit-verify --path ./workspace/storage/audit/audit.jsonl
cat-agent audit-export --path ... --output ./audit-export.jsonlFile paths in audit logs use SHA-256 hashes, not plaintext paths.
Offline regex redaction enabled by default at three points:
| Point | Env var | Default |
|---|---|---|
| RAG ingestion | CAT_AGENT_PII_REDACT_RAG |
on |
| Prompts to LLM | CAT_AGENT_PII_REDACT_PROMPTS |
on |
| Audit records | CAT_AGENT_PII_REDACT_AUDIT |
on |
Patterns: email, phone, IBAN, credit-card-like sequences, Turkish TC kimlik (checksum validated). Optional NER: pip install 'cat-agent[pii]'.
Build on a connected machine, transfer to regulated network:
cp deploy/.env.example deploy/.env
docker compose -f deploy/docker-compose.yml build
docker save cat-agent:on-prem | gzip > cat-agent-on-prem.tar.gz
# On air-gapped host:
docker load < cat-agent-on-prem.tar.gz
docker compose -f deploy/docker-compose.yml upSee deploy/README.md. Release SBOM: ./scripts/generate_sbom.sh sbom/.
Three libraries that sit beside agents (not RAG). Full notes and academic citations:
docs/tracing.md, docs/context.md,
docs/failure-analysis.md, docs/references.md.
Machine-readable run history for cost, debugging, and evaluation — separate from Loguru logging.
- JSONL append-only store, schema v1.0 (
Run/Step) - Step kinds:
llm_call,tool_call,handoff,context_op,error, … RunLimits— stop cleanly on max steps / tokens / wall clock / tool calls- Redaction — API keys never land in persisted
llm_config - Token & cost totals —
RunTotals(backend usage when present, otherwise estimated)
export CAT_AGENT_TRACE=1
export CAT_AGENT_TRACE_FILE=./traces.jsonl # optional; else in-memoryOr agent.run(..., trace=True, trace_store=JSONLTraceStore(...), run_limits=RunLimits(...)).
Example: examples/trace/.
Manages the growing conversation window during a long run. Separate from
cat_agent.memory (RAG over user files).
| Strategy | Role |
|---|---|
| Observation masking (default) | Elide old tool bodies; keep a compact residue (ids, repeated status tokens, salient mid-lines) |
| Summary compaction | Optional LLM summary of oldest blocks |
fold() |
Explicit scratch sub-task → one result message |
from cat_agent.context import ContextManager, ObservationMaskingStrategy
bot = Assistant(llm={...}, context_manager=ContextManager(
strategies=[ObservationMaskingStrategy(keep_recent=3)],
))
# Disable: context_manager=False or CAT_AGENT_CONTEXT=0Examples: examples/context/, examples/long_horizon_agent/.
Classify failure modes on a recorded trace using the MAST taxonomy (Cemri et al., arXiv:2503.13657) — 14 modes across system design, inter-agent misalignment, and task verification.
- Tier-1 (no LLM): deterministic detectors for 1.3 step repetition, 1.4 loss of history, 1.5 unaware of termination
- Tier-2 (opt-in): LLM-as-judge for the remaining modes — never sent without
an explicit
judge_llm=
python -m cat_agent.analysis traces.jsonl
python -m cat_agent.analysis traces.jsonl --judge # needs a configured LLMExamples: examples/analysis/, examples/failure_analysis/.
Silent by default (library-friendly). Activate with:
CAT_AGENT_LOG_LEVEL=INFO python my_script.py
CAT_AGENT_LOG_FORMAT=json python my_script.py # structured
CAT_AGENT_LOG_FILE=agent.log python my_script.py # rotating fileOr programmatically: from cat_agent.log import setup_logger.
Safe Python execution:
| Backend | Requires |
|---|---|
| WASM | [wasm] or [wasm-bundled] — no Docker |
| Docker | [code_interpreter] + Docker daemon |
Example: examples/wasm_code_interpreter/.
Expose agents as MCP servers: examples/mcp_service/.
Per-tool retries, timeouts, and rate limiting: examples/tool_resilience/.
Run from repo root after installing matching extras.
| Path | Topic |
|---|---|
trace/ |
Structured JSONL traces (RunTotals, RunLimits) |
context/ |
Observation masking + fold + residue |
analysis/ |
MAST Tier-1 (+ optional Ollama judge) |
long_horizon_agent/ |
Context masking quality A/B + traced prompt tokens |
failure_analysis/ |
MAST Tier-1 analysis on a failing loop |
tool_decorator/ |
@tool decorator + deploy yaml |
serve_fastapi/ |
HTTP serve + Nomad deploy |
multi_agent/ |
GroupChat, Router, 3-model team |
scheduling/ |
Report jobs (local + deploy) |
graph/ |
DAG workflow (StateGraph) |
async_agent/ |
arun + parallel tools |
native_parallel/ |
Live multi tool_calls against an OAI endpoint |
synthesis/from_draft/ |
Markdown → sandboxed tool |
synthesis/from_spec/ |
JSON ToolSpec → ToolSmith |
synthesis/promote/ |
Offline promote / share / adopt |
rag_keyword/ |
Rust BM25 keyword search |
rag_vector/ |
Native HNSW vector search |
rag_native/ |
Chunking + vector + truncation |
long_term_memory/ |
Cross-session memory |
doc_parser_agent/ |
Document Q&A |
parallel_doc_qa/ |
Exhaustive per-chunk ParallelDocQA |
observability/ |
Trace handlers |
langfuse/ |
OpenTelemetry → Langfuse UI |
llama_cpp_math_guy/ |
Local GGUF + tools |
llama_cpp_vision/ |
Multimodal LlamaCpp |
transformers_math_guy/ |
HuggingFace local |
mlx_lm_math_guy/ |
Apple Silicon MLX |
wasm_code_interpreter/ |
WASM sandbox |
mcp_service/ |
MCP server |
tool_resilience/ |
Retry, timeout, rate limit |
logging_demo/ |
Loguru configuration |
| Package | Description |
|---|---|
cat_agent.agent |
Base Agent (run / arun) |
cat_agent.agents |
Assistant, ReActChat, FnCallAgent, DocQA, GroupChat, Router |
cat_agent.multi_agent |
Blackboard, handoff, team tools |
cat_agent.graph |
StateGraph / GraphAgent DAG engine |
cat_agent.llm |
Chat backends (OAI, LlamaCpp, Transformers, MLX) |
cat_agent.tools |
@tool, RAG, DocParser, Storage, MCP, code interpreter |
cat_agent.memory |
Long-term encrypted memory |
cat_agent.scheduling |
Report jobs, channels, runner |
cat_agent.serve |
FastAPI HTTP invoke |
cat_agent.synthesis |
ToolSpec → ToolSmith → WASM validation |
cat_agent.platform |
Nomad deploy, manifest, HCL render |
cat_agent.security |
Offline guards, PII, encryption, audit |
cat_agent.observability |
Event hooks (Mermaid, OTel, Langfuse) |
cat_agent.trace |
Structured JSONL runs (Run / Step, RunLimits, redaction, totals) |
cat_agent.context |
Conversation-window masking / compaction / fold (not RAG) |
cat_agent.analysis |
MAST failure taxonomy over traces (Tier-1 + optional judge) |
cat_agent._native |
Rust: BM25, HNSW, PDF, tokenizer |
native/ |
Rust source (maturin/PyO3) |
examples/ |
Runnable demos (see index above) |
deploy/ |
Air-gap Docker + k8s CronJob |
benchmarks/ |
RAG / vector / truncation micro-benchmarks |
tests/ |
1600+ pytest functions |
pip install -e ".[test,local,rag,otel,synthesis]"
pytest
pytest --cov=./cat_agent --cov-report=term
cargo test --manifest-path native/Cargo.toml --no-default-featuresBenchmarks:
python benchmarks/benchmark_rag.py --chunks 1000 --queries 25
python benchmarks/benchmark_native_vector.py --chunks 2000 --queries 25
python benchmarks/benchmark_pdf_parser.py --pages 10 --repeats 3Wheels built for abi3 Python 3.10+ on Linux (x86_64, aarch64), macOS arm64, Windows amd64:
chmod +x release.sh
./release.sh 0.10.1Licensed under the Apache License 2.0.
Kemalcan Bora — kemalcanbora@gmail.com
GitHub: kemalcanbora/cat-agent