Deterministic Security Reference Monitor, Dynamic Taint Tracking & Adaptive Red-Teaming for Autonomous Tool-Calling AI Agents.
Proving that autonomous agent security cannot be entrusted to stochastic in-model prompts, but must be enforced by a deterministic reference monitor outside the model.
Autonomous tool-calling agents operate in untrusted environments where model inputs inevitably mix instructions with untrusted external data (invoices, emails, web search results). INTERPOSE addresses the fundamental flaw of in-model prompt guardrails: stochastic models cannot self-police.
By treating agent security as a formal Information Flow Control (IFC) problem, INTERPOSE enforces deterministic invariants outside the LLM execution loop:
- Provenance Tracking: Dynamic taint tracking with join semi-lattice guarantees.
- Deterministic AST Validation: Strict syntactic AST-level arguments verification before execution.
- Cryptographic HITL Sign-Off: HMAC-SHA256 tokens gating irreversible actions (wire transfers, database mutations).
┌────────────────────────────────────────────────────────────────────────┐
│ UNTRUSTED DATA SOURCES │
│ (Vendor Invoices, Web Scrapes, Support Tickets, PDF Files) │
└───────────────────────────────────┬────────────────────────────────────┘
│ Ingested & Tagged with Provenance
▼
┌────────────────────────────────────────────────────────────────────────┐
│ AGENT REASONING LOOP │
│ (Proposes Action: { tool_name, arguments }) │
└───────────────────────────────────┬────────────────────────────────────┘
│ Out-of-Band Interception
════════════════════════════════════╪═════════════════════════════════════
INTERPOSE DETERMINISTIC REFERENCE MONITOR (Zero Trust)
════════════════════════════════════╪═════════════════════════════════════
1. DYNAMIC TAINT LATTICE (Information Flow Control):
Ordering: SANITIZED < USER_TRUSTED < TOOL_UNTRUSTED_WEB < CRITICAL_SECRET
Lattice Join: lub(Trusted, Untrusted) = Untrusted
2. DECLARATIVE CAPABILITY POLICIES:
Rule A: TOOL_UNTRUSTED_WEB cannot flow to sensitive parameters
(e.g., send_email.recipient, execute_bash.cmd, sql_query.query).
Rule B: CRITICAL_SECRET cannot reach external network egress sinks.
3. AST & ARGUMENT SANITIZER:
Deterministic SQL AST check (SELECT-only; stacked queries & DROP blocked).
Shell AST lexical check (chaining operators ;, &&, ||, | strictly blocked).
4. CRYPTOGRAPHIC HITL GATE:
Irreversible actions (wire_transfer, delete_database_records) halt
execution and require signed HMAC-SHA256 challenge verification.
5. IMMUTABLE FORENSIC AUDIT ENGINE:
Structured JSONL logging of all provenance tags, rules, and latencies.
════════════════════════════════════╪═════════════════════════════════════
│
[ Policy Engine Verdict ]
┌────────────┴────────────┐
[PERMIT] [DENY]
│ │
▼ ▼
┌─────────────────────────────┐ ┌───────────────────────────────┐
│ Sandboxed Tool Sink │ │ Security Policy Exception │
│ (Tool executes safely) │ │ Action blocked; structured │
│ Safe output returned │ │ remediation sent to agent │
└─────────────────────────────┘ └───────────────────────────────┘
Evaluated on local NVIDIA GeForce RTX 4060 Laptop GPU against AgentDojo attack scenarios and PyRIT adversarial mutations across Qwen2.5-7B and Llama-3.2-3B:
| Defense Condition | Attacks Evaluated | Blocked | Attack Success Rate (ASR) | Benign Utility (%) | Median Latency (ms) |
|---|---|---|---|---|---|
| 1. No Defense | 36 | 0 | 100.0% | 100.0% | 0.000 ms |
| 2. Prompt Guardrail | 36 | 24 | 33.3% | 66.7% | 0.000 ms |
| 3. Argument Regex Only | 36 | 12 | 66.7% | 100.0% | 0.000 ms |
| 4. Full INTERPOSE | 36 | 36 | 0.0% | 100.0% | 0.038 ms |
The Next.js 14 frontend (frontend/) is a bespoke Cybersecurity Operations & Threat Simulation Console:
- Interactive Taint Lineage DAG Canvas: Real-time visual node flow from untrusted ingestion to policy enforcement, complete with forensic provenance inspection.
- Exploit Replay & Attack Playground: Side-by-side comparison showing an unprotected agent (Critical Breach) versus INTERPOSE (Attack Neutralized in <0.04 ms).
- HITL Authorization Queue: Interactive drawer managing irreversible tool invocations with cryptographic HMAC-SHA256 token verification.
- Pareto & Radar Charts: Recharts visualization demonstrating 0.0% ASR alongside 100% utility retention.
- Zero-Auth Simulator Mode: Pre-loaded with canonical AgentDojo scenarios for instant evaluation by recruiters with zero login walls.
- Python 3.11+
- Node.js 18+ & npm
uv(recommended)
# Clone and enter repository
git clone https://github.com/Hamza-HATTAB/interpose.git
cd interpose
# Install Python and Node dependencies
make install# Run unit and integration tests (37 tests)
make test
# Execute the 4-condition security benchmark matrix
make benchmark
# Execute full end-to-end regression suite
make verify# Terminal 1: Launch FastAPI reference monitor
make serve
# Terminal 2: Start Next.js 14 Sentinel HUD
make frontend-devOpen http://localhost:3002 in your browser.
interpose/
├── pyproject.toml # Python project configuration (uv managed)
├── Makefile # Standardized development & testing targets
├── README.md # Project overview & architecture guide
├── TECHNICAL_REPORT.md # Whitepaper, formal proofs & interview guide
├── LICENSE # MIT License
├── interpose/
│ ├── core/
│ │ ├── lattice.py # TaintLevel enum & ProvenanceTag lattice join
│ │ └── taint.py # TaintedStr subclass & dynamic propagation
│ ├── policy/
│ │ ├── models.py # SinkDefinition, PolicyVerdict, Decision
│ │ ├── validator.py # AST argument checks (SQL, Shell, Paths)
│ │ └── engine.py # Deterministic PolicyEngine (<0.04ms)
│ ├── gateway/
│ │ ├── hitl.py # Cryptographic HMAC-SHA256 approval gate
│ │ ├── sandbox.py # Isolated SQLite, bash, and email sandbox
│ │ ├── audit.py # Immutable JSONL forensic logger
│ │ └── proxy.py # InterposeGatewayProxy complete mediation
│ ├── redteam/
│ │ ├── adversary.py # PyRIT mutations (Base64, Delimiters, Homoglyphs)
│ │ ├── scenarios.py # AgentDojo benchmark scenario loader
│ │ ├── agent.py # Autonomous agent loop runner
│ │ └── benchmark.py # 4-condition evaluation harness
│ └── server/
│ └── app.py # FastAPI endpoints for live GPU & HUD
├── frontend/ # Next.js 14.2.35 Sentinel Cyber-Defense HUD
│ ├── package.json
│ ├── tsconfig.json
│ ├── components/ # TaintCanvas, AttackPlayground, HITLQueue, Radar
│ └── lib/
│ ├── types.ts
│ ├── mockData.ts # Pre-loaded zero-auth recruiter dataset
│ └── api.ts
├── tests/ # 37 unit, integration, and security regression tests
├── data/ # Scenarios, benchmark results, and audit logs
└── scripts/ # Benchmark, tunnel, and verification scripts
- Author: Hamza Riadh Hattab
- Email: hamza.riadh.htb@gmail.com
- GitHub: github.com/Hamza-HATTAB
- LinkedIn: linkedin.com/in/hamza-riadh-h-44a297345