AI Hardening Sandbox is a hands-on local lab for teaching LLM prompt-injection defenses through Red Team and Blue Team exercises. It combines Docker Desktop, Ollama, and a Python-based security gateway so students can compare vulnerable, hardened, and filtered model behavior in a controlled environment.
This repository is built around a single core learning objective: understand how layered defenses change the outcome of prompt injection and sensitive-data leakage attempts.
Students move through three protection layers:
- Phase 1: Model hardening
- strengthen the assistant's system prompt
- test the model without a gateway
- Phase 2: Static gateway filtering
- apply ingress and egress checks in
lab/scripts/filter_rules.py
- apply ingress and egress checks in
- Phase 3: OPA policy enforcement
- classify request context and evaluate policy logic in
policies/rules.jsonandpolicies/gateway.rego
- classify request context and evaluate policy logic in
The lab uses a fictional credit-union support assistant called Piper, with both a vulnerable and a hardened variant. The same gateway can route requests through either model and compare outcomes directly.
The lab runs as a three-service Docker stack defined in docker-compose.yml:
llmruns Ollama on port11434and stores model data in the persistentollama_storagevolume.webruns the Flask gateway on port5000, installs Python dependencies on startup, and communicates with Ollama viaOLLAMA_HOST=http://llm:11434.oparuns Open Policy Agent on port8181for policy-based context controls and decision checks.
flowchart LR
Browser["Browser<br/>localhost:5000"]
subgraph Compose["docker compose up -d"]
direction LR
subgraph web["web container — Flask gateway"]
GW["secure_gateway.py"]
FR["filter_rules.py<br/>(Phase 2 rules)"]
end
subgraph llm["llm container — Ollama :11434"]
VB["vulnerable_bot"]
HB["hardened_bot"]
CTX["llama3.2<br/>(context classifier)"]
end
subgraph opa["opa container — OPA :8181"]
REGO["gateway.rego<br/>(decision logic)"]
RULES["rules.json<br/>(Phase 3 policy data)"]
end
end
Browser -->|"HTTP requests"| GW
GW -->|"Phase 1 / 2 / 3"| VB
GW -->|"Phase 1 / 2 / 3"| HB
GW -.->|"Phase 3: classify prompt"| CTX
GW -->|"Phase 3: policy decision"| REGO
REGO --> RULES
GW --> FR
The gateway logic lives in lab/scripts/secure_gateway.py. In the normal lab flow, it is started by the Docker stack and does not need to be launched manually.
For the fastest path to the lab, start the Docker stack from the repo root:
docker compose up -dThen open the application in a browser:
http://localhost:5000
If you need to create or refresh the Ollama models explicitly:
docker compose exec llm ollama pull llama3.2
docker compose exec llm ollama create vulnerable_bot -f /app/lab/modelfiles/vulnerable.txt
docker compose exec llm ollama create hardened_bot -f /app/lab/modelfiles/hardened.txtFor full setup instructions, model build steps, and environment troubleshooting, use lab/SETUP_GUIDE.md. This README is intentionally focused on the architecture and learning flow.
Use the three phases in sequence for Red Team and Blue Team practice:
- Edit
lab/modelfiles/hardened.txtto strengthen the assistant's system instructions. - Rebuild the model with:
docker compose exec llm ollama create hardened_bot -f /app/lab/modelfiles/hardened.txt- Adjust
lab/scripts/filter_rules.pyto tuneINGRESS_BLACKLIST,EGRESS_SECRETS, andEGRESS_PATTERNS. - Refresh the browser; the gateway hot-reloads the file on each request.
- Use the UI mode Phase 3 - OPA context policy.
- Tune
policies/rules.jsonandpolicies/gateway.regoto control policy thresholds, contexts, and block reasons.
flowchart TD
Start(["Student submits a prompt"]) --> Mode{"Protection mode?"}
Mode -->|"Phase 1: Direct"| Model1["Send straight to the model"]
Model1 --> Resp1["Response shown as-is"]
Mode -->|"Phase 2: Static filters"| Ing2{"filter_rules.py<br/>INGRESS_BLACKLIST match?"}
Ing2 -->|Yes| Block2["Blocked before reaching the model"]
Ing2 -->|No| Model2["Send to the model"]
Model2 --> Eg2{"EGRESS_SECRETS / EGRESS_PATTERNS match?"}
Eg2 -->|Yes| BlockE2["Blocked before display"]
Eg2 -->|No| Resp2["Response shown"]
Mode -->|"Phase 3: OPA context policy"| Classify["llama3.2 classifies intent and risk"]
Classify --> OPAIn["OPA decides whether ingress is allowed"]
OPAIn -->|block| Block3["Blocked before reaching the model"]
OPAIn -->|allow| Model3["Send to the model"]
Model3 --> OPAOut["OPA decides whether egress is allowed"]
OPAOut -->|block| BlockE3["Blocked before display"]
OPAOut -->|allow| Resp3["Response shown"]
The sandbox uses a single scenario: Piper, a customer-support assistant for a fictional credit union. The environment contains:
vulnerable_bot— intentionally unsafe baselinehardened_bot— system-prompt defense baseline- a shared gateway that can route traffic through either model
This arrangement creates a clear comparison point: students can test attacks against the same scenario while toggling one defense layer at a time.
Use the tier switcher to set Phase 2 difficulty without manually editing files:
python lab/scripts/set_tier.py <tier>Available tiers:
calibrated(alias:demo) — default teaching baselinescaffolded(alias:student) — intermediate starter set with partial controlsblank(alias:advanced) — empty rule list for advanced scenarios
This script copies a preset into lab/scripts/filter_rules.py and the gateway hot-reloads it on each request.
| Path | Purpose |
|---|---|
docker-compose.yml |
Starts the llm, web, and opa services |
lab/modelfiles/vulnerable.txt |
Unhardened baseline model |
lab/modelfiles/hardened.txt |
Hardened model for Phase 1 testing |
lab/scripts/secure_gateway.py |
Browser gateway and request orchestration |
lab/scripts/filter_rules.py |
Phase 2 static filtering |
lab/scripts/set_tier.py |
Preset rule selection |
policies/rules.json |
Phase 3 policy data |
policies/gateway.rego |
Phase 3 OPA logic |
lab/SETUP_GUIDE.md |
Step-by-step installation and troubleshooting |
docs/assignments/ |
Red Team and Blue Team materials |
The lab is designed for guided walkthroughs, team exercises, and direct comparison of defense layers.
Included classroom materials:
This lab maps closely to OWASP GenAI/LLM Top 10 concerns, especially:
- LLM01 Prompt Injection
- LLM02 Sensitive Information Disclosure
- LLM08 Hidden Context Exposure
This README intentionally avoids long operational troubleshooting lists. For installation and environment issues, use lab/SETUP_GUIDE.md.
The main operational checks are:
- confirm Docker Desktop is running
- confirm
docker compose up -dshows the expected containers - confirm the models exist with
docker compose exec llm ollama list - check
docker compose logsforllm,web, oropaif the stack is unhealthy
Stop the stack and remove the persistent data when you want a fresh environment:
docker compose down -vIf you only need to reset the Ollama model state, remove the model entries before removing the storage volume.
The default scenario is intentionally tuned so Blue Team students have a meaningful challenge. If you want to reskin the lab for a different industry or fictional organization, the main files to adjust are:
lab/modelfiles/vulnerable.txtandlab/modelfiles/hardened.txt— persona, secrets, and system boundarieslab/scripts/filter_rules.py— static ingress/egress rulespolicies/rules.json— OPA context/risk datapolicies/gateway.rego— decision logic
Keep at least one meaningful bypass gap in the default Phase 2 rules so students still have a real defense task to complete.
This laboratory supports the AI Security and pedagogy curriculum published on Code and Cypher:
- Pedagogy Overview: My Teaching Philosophy: Building Accessible AI Security Education
- Hands-on Lab Series: Hands-on AI Security: Zero-Cost Local LLM Defense Lab