September 2026 re-run: current dependencies, test suite, measured cost - #1
Merged
Merged
Conversation
The workshop was delivered at COMPDES in July 2026. On 30 Sep 2026 I cloned the repo into a clean folder and ran it again from zero, to bring it up to date with today's versions of everything it depends on. First update: `pip install -r requirements.txt` now resolves mcp 2.2.0. The 2.0 release (28 Jul 2026, after the workshop) renamed FastMCP to MCPServer and moved host/port off the settings object, so servers written against 1.x no longer import. The 1.x line is still maintained, so pin `mcp>=1.8,<2`. Re-run with mcp 1.30.0 on Python 3.14: the three Hour 1 demo questions work end to end against gemini-3.5-flash-lite.
…sport
Updates from re-running the build phase on 30 Sep 2026:
- The description of consultar_inventario now states the table and column
names. With the current model the first demo question ("¿Cuánto stock
tenemos de cemento?") resolves in exactly one tool call, 8 runs out of 8.
- agent.py can log the tokens of every model call (set USAGE_LOG to a file),
so the cost of a session can be measured.
- The stdio/HTTP startup block moves to target/mcp/transporte.py, shared by
both MCP servers, so either one can be served to Onyx with --http.
- http_wrapper.py accepts an optional `historial`, so a multi-turn attack
(Lab 2.3) can be driven over HTTP, and AGENTE=seguro serves the hardened
agent through the same endpoint.
- The repo root is on sys.path, so the `from defenses.x import ...` snippets
of Hour 3 can be pasted into agent.py as the guide describes.
…wer key Updates from re-running the hardening phase on 30 Sep 2026 against today's libraries and model: - router.py: current semantic-router releases no longer export RouteLayer. The router is now ~15 lines on the local sentence-transformers model the RAG already uses: no extra dependency, no API spend. Threshold calibrated at 0.65 against the workshop's own prompts. - hitl.py covers actualizar_stock and any non-SELECT statement sent through consultar_inventario (turn 5 of the Crescendo lab), and denies when there is no human at the terminal. - inventory_mcp_server_seguro.py adds buscar_producto and a typed actualizar_stock, so the three Hour 1 business questions work after hardening, plus --http for the Onyx route. Writes connect as escritor_stock, reads as lector. - roles_seguros.sql is idempotent, and lector does not read costo_unit (the margin). - pin_descriptors.py ships the approved hash of the hardened server. - agent_seguro.py (new): the Hour 1 agent with all six controls applied, each block tagged [3.x]. It is the answer key for the hour and what the tests attack. - promptfooconfig.yaml targets the local wrapper at http://127.0.0.1:8000. Re-run with promptfoo 0.123.1: 4/4 pass against the hardened agent (exit 0), 3/4 fail against the vulnerable one (exit 100).
Each attack was re-run on 30 Sep 2026 against the vulnerable agent: - Lab 2.1 gets its own section in the README. The hidden query runs in about 8 of 10 attempts; the README says that attacks are probabilistic. - Lab 2.4 publishes http-echo's listening port (5678). The agent returns SECRETO-INTERNO-12345. - Lab 2.5: adds 2_5_garak_config.json and a prompt cap for garak 0.17.0. Measured: 60 prompts, 2.5 minutes, $0.014. The uncapped promptinject family is 768 prompts in this release. - OWASP ASI codes completed (ASI02, ASI04), and the PyRIT note follows its current API: the class is CrescendoAttack and it needs its own attacker model.
The Postgres container publishes its port on 127.0.0.1 only: its password is public (it is in this repo) and it has no reason to be reachable from the room's network. Both MCP servers and .env.example now default POSTGRES_HOST to 127.0.0.1. On Windows, `localhost` tries IPv6 first and adds about 2 s to every connection, which is every tool call: measured 2.04 s against 0.03 s.
…mand
tests/test_taller.py walks the whole workshop in three levels, each skipped
when its prerequisite is missing:
1. offline - PDFs, RAG, both MCP servers over stdio and HTTP, router,
HITL, descriptor pinning, every relative link in the docs
2. database - the vulnerable tools leak, destroy and SSRF; the roles and
the hardened server prevent it
3. live - Hour 1's three questions, then each Hour 2 attack against
the vulnerable agent and the same attack against the
hardened one (TALLER_LIVE=1, about 1.2 cents)
Attacks are retried up to three times because an LLM decides; defenses get
one try. tests/costo.py turns the usage log into dollars per phase.
Result on Windows 11, Python 3.14, mcp 1.30, gemini-3.5-flash-lite:
24 tests, six consecutive green passes, about 2.5 minutes and 39 model
calls each. Across eleven passes no defense test failed; in an earlier batch
of five, the Lab 2.2 attack did not land within its three tries twice.
Measured on its own it lands 29 times out of 30.
Onyx has moved since July, so docs/ONYX.md was walked again click by click on a fresh v4.8.2 deployment and rewritten to match: - The model is added under Language Models -> Custom Models with provider `gemini`; the dedicated Google Gemini tile is gone. - Onyx now ships SSRF protection that validates every outbound request, so a local MCP server is refused until Security & Hardening -> Network Safety -> SSRF Protection is set to Allow Private Network. New step 3b. - MCP servers are registered under MCP Actions, and tools are cached: switching to the hardened server needs "Refresh tools". - The deployment needs USER_AUTH_SECRET; onyx.sh and onyx.ps1 generate it. - OpenSearch replaced Vespa. Measured footprint: 7.7 GiB of RAM, 21 GB of images. Walked end to end: deploy, model, MCP server (3 tools), agent, the cement question (1,200), the stock update, the SSRF lab, and the switch to the hardened server over --http (SSRF rejected). Uploading the PDFs for RAG inside Onyx was not part of this pass; docs/VERIFICACION.md says so.
- docs/VERIFICACION.md (new): the record of the 30 Sep 2026 re-run. What was run, the environment and resolved versions, what had changed since July, what was added, and what was not covered. - docs/PRESUPUESTO.md: measured cost. One full pass of Hours 1-3 is 39 model calls and $0.012; Promptfoo $0.004; capped Garak $0.014. Per student: about 3 cents guided, 10-15 cents with repeats, about 30 cents with uncapped Garak, against a budget of $1.10. - Facts re-checked on primary sources: gemini-3.5-flash-lite is still $0.30 / $2.50 per 1M tokens (output includes thinking tokens); AI Studio minimum purchase is now $5; depleted credit is documented as HTTP 402, which check_key.py now recognises alongside 429; Promptfoo needs Node 22. - README: status section, links to the Hour 2 and Hour 3 guides, the real clone URL. SETUP: how to run the tests, and the separate environments for Garak and Promptfoo.
Installs the dependencies from scratch, at whatever versions resolve that day, starts the lab database and runs the test levels that need no API key. A dependency release that changes the workshop's behaviour, as mcp 2.0 did, shows up here on its own.
The first CI run (Python 3.12 on Linux) computed a different descriptor hash than Python 3.14: since 3.13 the interpreter strips the indentation of docstrings, so a multi-line tool description differs by whitespace only. hash_tools now collapses whitespace in the description before hashing, and APROBADO is the hash of the hardened server under that rule. A changed word, tool name or schema still changes the hash. The workflow runs on pull requests, on pushes to main and every Monday.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The workshop was delivered at COMPDES in July 2026. On 30 Sep 2026 I cloned it into a clean folder and ran it again from zero, to bring it up to date with today's versions of everything it depends on and to leave it easy to follow for anyone who finds the repo.
What had moved since July
FastMCPmcp>=1.8,<2(re-run with 1.30.0)semantic-routerno longer exportsRouteLayersentence-transformers, already a dependencydocs/ONYX.mdre-walked click by click and rewrittencheck_key.pyupdatedWhat this adds
tests/test_taller.py: 24 tests in three levels (offline, with the database, live against the model). Each Hour 2 attack is run against the vulnerable agent, then the same attack against the hardened one.defenses/agent_seguro.py: the Hour 1 agent with all six Hour 3 controls applied, as the answer key.tests/costo.py,docs/PRESUPUESTO.md): one full pass of Hours 1–3 is 39 model calls and $0.012. Per student: about 3 cents guided, 10–15 cents with repeats, about 30 cents with uncapped Garak.docs/VERIFICACION.md: the full record, including what this pass did not cover.Results
promptinjectcapped at 60 prompts: attack success 50 % / 10 % / 65 % per probe.Attack tests depend on an LLM and are retried up to three times; defense tests get one try and did not fail in any pass.
Not covered in this pass
RAG upload inside Onyx, Docker Desktop on Windows/macOS, the install scripts end to end on native Linux and macOS, PyRIT, and uncapped Garak (its cost is extrapolated). Details in
docs/VERIFICACION.md.To repeat it