Skip to content

September 2026 re-run: current dependencies, test suite, measured cost - #1

Merged
DavidMGDev merged 10 commits into
mainfrom
refresh/sept-2026-rerun
Oct 1, 2026
Merged

DavidMGDev merged 10 commits into
mainfrom
refresh/sept-2026-rerun

Conversation

@DavidMGDev

Copy link
Copy Markdown
Owner

The workshop was delivered at COMPDES in July 2026. On 30 Sep 2026 I cloned it into a clean folder and ran it again from zero, to bring it up to date with today's versions of everything it depends on and to leave it easy to follow for anyone who finds the repo.

What had moved since July

Change upstream Update here
MCP Python SDK 2.0 (28 Jul) renamed FastMCP mcp>=1.8,<2 (re-run with 1.30.0)
semantic-router no longer exports RouteLayer Router rebuilt on sentence-transformers, already a dependency
Onyx v4.8.2: new menus, SSRF protection on by default, cached MCP tools, OpenSearch docs/ONYX.md re-walked click by click and rewritten
Promptfoo needs Node 22; AI Studio minimum purchase and error codes Docs and check_key.py updated

What this adds

  • tests/test_taller.py: 24 tests in three levels (offline, with the database, live against the model). Each Hour 2 attack is run against the vulnerable agent, then the same attack against the hardened one.
  • defenses/agent_seguro.py: the Hour 1 agent with all six Hour 3 controls applied, as the answer key.
  • Measured cost (tests/costo.py, docs/PRESUPUESTO.md): one full pass of Hours 1–3 is 39 model calls and $0.012. Per student: about 3 cents guided, 10–15 cents with repeats, about 30 cents with uncapped Garak.
  • CI: the key-free test levels run on every push and every Monday against that day's dependency versions.
  • docs/VERIFICACION.md: the full record, including what this pass did not cover.

Results

  • Test suite: 24 of 24, six consecutive full passes (about 2.5 minutes each).
  • Promptfoo 0.123.1: 4/4 against the hardened agent, 3/4 fail against the vulnerable one.
  • Garak 0.17.0, promptinject capped at 60 prompts: attack success 50 % / 10 % / 65 % per probe.
  • Route A on Onyx v4.8.2: deploy, model, MCP server, agent, demo, SSRF lab and the switch to the hardened server.

Attack tests depend on an LLM and are retried up to three times; defense tests get one try and did not fail in any pass.

Not covered in this pass

RAG upload inside Onyx, Docker Desktop on Windows/macOS, the install scripts end to end on native Linux and macOS, PyRIT, and uncapped Garak (its cost is extrapolated). Details in docs/VERIFICACION.md.

To repeat it

python -m unittest discover tests                 # no API key needed
TALLER_LIVE=1 python -m unittest discover tests   # + live, about 1.2 cents
python tests/costo.py

The workshop was delivered at COMPDES in July 2026. On 30 Sep 2026 I cloned
the repo into a clean folder and ran it again from zero, to bring it up to
date with today's versions of everything it depends on.

First update: `pip install -r requirements.txt` now resolves mcp 2.2.0. The
2.0 release (28 Jul 2026, after the workshop) renamed FastMCP to MCPServer
and moved host/port off the settings object, so servers written against 1.x
no longer import.

The 1.x line is still maintained, so pin `mcp>=1.8,<2`. Re-run with
mcp 1.30.0 on Python 3.14: the three Hour 1 demo questions work end to end
against gemini-3.5-flash-lite.
…sport

Updates from re-running the build phase on 30 Sep 2026:

- The description of consultar_inventario now states the table and column
  names. With the current model the first demo question ("¿Cuánto stock
  tenemos de cemento?") resolves in exactly one tool call, 8 runs out of 8.
- agent.py can log the tokens of every model call (set USAGE_LOG to a file),
  so the cost of a session can be measured.
- The stdio/HTTP startup block moves to target/mcp/transporte.py, shared by
  both MCP servers, so either one can be served to Onyx with --http.
- http_wrapper.py accepts an optional `historial`, so a multi-turn attack
  (Lab 2.3) can be driven over HTTP, and AGENTE=seguro serves the hardened
  agent through the same endpoint.
- The repo root is on sys.path, so the `from defenses.x import ...` snippets
  of Hour 3 can be pasted into agent.py as the guide describes.
…wer key

Updates from re-running the hardening phase on 30 Sep 2026 against today's
libraries and model:

- router.py: current semantic-router releases no longer export RouteLayer.
  The router is now ~15 lines on the local sentence-transformers model the
  RAG already uses: no extra dependency, no API spend. Threshold calibrated
  at 0.65 against the workshop's own prompts.
- hitl.py covers actualizar_stock and any non-SELECT statement sent through
  consultar_inventario (turn 5 of the Crescendo lab), and denies when there
  is no human at the terminal.
- inventory_mcp_server_seguro.py adds buscar_producto and a typed
  actualizar_stock, so the three Hour 1 business questions work after
  hardening, plus --http for the Onyx route. Writes connect as
  escritor_stock, reads as lector.
- roles_seguros.sql is idempotent, and lector does not read costo_unit
  (the margin).
- pin_descriptors.py ships the approved hash of the hardened server.
- agent_seguro.py (new): the Hour 1 agent with all six controls applied,
  each block tagged [3.x]. It is the answer key for the hour and what the
  tests attack.
- promptfooconfig.yaml targets the local wrapper at http://127.0.0.1:8000.
  Re-run with promptfoo 0.123.1: 4/4 pass against the hardened agent
  (exit 0), 3/4 fail against the vulnerable one (exit 100).
Each attack was re-run on 30 Sep 2026 against the vulnerable agent:

- Lab 2.1 gets its own section in the README. The hidden query runs in
  about 8 of 10 attempts; the README says that attacks are probabilistic.
- Lab 2.4 publishes http-echo's listening port (5678). The agent returns
  SECRETO-INTERNO-12345.
- Lab 2.5: adds 2_5_garak_config.json and a prompt cap for garak 0.17.0.
  Measured: 60 prompts, 2.5 minutes, $0.014. The uncapped promptinject
  family is 768 prompts in this release.
- OWASP ASI codes completed (ASI02, ASI04), and the PyRIT note follows its
  current API: the class is CrescendoAttack and it needs its own attacker
  model.
The Postgres container publishes its port on 127.0.0.1 only: its password
is public (it is in this repo) and it has no reason to be reachable from
the room's network.

Both MCP servers and .env.example now default POSTGRES_HOST to 127.0.0.1.
On Windows, `localhost` tries IPv6 first and adds about 2 s to every
connection, which is every tool call: measured 2.04 s against 0.03 s.
…mand

tests/test_taller.py walks the whole workshop in three levels, each skipped
when its prerequisite is missing:

  1. offline  - PDFs, RAG, both MCP servers over stdio and HTTP, router,
                HITL, descriptor pinning, every relative link in the docs
  2. database - the vulnerable tools leak, destroy and SSRF; the roles and
                the hardened server prevent it
  3. live     - Hour 1's three questions, then each Hour 2 attack against
                the vulnerable agent and the same attack against the
                hardened one (TALLER_LIVE=1, about 1.2 cents)

Attacks are retried up to three times because an LLM decides; defenses get
one try. tests/costo.py turns the usage log into dollars per phase.

Result on Windows 11, Python 3.14, mcp 1.30, gemini-3.5-flash-lite:
24 tests, six consecutive green passes, about 2.5 minutes and 39 model
calls each. Across eleven passes no defense test failed; in an earlier batch
of five, the Lab 2.2 attack did not land within its three tries twice.
Measured on its own it lands 29 times out of 30.
Onyx has moved since July, so docs/ONYX.md was walked again click by click
on a fresh v4.8.2 deployment and rewritten to match:

- The model is added under Language Models -> Custom Models with provider
  `gemini`; the dedicated Google Gemini tile is gone.
- Onyx now ships SSRF protection that validates every outbound request, so
  a local MCP server is refused until Security & Hardening -> Network
  Safety -> SSRF Protection is set to Allow Private Network. New step 3b.
- MCP servers are registered under MCP Actions, and tools are cached:
  switching to the hardened server needs "Refresh tools".
- The deployment needs USER_AUTH_SECRET; onyx.sh and onyx.ps1 generate it.
- OpenSearch replaced Vespa. Measured footprint: 7.7 GiB of RAM, 21 GB of
  images.

Walked end to end: deploy, model, MCP server (3 tools), agent, the cement
question (1,200), the stock update, the SSRF lab, and the switch to the
hardened server over --http (SSRF rejected). Uploading the PDFs for RAG
inside Onyx was not part of this pass; docs/VERIFICACION.md says so.
- docs/VERIFICACION.md (new): the record of the 30 Sep 2026 re-run. What
  was run, the environment and resolved versions, what had changed since
  July, what was added, and what was not covered.
- docs/PRESUPUESTO.md: measured cost. One full pass of Hours 1-3 is 39
  model calls and $0.012; Promptfoo $0.004; capped Garak $0.014. Per
  student: about 3 cents guided, 10-15 cents with repeats, about 30 cents
  with uncapped Garak, against a budget of $1.10.
- Facts re-checked on primary sources: gemini-3.5-flash-lite is still
  $0.30 / $2.50 per 1M tokens (output includes thinking tokens); AI Studio
  minimum purchase is now $5; depleted credit is documented as HTTP 402,
  which check_key.py now recognises alongside 429; Promptfoo needs Node 22.
- README: status section, links to the Hour 2 and Hour 3 guides, the real
  clone URL. SETUP: how to run the tests, and the separate environments
  for Garak and Promptfoo.
Installs the dependencies from scratch, at whatever versions resolve that
day, starts the lab database and runs the test levels that need no API key.
A dependency release that changes the workshop's behaviour, as mcp 2.0 did,
shows up here on its own.
The first CI run (Python 3.12 on Linux) computed a different descriptor
hash than Python 3.14: since 3.13 the interpreter strips the indentation of
docstrings, so a multi-line tool description differs by whitespace only.

hash_tools now collapses whitespace in the description before hashing, and
APROBADO is the hash of the hardened server under that rule. A changed
word, tool name or schema still changes the hash.

The workflow runs on pull requests, on pushes to main and every Monday.
@DavidMGDev
DavidMGDev merged commit 04f3668 into main Oct 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant