A browser-based game show simulator: thirteen contestants draft trivia
domains, duel head-to-head on a chess clock, and fight to become sole owner
of the board for a $100,000,000 grand prize. Each player is backed by a
local LLM that makes their live in-show decisions and answers trivia; if a
call fails, that player transparently falls back to a scripted stand-in
agent so the show never stalls (see
src/dominion/agents/llm_agent.py).
This repo is a fork of the original Ollama-only prototype, in progress on
porting the inference layer to llama.cpp and raising the engineering
around it to a real service architecture, so the two can be run side by
side and compared on real numbers (latency, live-vs-fallback reliability,
accuracy) rather than by feel. See
docs/ENGINE_NOTES.md for the original engine's
full design history. Current status: both backends run end to end on the
new async service architecture, including llama.cpp's GBNF
grammar-constrained decisions (see "llama.cpp setup" below) -- pick either
one with the Ollama/llama.cpp toggle on the start screen (falls back to
scripted per-decision on any failure, same as always, regardless of which
one you pick). The comparison view segmenting the Standings page by
backend is the next milestone; until then, recorded history isn't yet
labeled by which backend produced it.
This is an intentional, explicit dependency tradeoff. The original
prototype was proud of being zero-pip-dependency, standard-library-only
Python. This fork trades that for production service scaffolding — an
async web framework, structured config, structured logging, metrics,
containerization — because that scaffolding is the actual point of this
fork. The game engine itself (src/dominion/engine/) still has zero
third-party imports; every dependency in pyproject.toml is
service/transport-layer only.
A short demonstration video of the original Ollama-only version is
available on Loom.
Dominion live
Want to see a full show end-to-end without running it yourself?
Dominion-Video/dominion-full-run-4x.mp4
(~30MB, optional — nothing in this repo depends on it) is a complete game
recorded on this fork, sped up 4x with audio pitch-corrected so it's a
six-minute watch instead of twenty-five. The 4x speed is a video edit, not
how the show actually plays — Dominion is TV-game-show-paced even at
normal speed (a Host, a separate Commentator, audience reactions, "pushing
on" cards between duels); the edit just compresses that same rhythm. An
earlier, unedited clip of the original prototype
(dominion-gameplay-2026-07-old.mp4, same folder) is kept for historical
comparison.
- Python 3.10+.
- Ollama — optional. Powers the Ollama backend's live agents; without it, every player just uses the scripted fallback.
- A llama.cpp
llama-serverbinary — optional. Powers the llama.cpp backend once you pick it via the start-screen toggle; see "llama.cpp setup" below. Without it running (or not picked at all), every player just uses the scripted fallback, same as Ollama.
python check_env.py
python -m venv .venv && source .venv/bin/activate # or .venv\Scripts\activate on Windows
pip install -e ".[dev]"check_env.py checks your Python version, and if Ollama is installed,
pulls the models the live agents use. It's a preflight check, not a
package installer — pip install -e ".[dev]" is what actually installs
the app and its dependencies (declared in pyproject.toml).
uvicorn dominion.server.app:app --port 8765Then open http://localhost:8765. Before clicking Start Show, the
start screen lists every model actually installed for the selected
backend (real Ollama tags via /api/tags, real .gguf files under
DOMINION_LLAMACPP_MODELS_DIR) with its on-disk size, pre-checked with
sensible defaults so Start stays one click if you don't touch it.
Checking/unchecking a model live-updates a memory budget against this
machine's actual available RAM (a conservative estimate, not a real
per-backend load/swap simulation -- see
server/model_catalog.py) and
greys out picks that would no longer fit. Set DOMINION_SCRIPTED_ONLY=1
first to skip live model calls entirely and run near-instantly, useful
for quick local iteration.
DOMINION_SCRIPTED_ONLY=1 uvicorn dominion.server.app:app --port 8765See docs/ENGINE_NOTES.md for the engine layout,
live-agent details, and environment variables
(DOMINION_SCRIPTED_ONLY, DOMINION_PORT), and
design/Game Show Sim - Design Document.docx
for the full rule set and design rationale.
The backend selector (the toggle on the start screen, or
?backend=llamacpp on /api/run-show directly) already picks this
backend for a real show end to end — you just need an actual
llama-server running for it to talk to.
- Build or download
llama-serverfrom ggml-org/llama.cpp. - Create a models directory and put GGUF files in it, named to match
inference/config.py'sLLAMACPP_MODELSregistry (edit that mapping if you'd rather name things differently, or to add/remove models — it's a plain dict, not generated). Roughly the same four families the Ollama backend uses, any Q4_K_M-or-similar quantization:Llama-3.2-3B-Instruct-Q4_K_M.ggufQwen2.5-3B-Instruct-Q4_K_M.ggufgemma-2-2b-it-Q4_K_M.ggufPhi-3-mini-4k-instruct-Q4_K_M.ggufQwen2.5-7B-Instruct-Q4_K_M.gguf— a real GGUF counterpart for whichever bigger Ollama models you've pulled beyond the curated four above (e.g.qwen2.5:7b); bigger than the rest of this roster (~4.7GB vs ~2GB), so bump--models-maxbelow if you add it.muse-glimmer-Q4_K_M.gguf— reserved slot for Meta's real "Muse Glimmer" 30B multimodal model; deliberately left as a placeholder here since its smallest public quant is 17GB+, wildly out of scale with everything else in this roster and a real risk to local memory/VRAM headroom if added casually.
- Run
llama-serverin router mode (no-m, so it serves every model in the directory from one process/port, the same way Ollama swaps between models on one daemon) —--models-maxshould match however many real GGUF files you've actually placed in the directory.-ngl 999(or--n-gpu-layers 999) is not optional if you have a GPU — without it, llama-server runs fully on CPU even with a CUDA-built binary, driver, and card all present and working (this bit Scott directly: RTX 2080, CUDA 13 driver, 8% GPU util the whole show). Make sure you actually downloaded a CUDA-tagged release from ggml-org/llama.cpp's releases too (not the plaincpubuild) — the flag alone can't offload if the binary itself has no CUDA support compiled in:One real tradeoff once GPU offload is actually on: an 8GB card can't hold every GGUF in this roster fully offloaded at once if you've added the biggerllama-server --models-dir ./models --models-max 4 --n-gpu-layers 999 --port 8080
qwen2.5-7b-instruct(~4.7GB) alongside the original four (~2GB each, ~8GB total) — pick a smaller subset per show via the model picker rather than always running all 5, or lower--n-gpu-layersfrom 999 to a smaller number to partially offload the biggest model instead of evicting others. DOMINION_LLAMACPP_URLdefaults tohttp://127.0.0.1:8080— override it if you ranllama-serveron a different host/port.
Every inference call (either backend) logs structured JSON to stdout (model, backend, retries, latency, outcome) and records to Prometheus, scraped at:
GET /metrics
dominion_inference_requests_total{backend,model,outcome},
dominion_inference_latency_seconds{backend,model}, and
dominion_inference_fallback_total{backend,model} (see
observability/metrics.py) --
no external Prometheus/Grafana setup is required to just look at the
endpoint directly.
docker compose upBrings up the app (port 8765) and a llama-server in router mode (port
8080, serving whatever GGUFs you've put in ./models -- see "llama.cpp
setup" above). Ollama itself isn't a compose service (containerizing it
well needs real GPU passthrough this file isn't trying to own) -- the app
container points at DOMINION_OLLAMA_URL=http://host.docker.internal:11434
by default, i.e. an Ollama already running on your host; override that
env var in docker-compose.yml if yours lives elsewhere. History persists
across restarts via the dominion-data named volume (DOMINION_DB_PATH,
set in the Dockerfile, points the SQLite file there instead of into the
installed package directory, which the container's non-root user can't
write to).
pytestpython scripts/run_show_batch.py --count 20Runs a batch of shows headlessly (no server, no browser) and prints the
same aggregate win-rate/reliability/latency stats web/stats.html shows,
so a real balance question can be answered from a real sample size
instead of watching one show at a time. --scripted-only skips live
model calls entirely for a fast, cheap sample of just the scripted
heuristics; --backend/--models match the start screen's own picker;
--db-path points at a scratch SQLite file instead of writing into your
real recorded history. --help for the full option list.
The steps above assume some command-line familiarity. If you just want to watch a show run without setting any of this up yourself, the demo video above is the easiest way in — running the actual server still needs Python and (optionally) Ollama installed and the commands above typed into Terminal (macOS/Linux) or PowerShell (Windows), there's no one-click installer for either backend yet.
Built by Scott A. Cole, an AI strategy consultant and the author of 31 books on AI strategy, GenAI, and
decision-making, including the five-book Stop Learning AI series for executives who need to
make good AI decisions without becoming technical themselves. The app includes an About page and a Books page listing every title and edition (src/dominion/web/about.html, src/dominion/web/books.html; /about.html and /books.html with the server running). More projects, including Aegis Vector and Palimpsest, are at
github.com/ScottColeSW.
As an Amazon Associate I earn from qualifying purchases.
MIT — see LICENSE.