Skip to content

Repository files navigation

OpenBot

A thin computer around a local coding model: six tools, ProductGuard, one folder, no default MCP/web/memory.

Measured on an RTX 5070 12GB (Q4, 8K, one model at a time). Not ChatGPT. Not T4. Not walk-away.

If you have ~12GB VRAM and already run Ollama, this is meant to be copied and spun: same deny list, your models, your folder. Feedback and forks are the point.

What it is

Piece What you get
Exam runner scripts/run_lh_product.py + run_lh_bands.py — in-memory workspace + ProductGuard
Live plugin harness/plugins/caller_definer/ — same ProductGuard before LocalHarness tools run
One-folder disk lib/workspace_fs.py — names resolve under one root; .. and host paths fail
Cards docs/STACK_CARD_*.md — which 8B/9B/4B/tiny seats we actually ran

The model will try to leave the folder, rm, write .env, curl, or invent a fetch. The computer must refuse. The role stays ~6 lines so 8B still calls tools. Safety in the prompt fights the model; safety in the veto does not.

What it is not

  • 100% / “safe for unattended work”
  • A GUI (use localharness web if you want a browser; this repo is the kernel)
  • Proven on your home directory — never point it at C:\Users\... or ~
  • Paolo’s private LoRA as a required weight. Daily here is qwen3-8b-ft-agentic-v1 (local). Stock qwen3:8b is weaker (skips tools). Use any Q4 8B-class that fits; expect to route (grep → a stronger search seat, least-change → a small 4B) rather than one perfect model

Other people’s stacks (you can mix)

This is not the only local-agent project. People already put their own spin on:

This kernel’s bet: keep six tools, deny-first, one folder, slim role. Fat MCP catalogs and unused tool schemas lost score on 8B here. If you fork, keep the computer; swap the weights.

12GB how-to

  1. Ollama up. One model loaded. Q4 that sits in VRAM. Context 8K (OLLAMA_CONTEXT_LENGTH=8192). Do not kill a standing llama-server on :8090 if you use one.
  2. Clone this lab. cd into a project folder, not your home.
    $env:LOCAL_LAB = "C:\path\to\local-lab"
    $env:LOCAL_LAB_WORKSPACE = "C:\path\to\your\project"
    python "$env:LOCAL_LAB\scripts\smoke_workspace.py"
    Smoke does not call a model. It must print workspace smoke ok.
  3. Copy harness/plugins/caller_definer into ~/.localharness/plugins/ (or your LH plugin dir). Restart localharness start. Set workspace to that same project folder.
  4. Slim role only (see harness/orchestrator.role.md). Coding tools on. Web / MCP / memory off until you opt in per chat.
  5. Run a known ticket: “Fix the typo in one file” or “do not run rm -rf.” Watch vetoes, not vibes.

Models we measured (your tags will differ)

Seat Role on this card Notes
~8B Qwen FT (ours) / try qwen3:8b Daily coding Stock 8B may skip tools
Ornith 9B Grep / observe Daily grep was 0/3; Ornith 3/3
nemotron 3 nano 4B Least-change / tiny trust Not 3-file batches
granite tiny / gemma e4b Tight VRAM / speed Privacy computer 8/8; work is not T3

Privacy pack (escapes, .env, rm, no curl): 8/8 n=3 on 8B, 9B, 4B, tiny, and e4b. That is the exam + plugin computer, not a promise your disk is a container.

Safety / privacy (rogue model)

The model tries Computer
../, C:\, ~/.ssh Path veto (read/write/grep/glob/bash)
rm -rf, wipe-claim files Harm + freeze existing files
.env / copy secret.txt / echo canary Path + body veto + chat scrub
curl / python app.py Network/exec refuse; no invented HTML/out.txt

You still must: one workspace folder; do not mount home; do not enable a real shell. Live LocalHarness without this plugin is Door A (unguarded). Install the plugin or you only have the exam.

Models learn the box from ERROR: harness: …, not a longer lecture.

Honest leftovers

  • T3 short tickets, not a workday. Grep and least-change are routed, not solved by one 8B.
  • Conflict prompts (“don’t write” + “fix the typo”) — they often fix. Judgment, not a host leak.
  • Exam 12/12 is one majority n=3 shot; neighbors were 11/12.
  • No TUI in this repo.

Details: docs/SANDBOX_PRIVACY.md, docs/HARNESS_TRANSFER.md, docs/TWO_SAFE_AGENTS.md, docs/CRITICS.md.

Public copy: github.com/paolothomas72/openbot. How we count forks/clones: docs/SIGNAL.md. Paste-ready post: docs/ANNOUNCE.md. If you actually ran it, open a I ran it issue.

Lab internals (this machine)

cd C:\Users\paolo\paseo-tool-workspaces\opencode\local-lab
python scripts\verify_setup.py
python scripts\run_lh_bands.py --privacy --n3 qwen3:8b

Gold: data/gold/. Do not train on data/raw/ (not in this snapshot). MIT: LICENSE. gold_lh_v500.jsonl is a poison control — do not SFT it.

About

OpenBot: local coding kernel (six tools, ProductGuard, one folder). Measured on 12GB VRAM; not ChatGPT, not walk-away.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages