Skip to content

FlyCoder 0.2 beta: pure Ollama model for MacBook and Mac mini, harness removed - #3

Merged
Tromset merged 4 commits into
mainfrom
claude-code/keen-wright-0clt9o
Oct 3, 2026
Merged

Tromset merged 4 commits into
mainfrom
claude-code/keen-wright-0clt9o

Conversation

@Tromset

@Tromset Tromset commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Summary

FlyCoder becomes only the model. The 0.1 harness (CLI, web UI, Electron app, FlyBrain routing controller) is removed. FlyCoder 0.2 beta is an Ollama model with two variants, installable with one terminal command and publishable on ollama.com.

Tag Base Weights Target Context
flycoder:0.2-beta (latest) Gemma 4 12B 7.7 GB, NVFP4 on Ollama's MLX engine Macs with 16 GB+ 32,768
flycoder:0.2-beta-fast (fast) Qwen3.5 4B (the 0.1 base) 4.0 GB, NVFP4 on MLX Macs with 8 GB+ 16,384

Measured results

20 hidden-test coding problems. Run on a CPU-only Linux server (Ollama 0.35, GGUF Q4_K_M, thinking off as in the 0.1 harness, one sample per problem). Full answers and metrics are in docs/bench/cpu-nothink.json.

Profile Passed Decode Tokens/answer Time/problem
flycoder0.1beta (0.1) 5/20 6.6 tok/s 581 94 s
flycoder:0.2-beta-fast 5/20 6.5 tok/s 537 89 s
flycoder:0.2-beta 14/20 3.4 tok/s 739 235 s
  • The default variant solves almost 3× more problems, but takes about 2.5× the time per problem on CPU. Multi-token prediction helps: 77% of 6,594 drafted tokens were accepted (3.35 tokens per verification step), and generation was +49% faster than with drafting disabled. Ollama reports about +90% on Apple Silicon, so the gap should narrow on a Mac. This is not measured yet.
  • The fast variant keeps 0.1's speed and pass rate, and writes 8% fewer tokens. On a Mac it also gets the MLX engine.
  • Seven default-variant problems and one fast-variant problem were rerun. On the first pass the server ran out of memory or the bench's 5-minute fetch timeout cut the request. These were infrastructure failures, not answers. The bench now streams answers and retries a crashed runner once.

Changes

  • Modelfile, Modelfile.fast: the two variants. They use each base model's official sampling, set an explicit context (Ollama otherwise defaults to 4K below 24 GB), and add a short code-quality system prompt.
  • install.sh: curl -fsSL https://raw.githubusercontent.com/Tromset/flycoder/main/install.sh | sh. It checks the Ollama version (0.31 or later), picks the variant from the Mac's memory, and falls back to GGUF bases on Intel Macs, Linux and Windows.
  • scripts/publish.sh <user>: pushes 0.2-beta, latest, 0.2-beta-fast and fast to ollama.com. It refuses to push non-MLX builds as the main tags.
  • bench/: the benchmark runner and problems (npm run bench). On macOS it uses sandbox-exec (no network, no writes outside the scratch dir). bench/baselines/ keeps the 0.1 profile so the two versions can be compared.
  • tests/ + .github/workflows/test.yml: every bench test is checked against a reference solution and must reject a stub. Modelfiles, installer and publisher are checked for consistency. CI runs on Linux and macOS.
  • Harness code, its docs and its dependencies are removed (package.json now has no dependencies). The independent Qwen 2.5 fly-language QLoRA pipeline is kept.

Validation

  • npm test: 52/52 locally. CI is green on ubuntu and macos; on macOS the reference solutions run inside the sandbox.
  • Ollama 0.35.0: ollama create succeeds for both MLX Modelfiles. The models report tools, thinking and vision, and keep the gemma4 / qwen3.5 renderers and parsers.
  • curl … | sh installer tested from this branch (GGUF path on Linux).

Not verified here

  • Speed on Apple Silicon: MLX models cannot run in this Linux container. Run npm run bench on a Mac. If the default variant is too slow there, ollama cp flycoder:0.2-beta-fast flycoder:latest makes the fast one the default.
  • Thinking-mode quality and length were not benchmarked.
  • Publishing needs an ollama.com account and ollama signin on the Mac.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid

claude added 4 commits October 2, 2026 11:04
- Drop the CLI, web UI, Electron app and FlyBrain controller harness
- Modelfile: Gemma 4 12B (NVFP4, MLX, multi-token prediction), 32K context,
  Google's sampling, code-quality system prompt
- Modelfile.fast: Qwen3.5 4B (NVFP4, MLX), 16K context, Qwen's coding sampling
- install.sh: one-command install that picks the variant from Mac memory,
  with a GGUF fallback for Intel Macs, Linux and Windows
- scripts/publish.sh: push both variants to ollama.com
- bench/: 20 hidden-test coding problems to measure pass rate and speed,
  each test proven by a reference solution and rejecting a stub

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
The macOS job runs every bench reference solution inside sandbox-exec,
the path users take on their Macs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
Non-streamed requests send no byte until the answer ends, and fetch stops
waiting for headers after 5 minutes, so long answers on slow machines
were reported as errors. Crashed runners are now retried once with every
model unloaded, and --num-ctx lets machines short on memory shrink the
context without changing answers to these short prompts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
Gemma 4 12B (0.2-beta) solves 14/20 hidden-test problems against 5/20 for
0.1, at about 2.5x the time per problem on a CPU-only server. The fast
variant keeps 0.1's speed (6.5 vs 6.6 tok/s) with the same pass rate
without thinking. Full answers and metrics are in docs/bench.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
@Tromset
Tromset marked this pull request as ready for review October 3, 2026 08:22
@Tromset
Tromset merged commit 3f06dbd into main Oct 3, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants