Repository navigation
FlyCoder 0.2 beta: pure Ollama model for MacBook and Mac mini, harness removed - #3
Merged
Merged
Conversation
- Drop the CLI, web UI, Electron app and FlyBrain controller harness - Modelfile: Gemma 4 12B (NVFP4, MLX, multi-token prediction), 32K context, Google's sampling, code-quality system prompt - Modelfile.fast: Qwen3.5 4B (NVFP4, MLX), 16K context, Qwen's coding sampling - install.sh: one-command install that picks the variant from Mac memory, with a GGUF fallback for Intel Macs, Linux and Windows - scripts/publish.sh: push both variants to ollama.com - bench/: 20 hidden-test coding problems to measure pass rate and speed, each test proven by a reference solution and rejecting a stub Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
The macOS job runs every bench reference solution inside sandbox-exec, the path users take on their Macs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
Non-streamed requests send no byte until the answer ends, and fetch stops waiting for headers after 5 minutes, so long answers on slow machines were reported as errors. Crashed runners are now retried once with every model unloaded, and --num-ctx lets machines short on memory shrink the context without changing answers to these short prompts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
Gemma 4 12B (0.2-beta) solves 14/20 hidden-test problems against 5/20 for 0.1, at about 2.5x the time per problem on a CPU-only server. The fast variant keeps 0.1's speed (6.5 vs 6.6 tok/s) with the same pass rate without thinking. Full answers and metrics are in docs/bench. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid
Tromset
marked this pull request as ready for review
October 3, 2026 08:22
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
FlyCoder becomes only the model. The 0.1 harness (CLI, web UI, Electron app, FlyBrain routing controller) is removed. FlyCoder 0.2 beta is an Ollama model with two variants, installable with one terminal command and publishable on ollama.com.
flycoder:0.2-beta(latest)flycoder:0.2-beta-fast(fast)Measured results
20 hidden-test coding problems. Run on a CPU-only Linux server (Ollama 0.35, GGUF Q4_K_M, thinking off as in the 0.1 harness, one sample per problem). Full answers and metrics are in
docs/bench/cpu-nothink.json.flycoder0.1beta(0.1)flycoder:0.2-beta-fastflycoder:0.2-betaChanges
Modelfile,Modelfile.fast: the two variants. They use each base model's official sampling, set an explicit context (Ollama otherwise defaults to 4K below 24 GB), and add a short code-quality system prompt.install.sh:curl -fsSL https://raw.githubusercontent.com/Tromset/flycoder/main/install.sh | sh. It checks the Ollama version (0.31 or later), picks the variant from the Mac's memory, and falls back to GGUF bases on Intel Macs, Linux and Windows.scripts/publish.sh <user>: pushes0.2-beta,latest,0.2-beta-fastandfastto ollama.com. It refuses to push non-MLX builds as the main tags.bench/: the benchmark runner and problems (npm run bench). On macOS it usessandbox-exec(no network, no writes outside the scratch dir).bench/baselines/keeps the 0.1 profile so the two versions can be compared.tests/+.github/workflows/test.yml: every bench test is checked against a reference solution and must reject a stub. Modelfiles, installer and publisher are checked for consistency. CI runs on Linux and macOS.package.jsonnow has no dependencies). The independent Qwen 2.5 fly-language QLoRA pipeline is kept.Validation
npm test: 52/52 locally. CI is green on ubuntu and macos; on macOS the reference solutions run inside the sandbox.ollama createsucceeds for both MLX Modelfiles. The models report tools, thinking and vision, and keep the gemma4 / qwen3.5 renderers and parsers.curl … | shinstaller tested from this branch (GGUF path on Linux).Not verified here
npm run benchon a Mac. If the default variant is too slow there,ollama cp flycoder:0.2-beta-fast flycoder:latestmakes the fast one the default.ollama signinon the Mac.🤖 Generated with Claude Code
https://claude.ai/code/session_01MDAB5YFn5ZSH37mcH3Tfid