A C++/CUDA inference engine for the Qwen3.5 model family (Qwen3.6 and Qwen3.8 27B, Qwen3.6-35B-A3B, their fine-tunes and Ternary Bonsai 2), Qwen3.8-Flash-Next and Gemma 4 26B-A4B, on the RTX 3090, RTX 4090, RTX 5090 and RTX PRO 6000 Blackwell. It serves the OpenAI, Anthropic and llama.cpp APIs from one GPU, or from up to eight as pipeline stages on Linux.
NInfer-all continues NInfer through the NInfer-3090 line and carries the work of several forks; CONTRIBUTORS.md credits each of them. The documentation map lists the guides and references.
Measured in September 2026, one card each, greedy, one request unless the row says otherwise. Setups and full tables: reference measurements.
| RTX 3090 | RTX 4090 | RTX 5090 | RTX PRO 6000 | |
|---|---|---|---|---|
| Ternary Bonsai 2 27B, short chat (DFlash2, 7 drafts) | 202 tok/s | 256 tok/s | 397 tok/s | 381 tok/s |
| decode after a 261K-token document (fastest drafter) | 90 tok/s | 123 tok/s | 218 tok/s | 218 tok/s |
| time to first token for a 261K-token prompt | 215 s | 102 s | 82 s | 78 s |
| largest context, filled and all three needles found | 970,752 | 958,464 | 978,944 | 1,048,576* |
| eight requests at once (MTP, 3 drafts), total | 551 tok/s | 824 tok/s | 1,063 tok/s | 1,155 tok/s |
| Qwen3.8-27B, short chat (DFlash2, 7 drafts) | 118 tok/s | 149 tok/s | 236 tok/s | 237 tok/s |
| largest context, filled and all three needles found | 417,792 | 405,504 | 872,448 | 1,048,576* |
| eight requests at once (MTP, 3 drafts), total | 329 tok/s | 442 tok/s | 690 tok/s | 739 tok/s |
* The engine's ceiling. The RTX PRO 6000 (96 GB) starts there with every KV format and drafter; filled to it, both models find two of the three needles.
- RTX PRO 6000 against RTX 5090. Re-measured in October beside an RTX 5090, both at 600 W, the PRO 6000 came back within 2% of September wherever the drafts accepted the same share. One request's decode is bound by memory bandwidth, and both cards have 1.79 TB/s of GDDR7 (Bonsai 2 without speculation: 167.7 against 170.4 tok/s). The PRO 6000's 188 SMs against 170 show where compute decides: the 261K prompt is 2% faster and eight requests at once 6 to 7% faster (re-check).
- Draft length. DFlash2 with seven drafts is fastest on short answers; after long documents the best count is three to seven. MTP runs up to fifteen drafts and is fastest at three to five.
- Past the native window. Filled to about 880K tokens, Bonsai 2 returned all three planted codes on every card. At 1,048,576 tokens (RTX 5090 and PRO 6000 only) it misses the one at 943K.
One request, greedy, October 2026. Each cell is decode of a short answer · prefill of a 4,463-token prompt, in tok/s. Decode after the long prompt, memory, host links and power limits: Qwen3.8-Flash-Next.
Q2_0 (35.9 GiB model + 26.8 GiB n-gram table):
| Experts | RTX PRO 6000 | RTX 5090 | RTX 4090 | RTX 3090 |
|---|---|---|---|---|
| on the GPU (two cards, except the PRO 6000) | 138 · 2,939 | 136 · 3,809 | 97 · 3,608 | 90 · 1,504 † |
| in pinned host memory | 55 · 1,297 | 74 · 1,677 | 53 · 1,138 | 49 · 842 |
| on disk, the files in the page cache | 76 · 2,024 | 68 · 1,739 | 42 · 723 | 47 · 630 |
| on disk, cold (pages evicted every second) | 35 · 1,000 | 25 · 694 | 18 · 160 | 17 · 195 |
† Two RTX 3090 Ti.
The host and disk rows were measured before the October 9 decode work. Host experts on an RTX 3090 now decode 87-94 tok/s in a different workload, 98-100 with a repeated prompt (host experts on one RTX 3090).
IQ3_S (51.9 GiB model + the same table):
| Experts | RTX PRO 6000 | RTX 5090 | RTX 3090 |
|---|---|---|---|
| on the GPU (two cards for the RTX 5090) | 125 · 2,541 | 122 · 3,326 | — |
| in pinned host memory | 29 · 803 | 41 · 1,100 | 34 · 533 |
| on disk, the files in the page cache | 66 · 1,661 | 55 · 439 | 19 · 209 ‡ |
| on disk, cold | 29 · 773 | 19 · 171 | 11 · 47 |
— Not measured: every expert needs three 24 GB cards. ‡ The host's 62 GB of RAM cached only part of the file.
Coder IQ1_M (28.4 GiB model + the same table), one NVIDIA L40S (48 GB):
| Experts | NVIDIA L40S |
|---|---|
| in pinned host memory | 35 · 1,648 |
| on disk, the files in the page cache | 42 · 1,938 |
- The host matters. Host and disk rows depend on the host as much as on the card. The RTX 5090 host had PCIe 5.0 x16; the others had PCIe 4.0 x16.
- The PRO 6000 caches almost everything. Its 96 GB device expert cache ends up holding nearly every expert.
- RTX 4090 host pinning. The RTX 4090 rows ran pinned to the GPUs' NUMA node in a two-socket VM. Unpinned, host experts decode there at 40 tok/s and disk experts at 32 to 33.
- Many requests at once. Six reasoning requests at once on two RTX 5090s produced 193 tok/s of output with Q2_0.
Models and weight formats
-
Qwen3.5 family. Dense and MoE (
Qwen3_5ForCausalLM,Qwen3_5MoeForCausalLM) with MTP, DFlash (35B-A3B) and DFlash2 (Qwen3.8-27B) drafters, Vision for images and video, andgroupwise-int, FP8, NVFP4 and ternary weights. NVFP4 expert banks prefill through W4A4 and need a120abuild. -
GGUF block formats. Releases that pick a ggml type per tensor, such as ISTA-DASLab's GSQ-RCO, convert without requantization and multiply their blocks in place. The 3.5-bit Qwen3.8-27B IQ3_S release holds 10.95 GiB of weights instead of 15.9, scores WikiText-2 perplexity 7.071 (the official artifact: 7.286) and decodes at 59.9 tok/s against the official artifact's 40.3 on an RTX 3090 (GGUF block formats).
-
Ternary Bonsai 2 27B. PrismML's ternary Qwen3.8-27B runs from
t2_g128_fp16weights: 2.125 bits per weight, imported without rounding, with the Hadamard rotations fused into the norms and gates and integer activations for its projections. Thebonsai2_27b_ternaryrecipe adds ProCreations' MTP head and DFlash2 adapter and an exact proposal head. -
Qwen3.8-Flash-Next. The 125B-parameter MoE (512 experts, about 6B active) from the GSQ-RCO GGUF releases, with its experts on the GPUs, in pinned host memory with a GPU cache (
--expert-residency host) or in the artifact's files read into a GPU cache (--expert-residency disk, under 1 GB of RAM). It serves up to eight requests with prefix reuse, structured output, images and video, and MTP drafting. On two RTX 5090s the Q2_0 release scores 93.3% on AIME 2025 and 86.4% on GPQA-Diamond (Qwen3.8-Flash-Next). -
Gemma 4 26B-A4B. Google's QAT Q4_0 GGUF release, imported block for block (13.4 GiB): the sliding-window and full-attention blocks, the dense MLP beside 128 routed experts, up to eight concurrent requests with prefix reuse, tool calls and structured output. On an RTX 3090 it prefills a 10,000-token prompt at 5,695 tok/s and decodes at 208 tok/s, 177 after 10,000 tokens, against llama.cpp's 4,427, 167 and 151 on the same GGUF (Gemma 4).
Speculative decoding
- MTP, DFlash and DFlash2 with 1 to 15 drafts; the artifact's proposal head drafts by default when it stores one. Under MTP and DFlash a sampled request keeps a draft only while it equals the model's own keyed draw, so every emitted token is the model's sample; DFlash2 verifies its proposal distribution by rejection sampling.
--adaptive-mtppicks each round's width from measured draft survival and round cost.- N-gram copy drafting beside the neural drafter (on with
--spec) verifies up to 15 tokens copied from earlier prompt, tool-result or output text (ngram copy proposals).
Context and memory
- Nine KV formats, from
bf16tork2v4-e8at 216 bytes per token and KV head;rk8v4is the all-round choice andrk4v4holds the native 262,144-token window of Qwen3.8-27B on a 24 GB card (context and memory). - Up to 1,048,576 tokens of context, four times the native window, with plain RoPE or YaRN
(
--rope-yarn). - A context cache of checkpoints on the device, in pinned host memory and in an optional disk tier
that survives restarts (
--disk-kv-path,--disk-kv-restore), with endpoint and branch anchors that let an edited or regenerated turn resume where it left the conversation. The hybrid prefix cache (--use-alt-prefix-caching) shares content-addressed KV blocks across requests instead. - Precision trades that buy context:
--embedding-q4,--lm-head-q6,--gdn-state-fp16(quality trades).
Serving
- OpenAI Responses and Chat Completions, Anthropic Messages, llama.cpp's
/completion,/tokenize,/detokenize,/apply-template,/v1/rerank,/metrics,/slotsand/props, streaming, tool calls, structured output through xgrammar (--structured-output), token log probabilities and prompt grafts (HTTP serving). - Model suspend.
--model-suspendlets an idle server give its device memory back without exiting and take it again on the next request, retained conversations included. Ternary Bonsai 2 27B on an RTX 3090: 7.9 GiB down to 0.3 GiB in 0.33 s, back in 1.1 s (model suspend). - Several models behind one server. With
--models-diror a llama.cpp-style--models-preset,ninfer-serveis a router with llama.cpp's model API; with--model-suspenda model sleeps instead of unloading: two 27B models on one 24 GB RTX 3090 swap in 1.8 s instead of an 11.6 s cold load (several models). GET /statsand a monitoring dashboard, request logs with rotation, and an optional compiled-in WebUI.
Kernels and devices
- Device route profiles. Each card's measured profile picks the kernel route for every
operation and width. Profiles are built in for the RTX 3090, 4090, 5090 and the three RTX PRO
6000 editions; any other GPU is measured once at first start, and
ninfer-calibratere-measures (device profiles). - Pipeline stages.
--devices 0,1,...splits the layers over up to eight GPUs on Linux, each stage owning its weights, KV and state (pipeline stages). - FP8 A8 and NVFP4 W4A4 tensor-core routes on every
120abuild, an FP4 tensor-core prompt kernel for an NVFP4 KV cache, and programmatic dependent launches on Blackwell.
Download an artifact from the table below and serve it from the Docker image or from a build. The server speaks the OpenAI and Anthropic APIs, and each card picks up its device profile on its own.
hf download WaveCut/Ternary-Bonsai-2-27B-NInfer-v3 Ternary-Bonsai-2-27B-ninfer-v3.ninfer --local-dir models
# The image: `serve`, the artifact under /models, the flags. Listens on http://localhost:8080/v1.
docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
-v "$PWD/models:/models" -v ninfer-cache:/cache \
ghcr.io/iamwavecut/ninfer-all serve /models/Ternary-Bonsai-2-27B-ninfer-v3.ninfer \
--model-id bonsai2-27b --max-context 262144 --kv-capacity 262144 --kv-dtype rk8v4 \
--gdn-state-fp16 --spec dflash2 --draft-tokens 5
# A build: the same flags. Listens on 127.0.0.1:8080 unless --host and --port say otherwise.
ninfer-serve models/Ternary-Bonsai-2-27B-ninfer-v3.ninfer --model-id bonsai2-27b \
--max-context 262144 --kv-capacity 262144 --kv-dtype rk8v4 --gdn-state-fp16 \
--spec dflash2 --draft-tokens 5The recipes below are the configurations the reference tables
use. They are written for a build; in the container, replace ninfer-serve models/ with the
docker run ... serve /models/ line above.
Ternary Bonsai 2 27B: DFlash2, MTP, the largest context, a disk tier
# Fastest single stream: DFlash2 with five drafts over the full 262,144-token window.
ninfer-serve models/Ternary-Bonsai-2-27B-ninfer-v3.ninfer --model-id bonsai2-27b \
--max-context 262144 --kv-capacity 262144 --kv-dtype rk8v4 --gdn-state-fp16 \
--spec dflash2 --draft-tokens 5
# MTP drafting through the proposal head.
ninfer-serve models/Ternary-Bonsai-2-27B-ninfer-v3.ninfer --model-id bonsai2-27b \
--max-context 262144 --kv-capacity 262144 --kv-dtype rk8v4 --gdn-state-fp16 \
--spec mtp --draft-tokens 3
# The largest context a 24 GB card holds: 958,464 tokens of rk4v4.
ninfer-serve models/Ternary-Bonsai-2-27B-ninfer-v3.ninfer --model-id bonsai2-27b \
--max-context 958464 --kv-capacity 958464 --kv-dtype rk4v4 --gdn-state-fp16 --rope-yarn
# Adaptive MTP and a disk tier that keeps evicted conversations.
ninfer-serve models/Ternary-Bonsai-2-27B-ninfer-v3.ninfer --model-id bonsai2-27b \
--max-context 198400 --kv-capacity 198400 --kv-dtype rk8v4 --gdn-state-fp16 \
--spec mtp --draft-tokens 5 --adaptive-mtp \
--disk-kv-path /var/cache/ninfer --disk-kv-gib 64 --disk-kv-restore- Draft count. Five drafts are the all-round choice. Seven are faster on short answers; after long documents three to seven win (draft length).
- Images.
--vision --vision-residency overlay --vision-max-merged 12288adds them. The encode borrows the drafter's memory, so the whole window still fits a 24 GB card. - Past 958,464 tokens. The RTX 5090 and the PRO 6000 hold the 1,048,576-token maximum with
rk4v4, DFlash2 or MTP included; the RTX 5090 holds 978,944 withrk8v4. Filled to 1,048,576 tokens, the model misses the code at 90% (about 943K); up to about 880K it found every code on every card.
Qwen3.8-27B: the official artifact and the GSQ-RCO IQ3_S release
# The official artifact on a 24 GB card: DFlash2 with five drafts over 245,760 tokens of rk4v4.
ninfer-serve models/qwen3_8_27b.ninfer --model-id qwen3.8-27b \
--max-context 245760 --kv-capacity 245760 --kv-dtype rk4v4 --gdn-state-fp16 \
--spec dflash2 --draft-tokens 5
# GSQ-RCO IQ3_S: MTP over 176,128 tokens of rk8v4.
ninfer-serve models/Qwen3.8-27B-GSQ-RCO-IQ3_S-ninfer-v3.ninfer --model-id qwen3.8-27b \
--max-context 176128 --kv-capacity 176128 --kv-dtype rk8v4 --gdn-state-fp16 \
--spec mtp --draft-tokens 3- The official artifact with
rk8v4. The same speculation fits 167,936 tokens on an RTX 4090 and 176,128 on an RTX 3090. An RTX 5090 takes the full 262,144 with either KV format. - IQ3_S. It drafts with
--spec dflash2 --draft-tokens 5as well, and--visionadds images.
Qwen3.8-Flash-Next: experts on the GPUs, in host memory or on disk
hf download WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3 \
Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer --local-dir models
hf download WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 \
Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer --local-dir models
M=models/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer
T=models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer
# Every expert on one RTX PRO 6000 (96 GB).
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --max-context 32768
# Every expert on two GPUs, one pipeline stage each (Linux): 24 GB cards for Q2_0, 32 GB for IQ3_S.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --max-context 32768 --devices 0,1
# One 24 GB GPU: the experts in pinned host memory, the most used of them cached on the GPU.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --max-context 32768 --expert-residency host
# One 24 GB GPU and little RAM: the experts stay in the file and stream into a GPU cache.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --max-context 32768 --expert-residency disk
# The container, experts in host memory.
docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
-v "$PWD/models:/models" -v ninfer-cache:/cache \
ghcr.io/iamwavecut/ninfer-all serve /models/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer \
--ngram-table /models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer \
--model-id qwen3.8-flash-next --max-context 32768 --expert-residency host- Other releases. IQ3_S and the Coder IQ1_M build take the same flags with their own file and the same table. The tables above show the placements measured for each.
- Memory. Host experts pin 34 GB of RAM for Q2_0, 50 GB for IQ3_S and 25 GB for the Coder
build. Disk experts need under 1 GB of RAM; the page cache does the rest. The GPU expert cache
takes what is free after startup, or
--expert-cache-mib. - The n-gram table. Its rows are read from the file, 16 per token, unless
--ngram-residency ramloads all 28.8 GB orram-hotkeeps the rows a profile ranks first in a RAM budget (the n-gram rows). A model started without its table is refused;--no-ngram-tableoverrides that, an experimental mode with no practical use. - Drafting.
--spec mtp --draft-tokens 3drafts with the MTP block the published artifacts carry. - Images and video.
--visionadds the Vision tower (0.9 GB on the GPU).
Gemma 4 26B-A4B: converted from Google's QAT GGUF
# The release's GGUF and its Hugging Face text files; one conversion writes the artifact.
mkdir -p models/gemma4/hf
for f in config.json tokenizer.json tokenizer_config.json chat_template.jinja generation_config.json; do
curl -fsSL -o models/gemma4/hf/$f https://huggingface.co/google/gemma-4-26B-A4B-it/resolve/main/$f
done
hf download google/gemma-4-26B-A4B-it-qat-q4_0-gguf gemma-4-26B_q4_0-it.gguf --local-dir models/gemma4
python -m tools.convert --model models/gemma4/hf --recipe gemma4_gguf \
--source gguf=models/gemma4/gemma-4-26B_q4_0-it.gguf --name Gemma-4-26B-A4B-it-QAT \
--out models/gemma-4-26B-A4B-it-qat-q4_0.ninfer --device cpu
# Four concurrent requests over 40K-token contexts; thinking stays off unless a request asks.
ninfer-serve models/gemma-4-26B-A4B-it-qat-q4_0.ninfer --model-id gemma4 \
--max-context 40960 --max-concurrency 4Qwen3.6-35B-A3B NVFP4: RTX 50 series and RTX PRO 6000 only
ninfer-serve models/Qwen3.6-35B-A3B-NVFP4-ninfer-v3.ninfer --model-id qwen3.6-35b-a3b \
--max-context 262144 --kv-capacity 262144 --kv-dtype rk8v4 --gdn-state-fp16 \
--spec mtp --draft-tokens 3 --visionThe NVFP4 experts prefill through W4A4 on Blackwell's FP4 tensor cores, so this needs a 120a
build (the image has one); sm_86 and sm_89 builds refuse the file. On an RTX 5090 it starts in
17 s and leaves 7.4 GiB of the card free.
Qwen3.8-27B fine-tunes: Huihui abliterated, HauhauCS Aggressive, MXFP8-CRACK
All three have the official Qwen3.8-27B identity and size, so the Qwen3.8-27B recipes above serve them too; CRACK has no DFlash2 adapter. Their cards' configurations:
# Huihui abliterated on one RTX 3090: 198,400 tokens with MTP and Vision.
ninfer-serve models/Huihui-Qwen3.8-27B-abliterated-ninfer-v3.ninfer --model-id qwen3.8-27b \
--max-context 198400 --kv-capacity 198400 --kv-dtype rk8v4 --gdn-state-fp16 \
--spec mtp --draft-tokens 3 \
--vision --vision-residency overlay --vision-max-merged 12288
# HauhauCS Aggressive: DFlash2 with seven drafts.
ninfer-serve models/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-DFlash2-ninfer-v3.ninfer \
--model-id qwen3.8-27b --max-context 32768 --kv-capacity auto \
--spec dflash2 --draft-tokens 7
# MXFP8-CRACK: MTP.
ninfer-serve models/Qwen3.8-27B-MXFP8-CRACK-ninfer-v3.ninfer --model-id qwen3.8-27b \
--max-context 32768 --kv-capacity auto --spec mtp --draft-tokens 3A card without a built-in profile
ninfer-calibrate --print > my-gpu.jsonThe engine measures it by itself at first start. Running it by hand refreshes the profile after a driver or clock change. See device profiles.
The Bonsai, GSQ-RCO and Flash-Next artifacts use formats that only this engine reads.
| model | artifact | size | contents |
|---|---|---|---|
| Ternary Bonsai 2 27B | WaveCut/Ternary-Bonsai-2-27B-NInfer-v3 | 8.87 GiB | ternary weights, Vision, Bonsai-trained MTP head and DFlash2 adapter, exact proposal head |
| Qwen3.8-27B | neroued/Qwen3.8-27B-NInfer | 19.03 GiB | the official groupwise-int (Q4/Q5) with MTP and DFlash2; the reference tables' artifact |
| Qwen3.8-27B GSQ-RCO IQ3_S | WaveCut/Qwen3.8-27B-GSQ-RCO-IQ3_S-NInfer-v3 | 13.99 GiB | ISTA-DASLab's 3.5-bit GGUF blocks byte for byte, Q6_K MTP head, Vision, DFlash2, proposal head |
| Qwen3.8-27B, abliterated | WaveCut/Huihui-Qwen3.8-27B-abliterated-NInfer-v3 | 19.03 GiB | the official qwen3_8_27b recipe with MTP, DFlash2 and a proposal head |
| Qwen3.8-27B Uncensored, HauhauCS Aggressive | WaveCut/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-DFlash2-NInfer-v3 | 19.03 GiB | groupwise-int with the tune's MTP head, DFlash2 and a proposal head |
| Qwen3.8-27B MXFP8-CRACK | WaveCut/Qwen3.8-27B-MXFP8-CRACK-NInfer-v3 | 16.96 GiB | groupwise-int with MTP and a proposal head; no DFlash2 |
| Qwen3.8-Flash-Next GSQ-RCO Q2_0 | WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3 | 38.49 GiB | ISTA-DASLab's 2.4-bit GGUF blocks byte for byte, Vision, shared-Q8_0 MTP; reads the n-gram table |
| Qwen3.8-Flash-Next GSQ-RCO IQ3_S | WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-NInfer-v3 | 54.50 GiB | the 3.5-bit release with Vision and shared-Q8_0 MTP |
| Qwen3.8-Flash-Next Coder GSQ-RCO IQ1_M | WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Coder-IQ1_M-NInfer-v3 | 31.02 GiB | the expert-pruned coding build (256 experts per layer), Vision, shared-Q8_0 MTP |
| Qwen3.8-Flash-Next n-gram table | WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 | 26.92 GiB | unchanged IQ4_NL rows plus an optional broad hot-row profile (--ngram-table) |
| Qwen3.6-35B-A3B NVFP4 | WaveCut/Qwen3.6-35B-A3B-NVFP4-NInfer-v3 | 20.39 GiB | RedHatAI's NVFP4 experts code for code, Q8 projections, Vision, MTP, proposal head; sm_120a GPUs only |
The official NInfer artifacts (Qwen3.6-27B, Qwen3.6-35B-A3B and Qwen3.8-27B, groupwise-int and
nvfp4) load too; the documentation map links them. Weight
conversion shows how the Bonsai and
GSQ-RCO artifacts are built.
ghcr.io/iamwavecut/ninfer-all:latest is built from every master commit that passes CI, on CUDA
13.4 and Ubuntu 26.04. It carries two builds and starts the one that matches the GPU: sm_86 for
the RTX 30 and RTX 40 series, sm_120a for the RTX 50 series and the RTX PRO 6000 Blackwell.
- Host. An NVIDIA driver of the CUDA 13 branch (580 or newer) and the NVIDIA Container Toolkit.
- Tags.
latest,sha-<commit>and theVERSION; the image is 1.6 GB compressed. - Volumes.
/modelsholds artifacts,/cachethe device profile measured on first start, and/grafts/<model>/optional prompt grafts. - Checked on an RTX 3090 and an RTX 5090 with driver 580.159.03.
The container's command chooses what runs:
| command | runs |
|---|---|
serve, ninfer, perplexity, calibrate [args] |
that binary with any artifact and flags; serve listens on 0.0.0.0:8080 unless given --host or --port |
run <model> [profile] (the default: run qwen38-27b) |
the launcher profiles of scripts/run.sh for the official qwen38-27b and qwen36-35b-a3b, with its NINFER_* overrides (-e NINFER_SPEC=mtp, -e NINFER_CONTEXT=131072, ...) |
download <model> |
scripts/download-model.sh into /models: qwen38-27b, qwen36-27b or qwen36-35b-a3b |
The official Qwen3.8-27B with the measured tuned profile, on http://localhost:8080/v1:
docker run --rm -v "$PWD/models:/models" ghcr.io/iamwavecut/ninfer-all download qwen38-27b
docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
-v "$PWD/models:/models" -v ninfer-cache:/cache \
ghcr.io/iamwavecut/ninfer-all run qwen38-27bcompose.yaml wires the GPU, the port and the volumes for the same:
docker compose run --rm ninfer download qwen38-27b, then docker compose up -d.
NINFER_IMAGE_ARCH=sm86|sm120a overrides the GPU detection, and docker build -t ninfer . builds
the image from source (--build-arg ARCHS=86 for one architecture).
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release -DCMAKE_CUDA_ARCHITECTURES=86
cmake --build build --target ninfer-serve ninfer-calibrateCMAKE_CUDA_ARCHITECTURES is 86 for the RTX 30 series (that build also runs the RTX 40 series),
89 for the RTX 40 series and 120a for the RTX 50 series and the RTX PRO 6000 Blackwell (on the
mma.sync compatibility path, which the ternary route needs). CUDA 12.8 or newer builds 86 and
89; a 120a build needs CUDA 13.1 or newer, since 12.8 and 12.9 miscompile sm_120a kernels and
configure refuses them. The Linux and Windows
guides cover dependencies, the opt-in build options and the launcher scripts;
tests and benchmarks have their own guides.
Apache-2.0, as upstream. The Bonsai artifact's weights come from PrismML, ProCreations and Qwen, all Apache-2.0; its card lists the notices. The Qwen3.8-Flash-Next artifacts carry the Qwen Community License 1.0 of their model.