MiniCPM5-2B IS INSANE! 128K Context & Deep Reasoning Runs Locally - OpenBMB MiniCPM5-2B local execution and context benchmarking
-
Updated
Sep 8, 2026 - Python
MiniCPM5-2B IS INSANE! 128K Context & Deep Reasoning Runs Locally - OpenBMB MiniCPM5-2B local execution and context benchmarking
One-click deployment of Qwen3.8-27B on a single RTX 3090 (24GB) on Windows 11: Q4_K_XL quantization + MTP speculative decoding + 128K context window, exposed as an OpenAI-compatible llama-server, averaging ~47 tok/s with ~1.5× lossless MTP speedup
To associate your repository with the 128k-context topic, visit your repo's landing page and select "manage topics."