OpenAI API server for OpenVino - when OVMS is too big
-
Updated
Jun 1, 2026 - Python
OpenAI API server for OpenVino - when OVMS is too big
Run any LLM locally on Intel Arc GPUs. Gemma, Qwen, vision and audio models optimized for budget hardware. · Dirty South Alpha™
Real-world LLM inference benchmarks on Intel Arc (Battlemage B60) — measured tok/s, quantization recipes, and MoE-vs-dense data, not synthetic claims. · Dirty South Alpha™
To associate your repository with the arc-b60 topic, visit your repo's landing page and select "manage topics."