Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
-
Updated
Aug 20, 2026 - Python
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
Run LLMs on AMD Vega8 / Vega10 APUs and GPUs (Ryzen 5700G / gfx90c and Vega 56/64): ROCm 7 gfx900 backport + Vulkan/RADV launchers, Docker images, benchmarks for llama.cpp and LM Studio
Monorepo of 4 healthcare AI agent skills: chart review, HEDIS NLP, HCC NLP, HIPAA compliance. Install: npx skills add Yar177/medical-chart-review-skill --skill '*'
Agentic auditor that flags unsubstantiated CMS-HCC risk-adjustment codes in clinical documentation, grounded against a deterministic Rust scoring engine as a verifiable oracle. Built on synthetic data.
Local speech-to-text toolkit for Arch Linux. Transcribes audio/video to clean Markdown via whisper.cpp (Vulkan/RADV). CLI-first: local files, YouTube, Telegram, dictation, LLM cleanup. Clean text by default, timestamps optional. GPU: AMD RX 580 via RADV
Add a description, image, and links to the radv topic page so that developers can more easily learn about it.
To associate your repository with the radv topic, visit your repo's landing page and select "manage topics."