Hand-written NVFP4 W4A16 CUDA kernels for Volta
-
Updated
Aug 19, 2026 - Python
Hand-written NVFP4 W4A16 CUDA kernels for Volta
The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card. Ready-to-run PXQ models in MODELS.md; benchmarks vs llama.cpp in the README.
NVIDIA Tesla V100 显卡驱动自动化管理套件
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
llama.cpp optimized for NVIDIA Tesla V100 (Volta, SM70): 2–6 GPU tensor parallelism, Qwen3.8-27B Q8_0/Q4 with DFlash2 speculative decoding, up to 512K context, multimodal (image/PDF/video) and concurrent serving.
Qwen3.8-27B in native NVFP4/FP8 on 2x PCIe Tesla V100-32GB (SM70): the PCIe runbook for v100-skinny + 1Cat-vLLM, with the 3 fixes that make it work without NVLink. 61-74 tok/s decode, MTP speculative decoding, OpenAI-compatible.
Agentic-coding oriented inference engine for N x V100-SXM2, based on SGLang: GLM-5.3-Flash, Qwen3.8-Flash-Next, DeepSeek-V4.1-Flash, MiniMax-H3.
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
Fork adding clock mode for GPUs without P-states (V100, P100, pre-Volta) — upstream PR #10. A daemon that automatically manages the performance states of NVIDIA GPUs.
Reproducible llama.cpp kernel and runtime optimization lab for dual NVIDIA Tesla V100 GPUs (SM70)
Running large LLMs on pre-Ampere NVIDIA hardware — Tesla V100 (sm_70), RTX 2080 Ti (sm_75), CMP 170HX. Measured benchmarks, vLLM forks, and the hardware side: NVLink on SXM2 carrier boards, driver traps, cooling, used-kit acceptance.
Unofficial third-party patches: run Strata (Qwen3.8-Flash-Next) on 2x Tesla V100 sm_70 - platform patches, a PLE correctness fix that garbles output if unpatched, dual-GPU tuning to 43 tok/s @128k
Open hardware desktop AI node: 4× Tesla V100, 128GB HBM2, PCIe/NVLink topology and V-Core liquid/air cooling.
Runbook + benchmarks: Qwen3.8-Flash-Next-ABLITERATED NVFP4 on 4× Tesla V100-32GB (reflashed SXM2→PCIe, 2+2 NVLink + PLX). 1Cat-vLLM 1.5.0, TP4 — 262,144-token context validated, 46 tok/s decode, 122 tok/s aggregate at 4 concurrent streams. Full E0–E17 optimization log with measured evidence.
SGLang fork for IBM POWER9 (ppc64le): Tesla V100 sm70, CUDA 12.4, Granite LLM inference. Triton attention, float16, OpenAI-compatible API.
Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md
[Unmaintained, use dg1kjd/sglang-sxm2] Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB with vLLM, a downstream fork of 1Cat-vLLM.
Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
云途觉晓多人云音创作平台:基于 MOSS-TTSD 的简体中文多人语音、参考音色与 V100 部署方案
To associate your repository with the tesla-v100 topic, visit your repo's landing page and select "manage topics."