Hand-written NVFP4 W4A16 CUDA kernels for Volta
-
Updated
Aug 19, 2026 - Python
Hand-written NVFP4 W4A16 CUDA kernels for Volta
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
Run Qwen3.6-27B on four Tesla V100s at 366 tok/s using hand-written NVFP4 CUDA kernels and chain-MTP speculation.
llama.cpp optimized for NVIDIA Tesla V100 (Volta, SM70): 2–6 GPU tensor parallelism, Qwen3.8-27B Q8_0/Q4 with DFlash2 speculative decoding, up to 512K context, multimodal (image/PDF/video) and concurrent serving.
Qwen3.8-27B in native NVFP4/FP8 on 2x PCIe Tesla V100-32GB (SM70): the PCIe runbook for v100-skinny + 1Cat-vLLM, with the 3 fixes that make it work without NVLink. 61-74 tok/s decode, MTP speculative decoding, OpenAI-compatible.
Qwen3.8-Flash-Next (125B MoE, NVFP4) on 4x V100-SXM2-32GB and DeepSeek-V4.1-flash on 8x V100 — a Volta port of SGLang for agentic coding.
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
Reproducible llama.cpp kernel and runtime optimization lab for dual NVIDIA Tesla V100 GPUs (SM70)
SGLang fork for IBM POWER9 (ppc64le): Tesla V100 sm70, CUDA 12.4, Granite LLM inference. Triton attention, float16, OpenAI-compatible API.
Runbook + benchmarks: Qwen3.8-Flash-Next-ABLITERATED NVFP4 on 4× Tesla V100-32GB (reflashed SXM2→PCIe, 2+2 NVLink + PLX). 1Cat-vLLM 1.5.0, TP4 — 262,144-token context validated, 46 tok/s decode, 122 tok/s aggregate at 4 concurrent streams. Full E0–E17 optimization log with measured evidence.
DP4A FlashAttention-2 and GEMM for Volta GPUs. 46 TOP/s INT8 on CMP 100-210 where tensor cores are firmware-disabled. The .superl8 format loads weights at memory speed.
V100 (sm_70) tuning kit for Ternary Bonsai 2 27B: q8_0-KV flash-attention direct read, D256 Split-D prefill kernel, PTQ1_0 planar mat-vec, MTP graft + draft micro-batch fix, ubatch presets. Prefill 175 -> 854 t/s, VRAM -1.9 GiB, PPL bit-identical (2-chunk). 中英双语 README (README.md / README.zh-CN.md).
Serve Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB (DGX-1, TP8) for agentic coding & ops — a downstream fork of 1Cat-vLLM.
VastLLM: a production-oriented FastLLM fork for native C++ inference, V100/SM70, long context, and Qwen3.8/3.6/3.5 series; upstream: ztxz16/fastllm
PyTorch 2.12 fork for IBM POWER9/POWER10 (ppc64le) with CUDA 12.4 and Triton. Tesla V100 sm70, GPU training and LLM inference.
Measured llama.cpp and quantization results on Tesla V100 (sm_70), Pascal GTX 1070, and RTX 4070 - hardware most projects don't test on
Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md
OpenAI Triton compiler fork for IBM POWER9/POWER10 (ppc64le) with CUDA 12.4. Tesla V100 sm70 GPU kernels for PyTorch and SGLang.
To associate your repository with the sm70 topic, visit your repo's landing page and select "manage topics."