Minimal SGLang patch for Qwen3.8-27B + DFlash2 on dual RTX 3080: TP-sharded Draft fc and static per-head FP8 KV.
-
Updated
Sep 2, 2026 - Python
Minimal SGLang patch for Qwen3.8-27B + DFlash2 on dual RTX 3080: TP-sharded Draft fc and static per-head FP8 KV.
Local LLM benchmarks on one RTX 3080 (10 GB): max stable context, quality, accelerators — EN/RU/DE. Author: https://homensai.com/
To associate your repository with the rtx-3080 topic, visit your repo's landing page and select "manage topics."