Bonsai 2 27B: 262k context (trained max) on 12 GB, 54 to 97 tok/s decode, 2x prefill, q8_0 KV recipe, grammar tool calls 9/9. Same ternary weights. Windows bundle + receipts.
-
Updated
Sep 30, 2026 - Python
Bonsai 2 27B: 262k context (trained max) on 12 GB, 54 to 97 tok/s decode, 2x prefill, q8_0 KV recipe, grammar tool calls 9/9. Same ternary weights. Windows bundle + receipts.
Prebuilt CUDA wheels (sm_89) + build guide to run Microsoft TRELLIS.2 and TencentARC Pixal3D image-to-3D in ComfyUI on NVIDIA RTX 40-series (Ada) under Debian 13 / Linux. Fixes the 'no kernel image is available' (sm_120 vs sm_89) problem.
Fork of karpathy/nanochat pushed onto one RTX 4070 12GB: a 910M-param base model from 530 GPU-hours, OOM-safe training automation, and 14 documented negative results.
Reproducible Qwen3.5 MTP and D-Flash benchmarks on an RTX 4070: throughput, acceptance rate, quantization, and language effects.
Measured llama.cpp and quantization results on Tesla V100 (sm_70), Pascal GTX 1070, and RTX 4070 - hardware most projects don't test on
Локальный LLM-провайдер для проекта slovo — Ollama c Laguna XS.2 (33B/3B MoE) в Docker с проброшенной NVIDIA GPU. Оптимизирован под single-user агентские сессии на RTX 4070 Ti SUPER + i9-11900K.
To associate your repository with the rtx-4070 topic, visit your repo's landing page and select "manage topics."