You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
keys-TensorFold Studio (Mac+DGXSpark 64GB VRam Min): Qwen-Image-2.1-Turbo images and MiniMax H3 / FastH3 video with sound on TensorFold. Apple silicon (Metal) and NVIDIA DGX Spark (CUDA), one install command.
NVFP4 BIZ: nvidia/GLM-5.3-Flash-NVFP4 as distributed, on DGX Spark-class GB10 systems. 2.x serves it with TensorFold (TP=2 or TP=3, FP8 KV, MTP, drafted replies equal serial); 1.x with a pinned vLLM (TP=2 or TP=3, images, optional AXL repack). Apache-2.0 code, MIT weights fetched separately. BIZ = business-use intent, not support or certification.
MiniMax H3 audio-video model on TensorFold NVFP4 kernels: a faster ComfyUI loader for one RTX 50-series GPU (2.3x stock at 20 steps, 81 s per 5 s 1344x768 video with Turbo + sparse attention on an RTX 5070 Ti)
Keep every AI model on a NAS; insert any of them on your GPU box with one click (Ollama / vLLM) — never evicts a model you didn't name. Stdlib Python, demo mode, web UI.
Unofficial fork of TensorFold v0.6.5 (Python engine line; upstream main is now Zig): EXL3 3.05 bpw + a prompt-lookup (suffix) drafter for Qwen3.8-Flash-Next on CUDA — ~2.4x EXL3 prefill, native vision/video, image + long-prefix caching, every result byte-exact. 14 patches, fixtures, token-hash receipts (provenance: issue 444).
GLM-5.3-Flash on TensorFold, 4x RTX PRO 6000: two custom patches, tuned knobs, bench scripts and every measured result (incl. the ones that lost) on top of the Aevonix recipe
Mia's Qwen3.8-Flash-Next recipe for two DGX Sparks (TensorFold Zig engine, NVFP4) plus a CJK draft vocabulary: Chinese decode +38-107%, English -2 to -3%, replies byte-identical.
One command: GLM-5.3-Flash with the Dealign o_proj abliteration transplant on Mia's TensorFold recipe (2x DGX Spark). Byte-verified, stock-speed decode.
GLM-5.3-Flash on 1x RTX 5090 + 2x DGX Spark: attention on the 5090, routed experts on the Sparks over MCDMA RoCE, on TensorFold v0.6.5 (+ 0.6.6's commits) with Mia's GLM work. 2.0 at tag v2.0, v1.0 (glm53f-afd) at tag v1.0. Deployment recipe, as-is.
Helm chart + Ansible to serve GLM-5.3-Flash-EXL3 (EXL3 4bpw, FP8 KV, DFlash2) tensor-parallel on 2× DGX Spark (GB10) joined to an existing k3s cluster via TensorFold v0.6.0