🎯
Focusing
AI Engineer / Solution Architect · LLM & dLLM researcher · AWS Integrations
Pinned Loading
-
llm2vec-8gb-offload-lossless-batching
llm2vec-8gb-offload-lossless-batching PublicRun LLM2Vec (Llama-3-8B) on an 8 GB GPU via RAM offload, and batch it without changing the embeddings (padding shifts them: cosine 0.76 vs 0.999956)
Python
-
vergilius
vergilius PublicAssistente AI locale per capire cosa succede nel mondo, ovunque: modelli piccoli, macchine modeste, nessun fornitore esterno. Odysseus + ShadowBroker.
Python
-
blackwell6000-qwen3.8-flash-next
blackwell6000-qwen3.8-flash-next PublicBlackwell 6000 + qwen 3.8 flash next: Swift 1.5 Qwen3.8 Flash-Next NVFP4 on one RTX PRO 6000 with vLLM, 4 users x 227k context, ~100 tok/s each, exact kernels + determinism fix
Cuda
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.