# FreeToken RoadMap - 2026 ### New Model Support - [x] GLM-5.3-Flash @andy-yang-1 #332 - [x] Qwen3.8-Flash-Next @jason-fxz - [x] pin PLE in RAM — landed in #257 - [x] offload PLE to Disk #311 - [ ] DeepSeek-V4.1-Flash ### Hardware Support - [ ] macOS: a native Metal engine on Apple Silicon Macs. - [ ] AMD: support for AMD GPUs via ROCm. - [ ] DGX Spark: aarch64 wheels, sm_121 kernels, and a unified-memory mode. ### Multimodal - [ ] Support Image input for vision-language models - [ ] Qwen3.5 / Qwen3.6 / Qwen3.8 family @jason-fxz - [ ] Qwen3.8-Flash-Next - [ ] Gemma-4 family ### Multi-GPU - [ ] DGX Spark: multi-GPU optimisation for linked DGX Spark units. - [ ] Tensor parallelism (TP): serve one model across multiple GPUs on one machine. ### Quantization - [x] Quant Layer Refactor: - [x] #418 - [x] #427 - [ ] GGUF: support GGUF checkpoints and their quantization types across model architectures. ### Speculative Decoding - [ ] Speculative decoding, including MTP / DFlash / Dspark. --- Any contribution is welcome. Join the [Developer Slack](https://join.slack.com/t/flashml/shared_invite/zt-3zpdh5j10-9dwTXrgLiqpVxizhA9KVbA) if you want to get involved. --- TODO: This is still WIP.
FreeToken RoadMap - 2026
New Model Support
Hardware Support
Multimodal
Multi-GPU
Quantization
Speculative Decoding
Any contribution is welcome. Join the Developer Slack if you want to get involved.
TODO: This is still WIP.