FreeToken Supported Models
Complete list of officially supported models including GLM-5.2, DeepSeek-V4-Flash, Qwen3.6, and gpt-oss. Compare sizes, formats, and quantization options.
What are Supported Models?
The supported models reference lists every model known to run well on FreeToken, from GLM-5.2 and DeepSeek-V4-Flash to Qwen3.6 MoE and gpt-oss. Each entry covers checkpoint format, available quantizations, and which FreeToken MoE backend serves it best, so you can pick a model that fits your hardware on the first try.
Why check supported models?
Find Known-Good Checkpoints
Skip trial and error with officially supported and verified model builds
Match Format to Backend
Pair checkpoint formats and quantizations with the right MoE backend
Fit Models to Hardware
Compare model sizes against your VRAM before downloading gigabytes of weights
Featured & Essential
FreeToken qwen3 8 27b: Setup Guide & Safe Testing
Learn how to evaluate, configure, and safely test a FreeToken qwen3 8 27b endpoint with practical prompts, limits, and troubleshooting steps.
FreeToken qwen: Local Qwen MoE Setup Guide
Learn how FreeToken runs Qwen MoE models locally, manages VRAM limits, and improves inference for coding and agent workflows.
All Model Guides
FreeToken deepseek v4 flash: Local Setup Guide 2026
Set up DeepSeek V4 Flash with FreeToken for local AI inference, including RAM planning, GPU expectations, Open WebUI access, and troubleshooting.
FreeToken deepseek: Local AI Setup Guide & Benchmarks
Learn how FreeToken serves DeepSeek locally with MoE caching, CPU-GPU offload, hardware guidance, setup steps, and performance notes.
FreeToken glm 5 2: Access Guide and Model Overview
Learn how to evaluate the FreeToken glm 5 2 search term, verify model access, compare capabilities, and use AI tools safely.