FreeToken Benchmarks
Real tokens-per-second results across consumer GPUs, laptops, and workstations. Prefill and decode throughput with exact test conditions.
What are Benchmarks?
Benchmarks measure FreeToken performance across consumer laptops, gaming PCs, and workstation GPUs, focusing on tokens per second, prefill and decode speed, memory use, and CPU-GPU bandwidth. Every result lists the exact test conditions, and ft bench bw helps you decide between offload and hybrid MoE execution on your own machine.
Why read benchmarks?
Set Realistic Expectations
See the actual tokens/s you can expect from your GPU class
Compare Before You Buy
Weigh RTX 3090, 4090, and 5090 throughput against your budget
Tune Your Setup
Use ft bench bw results to choose between offload and hybrid execution
Featured & Essential
FreeToken tokens per second: Benchmark Guide for 2026
Learn how to read FreeToken tokens per second results, compare hardware, and improve local MoE inference performance in 2026.
FreeToken single gpu: Local AI Setup Guide
Learn how FreeToken uses one GPU, system RAM, and adaptive MoE execution to run large local AI models.
All Benchmarks
FreeToken 290b model: 2026 Edge MoE Setup Guide
Learn how FreeToken serves frontier-scale MoE models on consumer hardware through adaptive caching, bandwidth scheduling, and elastic memory.
FreeToken 753b model: Setup Guide and MoE Tuning Tips
Learn how FreeToken serves 753B GLM-5.2 locally, including setup, memory paging, MoE caching, benchmarks, and practical tuning advice.
FreeToken benchmark: Local MoE Serving Setup Guide
Review the FreeToken benchmark, local MoE serving design, hardware requirements, installation flow, and measured performance results.