The results
| Model | Worker and runtime | Measured result | Workload |
|---|---|---|---|
| Qwen3.8-Flash-Next NVFP4 | 1× RTX 5090 · FreeToken offload | c8: 115.249 aggregate output tok/s c16: 123.079 aggregate output tok/s | All requests passed. c8 delivered 32.428 median steady decode tok/s/request; c16 delivered 11.614. The c16 worker admitted all 16 requests without queueing. |
| DeepSeek-V4-Flash | 3× RTX 5090 · FreeToken | 30.62 end-to-end tok/s ~33.3 decode tok/s | Warmed 255-token probe |
| Qwen3.6 35B-A3B | 3× RTX 5090 · FreeToken | 33.76 end-to-end tok/s ~41.0 decode tok/s | Warmed 256-token probe |
| GLM-5.2 NVFP4 | 4× RTX PRO 6000 Blackwell Max-Q · FreeToken Triton offload | 7.158 end-to-end tok/s 9.235 steady decode tok/s | Warmed 255-token probe; served successfully, below the 12 tok/s qualification target |
| GLM-5.3-Flash NVFP4 | 2× RTX PRO 6000 Blackwell Server Edition · two independent FreeToken workers | c8: 40.048 aggregate output tok/s c16: 64.998 aggregate output tok/s | Requests split evenly across the workers; 24/24 passed. Median steady decode: 10.214 tok/s/request at c8 and 7.793 at c16. |
| GLM-5.3 full NVFP4 | 1× H200 NVL · FreeToken 0.1.2 · Vast Serverless | c8: 17.782 aggregate output tok/s c16: 18.023 aggregate output tok/s | Authenticated 128-token requests; 24/24 passed. Eight active slots meant the second eight c16 requests queued. |
| Kimi-K3 0.40B dev checkpoint | 1× RTX 3090 · FreeToken offload | 567.043 serial output tok/s 32/32 requests passed in the first concurrency ladder | Architecture proof only; not comparable to the full 2.8T model |
| Kimi-K3 full checkpoint | 1 of 4 B300 GPUs · FreeToken development revision | Did not become inference-ready | All 96 resident shards loaded; serial expert-bank construction remained active at the spend cutoff |
| Kimi-K3 full checkpoint | 1 of 4 NVIDIA H200 GPUs · FreeToken, parallel expert-bank load | 1.211 end-to-end tok/s 1.690 steady decode tok/s | Three warmed 127-token streams; 16/16 requests completed at the tested concurrency ceiling |
| Qwen3.6 35B-A3B | 8× RTX 5060 Ti · vLLM tensor parallelism 8 | 1,574 aggregate output tok/s | 64 concurrent requests; two successful trials |
| GPT-OSS 120B | 8× RTX 5060 Ti · vLLM tensor parallelism 8 | 1,451 aggregate output tok/s | 64 concurrent requests; two successful trials |
| Gemma4 31B NVFP4 | 8× RTX 5060 Ti · vLLM tensor parallelism 8 | Did not complete | Out of memory during warm-up at an 8k context window |
Do not turn this table into one ranking
The FreeToken rows include warmed single-request probes on multi-GPU workers and a separate concurrency test on one RTX 3090. The vLLM rows measure aggregate output across a fixed ladder on eight lower-memory GPUs. Those workloads are not interchangeable.
Qwen3.8-Flash-Next: more than 80 decode tok/s on one RTX 5090
On 20 September, FreeToken served RadixArk/Qwen3.8-Flash-Next-NVFP4 on one 32 GB RTX 5090 in New Jersey. The worker used QSA sparse attention, NVFP4 Triton kernels, disk-backed expert loading, and automatic MoE offload. Its measured PCIe gather bandwidth was 51.92 GB/s.
Two warmed 1,023-token streams measured 80.918 median steady decode tok/s, 73.832 median end-to-end tok/s, and 1.226 seconds median time to first token. Three shorter 127-token streams measured 75.057 median steady decode tok/s.
At c8, all requests ran concurrently and produced 740 output tokens in 6.421 seconds: 115.249 aggregate output tok/s. Median steady decode was 32.428 tok/s/request, median time to first token was 3.061 seconds, and the 528 prompt tokens imply 172.5 effective aggregate prefill tok/s.
A separate c16 run raised the worker limit and CUDA graph ceiling to 16. All 16 requests were admitted with no queue and produced 1,543 output tokens in 12.537 seconds: 123.079 aggregate output tok/s. Median steady decode fell to 11.614 tok/s/request while aggregate output rose only 6.8% over c8. For this worker and workload, c8 is the practical operating point.
Evidence: the machine-readable benchmark summary includes the hardware profile, single-stream results, and c1/c2/c4/c8/c16 concurrency ladder.
GLM-5.3-Flash: independent workers and Serverless qualification
On 20 September, two independent FreeToken workers ran on separate 96 GB RTX PRO 6000 Blackwell Server Edition GPUs in Czechia. Requests were assigned round-robin, four per worker at c8 and eight per worker at c16. The c8 batch produced 1,134 output tokens in 28.316 seconds for 40.048 aggregate output tok/s. The c16 batch produced 2,229 output tokens in 34.293 seconds for 64.998 aggregate output tok/s. All 24 requests completed successfully.
At c8, median prefill was 10.080 seconds, observed prefill throughput was 2.587 prompt tok/s, and median steady decode was 10.214 tok/s/request. At c16, median prefill was 7.728 seconds, observed prefill throughput was 3.365 prompt tok/s, and median steady decode was 7.793 tok/s/request. This topology served separate sessions on each GPU; it did not use tensor parallelism.
On 30 August, FreeToken loaded LibertAIDAI/GLM-5.3-Flash-NVFP4 on one 32 GB RTX 5090 with hybrid expert offload. The qualified run measured 16.840 steady decode tok/s over 126 steps, 59.38 ms/token, 7.13 seconds warm TTFT, and 29.21 GiB peak server VRAM.
A separate RTX PRO 6000 Workstation Edition worker produced the exact requested semantic answer and completed every request in the concurrency test. At c8, 376 output tokens took 14.452 seconds for 26.017 aggregate output tok/s. At c16, 752 output tokens took 24.042 seconds for 31.279 aggregate output tok/s. The worker had eight resident request slots, so the c16 result includes queueing.
Vast Serverless also completed the lifecycle test on an RTX PRO 6000 S worker in the Netherlands. Cold loading took 262.69 seconds. Cached resume reached readiness in 109 seconds, inside the 300-second startup deadline, and the authenticated semantic probe returned the exact expected sentinel. The Serverless platform benchmark reported 20.572 measured workload units per second at concurrency four; that platform metric is separate from the output-token rates above.
Evidence: the 20 September dual-worker record, RTX 5090 qualification, earlier RTX PRO 6000 concurrency results, semantic result, and Serverless lifecycle record are published as machine-readable JSON. The implementation and validation history remain available in FreeToken pull request 292.
Full GLM-5.3: authenticated Serverless qualification
On 7 September, FreeToken 0.1.2 loaded the full 432.9 GiB LibertAIDAI/GLM-5.3-NVFP4 checkpoint on one H200 NVL and served it through an authenticated Vast Serverless endpoint. Autoscaled worker creation, model loading, and Serverless routing all completed successfully.
At c8, all 8 requests completed with 1,016 reported output tokens in 57.136 seconds: 17.782 aggregate output tok/s. At c16, all 16 completed with 2,032 tokens in 112.747 seconds: 18.023 aggregate output tok/s. The worker admitted at most eight active requests, so the second eight c16 requests queued. Nearly doubling wall time without a corresponding throughput gain shows that this worker saturated around c8 for this workload.
A separate semantic smoke test returned exactly 42 with a stop finish reason in 10.402 seconds. Every benchmark request used a 128-token output cap, so these results measure authenticated routing and throughput, not completed-answer quality.
Evidence: all 24 request results, the semantic smoke result, and the qualification record. Idle stop and cached restart remain unverified because the instance was deleted after the successful live tests. The model-support work remains tracked in FreeToken PR #406.
What the Kimi run proved
FreeToken loaded the public eight-layer, 0.40B MXFP4 Kimi-K3 development checkpoint and served it through the OpenAI-compatible API on a 24GB RTX 3090. A 256-token serial request produced 255 reported output tokens in 0.450 seconds, or 567.043 tokens/s. The first ladder completed all 32 simultaneous requests.
This proves the new KDA/MLA, latent-MoE, SiTU, MXFP4 loading, batching, and API path. It says nothing about how fast the full 2.8T Kimi-K3 model will run.
During a sustained follow-up, the live-swapped NVIDIA driver lost the passed-through GPU on the second 16-request sample. We kept the completed samples and recorded the failure. We did not turn the earlier burst peak into a sustained-throughput claim.
The full 1.56 TB checkpoint downloaded and all 96 resident-weight shards loaded on one 275040 MiB B300. Resident weights used 111134 MiB of GPU memory. FreeToken then reported that Kimi had no parallel expert-bank reader and used its serial build path. The API still returned 503 model is still loading after 5 minutes 52 seconds in the worker, so the spend guard stopped the run before generation. Throughput and maximum concurrency therefore remain unmeasured.
On 24 August, a separate 4× H200 qualification run used FreeToken's parallel expert-bank loader. The full, machine-readable record is versioned with the FreeToken change; the values below are included so this page remains independently auditable.
| Cold-start milestone | Measured time from launch |
|---|---|
| Resident weights loaded | 55 s |
| Expert banks loaded | 19 m 39 s |
| CUDA graph capture | 31 s |
First successful /v1/models | 21 m 30 s |
| Concurrent requests | Aggregate completion tok/s | Median TTFT | Batch elapsed |
|---|---|---|---|
| 1 | 0.645 | 30.295 s | 48.047 s |
| 2 | 0.971 | 32.846 s | 63.845 s |
| 4 | 1.339 | 35.007 s | 92.637 s |
| 8 | 1.822 | 38.433 s | 136.126 s |
| 16 | 2.400 | 41.698 s | 206.627 s |
Warmup used a 101-token prompt and generated 15 tokens: 34.803 s TTFT, 42.665 s elapsed, and 1.781 steady decode tok/s. Three subsequent 127-token streams had 30.350 s median TTFT, 1.211 median end-to-end tok/s, and 1.690 median steady decode tok/s. Every request in the 1/2/4/8/16-client ladder completed (31/31 total).
An earlier H100 run failed because FreeToken selected tensor parallel size 1 and the 106.55 GiB resident set could not fit in 79.18 GiB of usable GPU memory. A Canadian H200 attempt hit repeatable provider-volume I/O errors. Those failed instances were deleted with operator approval. The B300 instance is stopped and its 1.75 TB checkpoint volume is preserved for a follow-up after the serial startup path is fixed.
Review the FreeToken implementation and H200 evidence, the machine-readable H200 record, and the development-checkpoint artifacts.
What this changes for routing
Parameter count is not enough. Placement needs the weight format, active memory footprint, context limit, runtime, and expected concurrency.
A cheap community worker can be right for one workload and unable to load another. We keep raw results and failed runs so that decision can be checked later.
Qwen3.8-Flash-Next now has a measured RTX 5090 operating point. Qwen3.8 27B remains a separate candidate for a 24 GB RTX 3090.
Method notes
The vLLM runs used fixed 128-token completions, temperature 0, a fixed prompt seed, tensor parallelism across eight GPUs, and two trials at each concurrency level through 64. FreeToken requests streamed over the OpenAI-compatible API. The Kimi run ignored EOS to exercise the requested token budget. We saved metrics, not generated text or credentials. These are throughput tests, not quality scores or maximum-capacity claims.