Benchmark notes · Updated 20 September 2026

Open-model inference, measured on real workers.

These are our passing runs and our failures. They are not a universal leaderboard: different runtimes, GPUs, context lengths, and request patterns answer different questions.

The results

ModelWorker and runtimeMeasured resultWorkload
Qwen3.8-Flash-Next NVFP41× RTX 5090 · FreeToken offloadc8: 115.249 aggregate output tok/s
c16: 123.079 aggregate output tok/s
All requests passed. c8 delivered 32.428 median steady decode tok/s/request; c16 delivered 11.614. The c16 worker admitted all 16 requests without queueing.
DeepSeek-V4-Flash3× RTX 5090 · FreeToken30.62 end-to-end tok/s
~33.3 decode tok/s
Warmed 255-token probe
Qwen3.6 35B-A3B3× RTX 5090 · FreeToken33.76 end-to-end tok/s
~41.0 decode tok/s
Warmed 256-token probe
GLM-5.2 NVFP44× RTX PRO 6000 Blackwell Max-Q · FreeToken Triton offload7.158 end-to-end tok/s
9.235 steady decode tok/s
Warmed 255-token probe; served successfully, below the 12 tok/s qualification target
GLM-5.3-Flash NVFP42× RTX PRO 6000 Blackwell Server Edition · two independent FreeToken workersc8: 40.048 aggregate output tok/s
c16: 64.998 aggregate output tok/s
Requests split evenly across the workers; 24/24 passed. Median steady decode: 10.214 tok/s/request at c8 and 7.793 at c16.
GLM-5.3 full NVFP41× H200 NVL · FreeToken 0.1.2 · Vast Serverlessc8: 17.782 aggregate output tok/s
c16: 18.023 aggregate output tok/s
Authenticated 128-token requests; 24/24 passed. Eight active slots meant the second eight c16 requests queued.
Kimi-K3 0.40B dev checkpoint1× RTX 3090 · FreeToken offload567.043 serial output tok/s
32/32 requests passed in the first concurrency ladder
Architecture proof only; not comparable to the full 2.8T model
Kimi-K3 full checkpoint1 of 4 B300 GPUs · FreeToken development revisionDid not become inference-readyAll 96 resident shards loaded; serial expert-bank construction remained active at the spend cutoff
Kimi-K3 full checkpoint1 of 4 NVIDIA H200 GPUs · FreeToken, parallel expert-bank load1.211 end-to-end tok/s
1.690 steady decode tok/s
Three warmed 127-token streams; 16/16 requests completed at the tested concurrency ceiling
Qwen3.6 35B-A3B8× RTX 5060 Ti · vLLM tensor parallelism 81,574 aggregate output tok/s64 concurrent requests; two successful trials
GPT-OSS 120B8× RTX 5060 Ti · vLLM tensor parallelism 81,451 aggregate output tok/s64 concurrent requests; two successful trials
Gemma4 31B NVFP48× RTX 5060 Ti · vLLM tensor parallelism 8Did not completeOut of memory during warm-up at an 8k context window

Do not turn this table into one ranking

The FreeToken rows include warmed single-request probes on multi-GPU workers and a separate concurrency test on one RTX 3090. The vLLM rows measure aggregate output across a fixed ladder on eight lower-memory GPUs. Those workloads are not interchangeable.

64successful concurrent requests reached by Qwen3.6 and GPT-OSS in the vLLM test ceiling.
8kcontext was enough to prevent Gemma4 from fitting on the 16GB-per-GPU worker.
2runtimes tested: FreeToken for hybrid MoE placement and vLLM for tensor-parallel serving.

Qwen3.8-Flash-Next: more than 80 decode tok/s on one RTX 5090

On 20 September, FreeToken served RadixArk/Qwen3.8-Flash-Next-NVFP4 on one 32 GB RTX 5090 in New Jersey. The worker used QSA sparse attention, NVFP4 Triton kernels, disk-backed expert loading, and automatic MoE offload. Its measured PCIe gather bandwidth was 51.92 GB/s.

Two warmed 1,023-token streams measured 80.918 median steady decode tok/s, 73.832 median end-to-end tok/s, and 1.226 seconds median time to first token. Three shorter 127-token streams measured 75.057 median steady decode tok/s.

At c8, all requests ran concurrently and produced 740 output tokens in 6.421 seconds: 115.249 aggregate output tok/s. Median steady decode was 32.428 tok/s/request, median time to first token was 3.061 seconds, and the 528 prompt tokens imply 172.5 effective aggregate prefill tok/s.

A separate c16 run raised the worker limit and CUDA graph ceiling to 16. All 16 requests were admitted with no queue and produced 1,543 output tokens in 12.537 seconds: 123.079 aggregate output tok/s. Median steady decode fell to 11.614 tok/s/request while aggregate output rose only 6.8% over c8. For this worker and workload, c8 is the practical operating point.

Evidence: the machine-readable benchmark summary includes the hardware profile, single-stream results, and c1/c2/c4/c8/c16 concurrency ladder.

GLM-5.3-Flash: independent workers and Serverless qualification

On 20 September, two independent FreeToken workers ran on separate 96 GB RTX PRO 6000 Blackwell Server Edition GPUs in Czechia. Requests were assigned round-robin, four per worker at c8 and eight per worker at c16. The c8 batch produced 1,134 output tokens in 28.316 seconds for 40.048 aggregate output tok/s. The c16 batch produced 2,229 output tokens in 34.293 seconds for 64.998 aggregate output tok/s. All 24 requests completed successfully.

At c8, median prefill was 10.080 seconds, observed prefill throughput was 2.587 prompt tok/s, and median steady decode was 10.214 tok/s/request. At c16, median prefill was 7.728 seconds, observed prefill throughput was 3.365 prompt tok/s, and median steady decode was 7.793 tok/s/request. This topology served separate sessions on each GPU; it did not use tensor parallelism.

On 30 August, FreeToken loaded LibertAIDAI/GLM-5.3-Flash-NVFP4 on one 32 GB RTX 5090 with hybrid expert offload. The qualified run measured 16.840 steady decode tok/s over 126 steps, 59.38 ms/token, 7.13 seconds warm TTFT, and 29.21 GiB peak server VRAM.

A separate RTX PRO 6000 Workstation Edition worker produced the exact requested semantic answer and completed every request in the concurrency test. At c8, 376 output tokens took 14.452 seconds for 26.017 aggregate output tok/s. At c16, 752 output tokens took 24.042 seconds for 31.279 aggregate output tok/s. The worker had eight resident request slots, so the c16 result includes queueing.

Vast Serverless also completed the lifecycle test on an RTX PRO 6000 S worker in the Netherlands. Cold loading took 262.69 seconds. Cached resume reached readiness in 109 seconds, inside the 300-second startup deadline, and the authenticated semantic probe returned the exact expected sentinel. The Serverless platform benchmark reported 20.572 measured workload units per second at concurrency four; that platform metric is separate from the output-token rates above.

Evidence: the 20 September dual-worker record, RTX 5090 qualification, earlier RTX PRO 6000 concurrency results, semantic result, and Serverless lifecycle record are published as machine-readable JSON. The implementation and validation history remain available in FreeToken pull request 292.

Full GLM-5.3: authenticated Serverless qualification

On 7 September, FreeToken 0.1.2 loaded the full 432.9 GiB LibertAIDAI/GLM-5.3-NVFP4 checkpoint on one H200 NVL and served it through an authenticated Vast Serverless endpoint. Autoscaled worker creation, model loading, and Serverless routing all completed successfully.

At c8, all 8 requests completed with 1,016 reported output tokens in 57.136 seconds: 17.782 aggregate output tok/s. At c16, all 16 completed with 2,032 tokens in 112.747 seconds: 18.023 aggregate output tok/s. The worker admitted at most eight active requests, so the second eight c16 requests queued. Nearly doubling wall time without a corresponding throughput gain shows that this worker saturated around c8 for this workload.

A separate semantic smoke test returned exactly 42 with a stop finish reason in 10.402 seconds. Every benchmark request used a 128-token output cap, so these results measure authenticated routing and throughput, not completed-answer quality.

Evidence: all 24 request results, the semantic smoke result, and the qualification record. Idle stop and cached restart remain unverified because the instance was deleted after the successful live tests. The model-support work remains tracked in FreeToken PR #406.

What the Kimi run proved

FreeToken loaded the public eight-layer, 0.40B MXFP4 Kimi-K3 development checkpoint and served it through the OpenAI-compatible API on a 24GB RTX 3090. A 256-token serial request produced 255 reported output tokens in 0.450 seconds, or 567.043 tokens/s. The first ladder completed all 32 simultaneous requests.

This proves the new KDA/MLA, latent-MoE, SiTU, MXFP4 loading, batching, and API path. It says nothing about how fast the full 2.8T Kimi-K3 model will run.

During a sustained follow-up, the live-swapped NVIDIA driver lost the passed-through GPU on the second 16-request sample. We kept the completed samples and recorded the failure. We did not turn the earlier burst peak into a sustained-throughput claim.

The full 1.56 TB checkpoint downloaded and all 96 resident-weight shards loaded on one 275040 MiB B300. Resident weights used 111134 MiB of GPU memory. FreeToken then reported that Kimi had no parallel expert-bank reader and used its serial build path. The API still returned 503 model is still loading after 5 minutes 52 seconds in the worker, so the spend guard stopped the run before generation. Throughput and maximum concurrency therefore remain unmeasured.

On 24 August, a separate 4× H200 qualification run used FreeToken's parallel expert-bank loader. The full, machine-readable record is versioned with the FreeToken change; the values below are included so this page remains independently auditable.

Cold-start milestoneMeasured time from launch
Resident weights loaded55 s
Expert banks loaded19 m 39 s
CUDA graph capture31 s
First successful /v1/models21 m 30 s
Concurrent requestsAggregate completion tok/sMedian TTFTBatch elapsed
10.64530.295 s48.047 s
20.97132.846 s63.845 s
41.33935.007 s92.637 s
81.82238.433 s136.126 s
162.40041.698 s206.627 s

Warmup used a 101-token prompt and generated 15 tokens: 34.803 s TTFT, 42.665 s elapsed, and 1.781 steady decode tok/s. Three subsequent 127-token streams had 30.350 s median TTFT, 1.211 median end-to-end tok/s, and 1.690 median steady decode tok/s. Every request in the 1/2/4/8/16-client ladder completed (31/31 total).

An earlier H100 run failed because FreeToken selected tensor parallel size 1 and the 106.55 GiB resident set could not fit in 79.18 GiB of usable GPU memory. A Canadian H200 attempt hit repeatable provider-volume I/O errors. Those failed instances were deleted with operator approval. The B300 instance is stopped and its 1.75 TB checkpoint volume is preserved for a follow-up after the serial startup path is fixed.

Review the FreeToken implementation and H200 evidence, the machine-readable H200 record, and the development-checkpoint artifacts.

What this changes for routing

Parameter count is not enough. Placement needs the weight format, active memory footprint, context limit, runtime, and expected concurrency.

A cheap community worker can be right for one workload and unable to load another. We keep raw results and failed runs so that decision can be checked later.

Qwen3.8-Flash-Next now has a measured RTX 5090 operating point. Qwen3.8 27B remains a separate candidate for a 24 GB RTX 3090.

Method notes

The vLLM runs used fixed 128-token completions, temperature 0, a fixed prompt seed, tensor parallelism across eight GPUs, and two trials at each concurrency level through 64. FreeToken requests streamed over the OpenAI-compatible API. The Kimi run ignored EOS to exercise the requested token budget. We saved metrics, not generated text or credentials. These are throughput tests, not quality scores or maximum-capacity claims.

Run a community provider or start with the API quickstart.