Strata vs Unsloth: Qwen3.8 Flash Next on an RTX 5090


I'm running Qwen3.8 Flash Next locally with Strata. On my RTX 5090 it generated 115.2 tokens/s, compared with 22.7 in Unsloth Studio, with a nearly full 256K context. Just over five times faster at generating the answer.

Median generation speed at 127, 2048, 8192 and 261880 input tokens. Strata: 123.2, 135.6, 138.3, 115.2 tokens per second. Unsloth: 45.3, 45.0, 61.2, 22.7.
Median of three runs. Dots show individual runs. Each response contains 256 output tokens.

My machine has a Ryzen 9 9950X3D, RTX 5090 with 32 GB VRAM, and 128 GB RAM, running Omarchy. The model is OrcaRouter's Qwen3.8 Flash Next Uncensored, IQ3_XXS. Its 85.2 GB download needs more than GPU memory, so Strata uses system RAM and caches experts on the GPU.

I added a compatibility option to convert 460 small tensors to BF16 while keeping the expert weights quantized. The code and setup instructions are in my fork.

Both servers had 262,144 tokens of context. Strata used MTP speculation and INT8 KV streaming. Unsloth used its automatic GPU fitting, q8_0 KV and n-gram speculation, without a separate MTP draft.

I ran each engine alone on the GPU: three prompts per size, 256 output tokens, temperature zero and thinking off. The prompts ask for a hash-table explanation, with synthetic records filling the longer inputs. Servers stayed warm, with no reused input tokens; speculative lookup state could carry between requests.

At the full window, getting the first token still took 4m 06s in Strata and 10m 32s in Unsloth. Most of the wait is reading the prompt.

At 261,880 input tokens, median first-token latency is 246.11 seconds for Strata and 632.06 for Unsloth. Generation speed is 115.2 versus 22.7 tokens per second.
261,880 input tokens, with a 262,144-token window on both servers.

At 8K, the wait was 6.96 seconds vs 19.27.

Median time to first token at 127, 2048 and 8192 input tokens. Strata: 0.27, 1.73, 6.96 seconds. Unsloth: 0.90, 4.98, 19.27.
Time to first token. Rings mark medians; dots show individual runs.

Including the answer, that 8K request took 8.80 seconds in Strata and 23.40 in Unsloth.

Total time for 256 output tokens at 127, 2048 and 8192 input tokens. Strata: 2.32, 3.61, 8.80 seconds. Unsloth: 6.52, 10.65, 23.40.
Total time for 256 output tokens. Each bar shows the run with the median total time.

I also wanted to check the answers. I ran 100 fixed GSM8K math questions and all 164 HumanEval coding problems, with medium thinking and an 8,192-token output limit on both. These used short prompts, one answer per problem, and no retries.

Measured answer accuracy. Math: Strata 96 of 100; Unsloth 96 of 100. Code: Strata 162 of 164; Unsloth 160 of 164.
One answer per problem, scored against the math answer key or the official coding tests.

Both got 96/100 on math. Strata passed 162/164 coding problems, Unsloth 160/164. Strata passed three that Unsloth missed; Unsloth passed one where Strata spent all 8,192 tokens in a reasoning loop. Both failed one.

For long context, I planted a unique code at 10%, 50% and 90% of prompts containing 8,192, 32,768 and 258,040 tokens. Both had medium thinking and 4,096 output tokens to find it.

Exact code retrieval at 8192, 32768 and 258040 input tokens, with answers at 10, 50 and 90 percent depth. Strata passes 9 of nine cases; Unsloth passes 8 of nine.
The same prompt goes to both engines. PASS requires the exact hidden code.

Strata found 9/9, Unsloth 8/9. Unsloth got stuck in a reasoning loop when the code was in the middle of the longest prompt. Strata returned the right code in 195 output tokens, including thinking.

I'm using Strata through OpenCode and OMP now. Both clients completed a real file-read tool call.

The timing data and settings, timing prompts, accuracy results, scores and accuracy prompts are here if you want to check the numbers.

comments