Strata vs Unsloth: Qwen3.8 Flash Next on an RTX 5090
I'm running Qwen3.8 Flash Next locally with Strata. On my RTX 5090 it generated 115.2 tokens/s, compared with 22.7 in Unsloth Studio, with a nearly full 256K context. Just over five times faster at generating the answer.
My machine has a Ryzen 9 9950X3D, RTX 5090 with 32 GB VRAM, and 128 GB RAM, running Omarchy. The model is OrcaRouter's Qwen3.8 Flash Next Uncensored, IQ3_XXS. Its 85.2 GB download needs more than GPU memory, so Strata uses system RAM and caches experts on the GPU.
I added a compatibility option to convert 460 small tensors to BF16 while keeping the expert weights quantized. The code and setup instructions are in my fork.
Both servers had 262,144 tokens of context. Strata used MTP speculation and INT8 KV streaming. Unsloth used its automatic GPU fitting, q8_0 KV and n-gram speculation, without a separate MTP draft.
I ran each engine alone on the GPU: three prompts per size, 256 output tokens, temperature zero and thinking off. The prompts ask for a hash-table explanation, with synthetic records filling the longer inputs. Servers stayed warm, with no reused input tokens; speculative lookup state could carry between requests.
At the full window, getting the first token still took 4m 06s in Strata and 10m 32s in Unsloth. Most of the wait is reading the prompt.
At 8K, the wait was 6.96 seconds vs 19.27.
Including the answer, that 8K request took 8.80 seconds in Strata and 23.40 in Unsloth.
I also wanted to check the answers. I ran 100 fixed GSM8K math questions and all 164 HumanEval coding problems, with medium thinking and an 8,192-token output limit on both. These used short prompts, one answer per problem, and no retries.
Both got 96/100 on math. Strata passed 162/164 coding problems, Unsloth 160/164. Strata passed three that Unsloth missed; Unsloth passed one where Strata spent all 8,192 tokens in a reasoning loop. Both failed one.
For long context, I planted a unique code at 10%, 50% and 90% of prompts containing 8,192, 32,768 and 258,040 tokens. Both had medium thinking and 4,096 output tokens to find it.
Strata found 9/9, Unsloth 8/9. Unsloth got stuck in a reasoning loop when the code was in the middle of the longest prompt. Strata returned the right code in 195 output tokens, including thinking.
I'm using Strata through OpenCode and OMP now. Both clients completed a real file-read tool call.
The timing data and settings, timing prompts, accuracy results, scores and accuracy prompts are here if you want to check the numbers.
comments