Strata has become one of the most interesting inference engines I have tested for running very large MoE models on workstation GPUs.
Over several days, I used Strata to run Qwen3.8-Flash-Next IQ3_S on a single NVIDIA RTX A4500 with 20GB of VRAM. This was not a simple single-chat benchmark. Strata was serving multiple long-lived OpenClaw agents, many with conversations well above 100K tokens, while simultaneously handling vision, KV streaming, conversation switching, and concurrent requests.
What began as an experiment in pushing Strata to a one-million-token context eventually became a more useful question:
How should Strata divide limited GPU memory between KV state, MoE experts, vision, speculative decoding, and parallel sessions?
The answer was not what I initially expected.
For my workload, the best Strata configuration was not the one with the largest context window or the largest expert cache. It was a native 262K-context configuration with parallel inference and a 128GB conversation cache, which eventually reused 88% of all prompt tokens.
My Strata Test System
The machine is a Dell Precision 5860 with an Intel Xeon w3-2435, 376GB of system RAM, and a single NVIDIA RTX A4500 20GB GPU.
The software stack used for the latest tests was:
OS: Ubuntu 22.04.5 LTS
GPU: NVIDIA RTX A4500 20GB
CUDA: 13.0
Driver: 580.173.02
Model:
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S
Inference engine:
Strata 0.1.39Strata 0.1.39 is particularly relevant to this workload because it introduces faster decode, changes to the long-prompt streamed-expert path, optimizations for very long attention contexts, and optional parallel request execution.
Strata Configuration Comparison
These four configurations turned out to be the most informative:
| Strata setup | Context | Vision | Parallel | GPU-resident experts | Host KV |
|---|---|---|---|---|---|
| 1M context test | 1,048,576 | GPU | 1 | ~5,174 | ~12.38 GiB |
| Native-context baseline | 262,144 | GPU | 1 | ~5,896 | ~3.09 GiB |
| CPU-vision test | 262,144 | CPU | 1 | ~6,542 | ~3.09 GiB |
| Current multi-agent setup | 262,144 | GPU | 2 | ~4,877 | ~3.09 GiB |
This table summarizes most of the tuning story.
Strata at 1M Context: It Worked
The first goal was ambitious: run Qwen3.8-Flash-Next with a 1,048,576-token context on a 20GB GPU.
Strata handled it.
The relevant configuration was:
--max-context 1048576
--kv int8
--kv-resident 32768Only 32,768 KV cells remained resident on the GPU, while the rest were streamed from pinned host memory.
At 1M context, Strata reported approximately:
Pinned KV memory: 12.38 GiB
Expert cache: 5,174 slots
Expert VRAM: 9.84 GiBTechnically, this was impressive.
Operationally, however, it was not optimal for me.
Most real OpenClaw conversations were in the 100K–250K range. Keeping the engine configured for a possible 1M-token request permanently reduced the space available for expert residency.
That led me back to the model’s native 262K context.
Why 262K Was Better for My Strata Workload
At:
--max-context 262144the pinned KV allocation dropped from approximately:
12.38 GiBto:
3.09 GiBThat immediately changed the memory balance.
The model no longer needed YaRN extension, the working context matched its native 262K training context, and considerably more expert capacity became available.
Strata 0.1.39 also contains a dedicated optimization for attention top-k in the 262K–524K-cell range. The release notes report a 26% improvement on a 243K-token prompt on an RTX 3060, making the 262K region particularly relevant for real long-context use.
For an agent system, I found that a large context window and persistent conversation state solve different problems.
A 1M context helps one enormous conversation.
A conversation cache helps many large conversations.
My workload needed the second much more often.
Understanding Strata Expert Cache Behavior
One of the more interesting parts of this experiment was tracking Strata’s GPU expert cache.
At one point I remembered seeing more than 6,500 resident experts, while later runs showed fewer than 5,900.
I initially suspected a Strata regression, YaRN, KV configuration, or the multi-conversation cache.
The actual explanation was much simpler.
It was vision.
How Much VRAM Does Strata GPU Vision Cost?
I performed a controlled A/B test, changing only Strata vision from GPU to CPU.
With GPU vision enabled:
strata-vision VRAM: ~1,248 MiB
Expert cache:
5,896 slots
11.20 GiBWith CPU vision:
strata-vision VRAM: 0
Expert cache:
6,542 slots
12.42 GiBThe difference matched almost perfectly:
Expert-cache difference:
12.42 - 11.20 ≈ 1.22 GiB
Vision GPU allocation:
~1.22 GiB
Expert difference:
6,542 - 5,896 = 646In other words, Strata was not losing experts. I had simply chosen to spend roughly 1.2GB of VRAM on fast vision instead.
This illustrates one of the most important aspects of Strata tuning: VRAM is a shared budget.
Memory allocated to the image encoder cannot simultaneously hold MoE experts.
I ultimately kept GPU vision because my OpenClaw workload frequently handles screenshots, browser images, and product imagery. The additional expert residency from CPU vision was useful, but the much slower image path was not worth it for this system.
Strata 0.1.39 Parallel Inference
The biggest practical change came with Strata 0.1.39.
Parallel request handling is enabled through:
"parallel": 2The Strata server converts this into:
--batch 2and maintains two conversation slots.
This feature is not intended to double single-GPU throughput. Its primary purpose is to reduce waiting time when several conversations compete for the same engine.
Strata’s own release notes describe the trade-off clearly: additional slots consume session memory that would otherwise be available to the expert cache, and on smaller cards parallel execution may reduce single-request performance even while improving queue latency.
That is exactly what I observed.
With GPU vision and no parallel execution:
~5,896 expertsWith GPU vision and:
"parallel": 2the cache dropped to roughly:
4,877 expertsThat is about a 17% reduction in resident experts.
Surprisingly, real decode performance did not drop nearly as much.
Strata Parallel=2 Was Actually Working
The logs made it clear that this was genuine concurrent execution rather than simple queued requests.
For example:
slot 1 takes 139730 tokens
slot 0 takes 117766 tokensBatch windows frequently reported:
avg 1.99 rows
avg 2.00 rows
avg 1.93 rowsand long prompts could yield while another slot continued decoding:
the prompt was read in 2 parts,
the slots decoding 2397 ms between themThis behavior matters enormously in a multi-agent system.
Without parallel execution, one long OpenClaw request can block every other session behind it.
With parallel=2, another active agent can continue making progress while a large prompt is being ingested.
Strata 0.1.39 was explicitly designed to allow a long prompt to yield at chunk boundaries while another slot continues.
Parallel Inference Does Not Mean Twice the Tokens per Second
This distinction is important.
If a single conversation produces 50 tokens/s, two parallel conversations do not automatically produce:
50 + 50 = 100 tokens/sThey still share one GPU.
The value of parallel=2 is responsiveness under contention.
Instead of:
Session A: running
Session B: waitingthe engine can operate more like:
Session A: running
Session B: runningEach may be slower while both are active, but the second request no longer has to wait until the first completes.
For interactive agent workloads, that can be more valuable than maximizing the benchmark speed of one request.
Where Strata Spends Time During Parallel Decode
The 0.1.39 batch profiler made the trade-off visible.
Typical two-slot windows showed approximately:
Total run: 38–46 ms
GPU-reach wait: 15–16 ms
CPU experts: 21–29 ms
PCIe: 1–2 msThe interesting part is that PCIe was not the dominant bottleneck.
The larger cost came from experts that were no longer resident in GPU memory and therefore had to execute on the CPU.
That makes sense: parallel=2 reduces expert residency to create room for another conversation slot.
Even so, the user-visible slowdown was smaller than the expert-count reduction suggested.
This is a good reminder that expert count alone is not a performance metric. Routing distribution and cache hit rate matter much more.
Strata KV Streaming at Long Context
KV streaming also performed well.
Across long-running sessions, I commonly saw:
96%–99.9%of KV block reads hit GPU-resident data.
Examples included:
96.09%
98.32%
98.86%
99.49%
99.88%Some log entries showed more than 14GB read from RAM, but those values represent cumulative KV traffic, not a single 14GB transfer.
The meaningful metric is the GPU hit rate.
For conversations well above 100K tokens, consistently staying in the high-90% range is a strong result.
Strata Conversation Cache Was the Biggest Win
The most valuable optimization in the entire setup may actually be Strata’s multi-conversation cache.
My current configuration is:
--conversation-cache-mib 131072
--conversation-cache-slots 64That gives Strata 128GB of host RAM for parked conversation state.
After sustained OpenClaw use, the monitor reported:
472 / 500 requests reused part of their prompt
50,772,040 tokens reused
88% of all prompt tokens reusedThe most recent request reused:
82,242 of 97,270 prompt tokensThe cache itself was nearly full:
62 / 64 parked conversations
125.8 / 128.0 GB RAM usedand had accumulated:
438 evictionsAt first glance, hundreds of evictions might look unhealthy.
They are not.
The more important statistic is:
88% prompt-token reuseThat means the cache is still keeping the active working set hot.
Instead of repeatedly pre-filling hundreds of thousands of tokens when switching between agents, Strata restores the parked conversation state and processes only the newly appended portion.
More than 50 million prompt tokens had already been avoided.
That benefit is far larger than chasing a few additional decode tokens per second.
Why I Have Not Increased the Strata Conversation Cache Further
The system has almost 376GB of RAM, so increasing the cache beyond 128GB would be possible.
But cache occupancy is not itself a problem.
Cache misses are.
At 88% token reuse, the current cache is still extremely effective.
The rest of the machine also needs memory for the expert arena, PLE, KV state, Linux file cache, OpenClaw, and normal system activity.
For now, 128GB and 64 slots appear to be a good operating point.
I would only increase it if prompt reuse started falling materially.
My Current Strata Configuration
The resulting setup looks approximately like this:
{
"parallel": 2,
"vision": {
"gpu": true,
"max_tokens": 1024
}
}with the main engine arguments:
--max-context 262144
--kv int8
--kv-resident 32768
--conversation-cache-mib 131072
--conversation-cache-slots 64
--ple-io ram
--vram-reserve-mib 700
--pcie-frac 0.20
--spec-min-p 0.70This produces a Strata configuration with:
Native 262K context
GPU vision
Two active inference slots
~4,877 GPU-resident experts
128GB multi-conversation cache
64 parked conversation slotsFor my workload, this has been substantially more useful than maximizing any single resource.
One Startup Issue I Hit with Strata
I also encountered one operational issue after upgrading.
A generated launcher used:
python serve/server.py ... --openand sometimes exited unexpectedly.
The stable form was:
.venv/bin/python -u -m serve.server \
--engine strata \
--config strata-iq3_s.json \
--host 10.91.133.47 \
--port 8085Since this Strata instance is used as an API backend, automatically opening the browser serves no purpose anyway.
Removing --open produced a more predictable server-style launcher.
One Problem That Was Not Strata
During the same testing period, OpenClaw image generation occasionally appeared to stall after a successful image generation.
The relevant message was:
Image generation completion wake failed after successful generationThe image already existed.
The failure happened when the asynchronous image task attempted to wake the original agent conversation.
Strata itself remained available and continued serving requests.
That distinction matters because agent orchestration failures can easily look like model-server failures when several layers are involved.
For image-heavy iterative workflows, I am now considering a synchronous ComfyUI tool so the agent can wait for image generation and continue within the same workflow, rather than relying on a later asynchronous wake.
What I Learned from Tuning Strata
The most useful lesson was that optimizing Strata is not about maximizing one number.
A 1M context is not automatically better than 262K.
6,500 resident experts are not automatically more useful than 4,877.
A single request running at maximum speed is not automatically better than two responsive conversations.
And a conversation cache that is almost completely full can still be working extremely well.
For a real multi-agent workload, the current balance is:
Strata at native 262K context, GPU vision, parallel=2, KV streaming, and a large host-side conversation cache.
The system sacrifices some GPU expert residency, but gains concurrent responsiveness and avoids enormous amounts of repeated prompt processing.
The most impressive number in the final setup is not the context length or even the decode speed.
It is this:
88% of all prompt tokens were reused.
For agent workloads, eliminating redundant computation may ultimately matter more than making each individual token a few percent faster.

Leave a Reply