Strata LLM Inference on a 20GB RTX A4500: 262K Context, Parallel Agents, and 88% Prompt Reuse


Strata has become one of the most interesting inference engines I have tested for running very large MoE models on workstation GPUs.

Over several days, I used Strata to run Qwen3.8-Flash-Next IQ3_S on a single NVIDIA RTX A4500 with 20GB of VRAM. This was not a simple single-chat benchmark. Strata was serving multiple long-lived OpenClaw agents, many with conversations well above 100K tokens, while simultaneously handling vision, KV streaming, conversation switching, and concurrent requests.

What began as an experiment in pushing Strata to a one-million-token context eventually became a more useful question:

How should Strata divide limited GPU memory between KV state, MoE experts, vision, speculative decoding, and parallel sessions?

The answer was not what I initially expected.

For my workload, the best Strata configuration was not the one with the largest context window or the largest expert cache. It was a native 262K-context configuration with parallel inference and a 128GB conversation cache, which eventually reused 88% of all prompt tokens.

My Strata Test System

The machine is a Dell Precision 5860 with an Intel Xeon w3-2435, 376GB of system RAM, and a single NVIDIA RTX A4500 20GB GPU.

The software stack used for the latest tests was:

OS:        Ubuntu 22.04.5 LTS
GPU:       NVIDIA RTX A4500 20GB
CUDA:      13.0
Driver:    580.173.02

Model:
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S

Inference engine:
Strata 0.1.39

Strata 0.1.39 is particularly relevant to this workload because it introduces faster decode, changes to the long-prompt streamed-expert path, optimizations for very long attention contexts, and optional parallel request execution.

Strata Configuration Comparison

These four configurations turned out to be the most informative:

Strata setupContextVisionParallelGPU-resident expertsHost KV
1M context test1,048,576GPU1~5,174~12.38 GiB
Native-context baseline262,144GPU1~5,896~3.09 GiB
CPU-vision test262,144CPU1~6,542~3.09 GiB
Current multi-agent setup262,144GPU2~4,877~3.09 GiB

This table summarizes most of the tuning story.

Strata at 1M Context: It Worked

The first goal was ambitious: run Qwen3.8-Flash-Next with a 1,048,576-token context on a 20GB GPU.

Strata handled it.

The relevant configuration was:

--max-context 1048576
--kv int8
--kv-resident 32768

Only 32,768 KV cells remained resident on the GPU, while the rest were streamed from pinned host memory.

At 1M context, Strata reported approximately:

Pinned KV memory:  12.38 GiB
Expert cache:      5,174 slots
Expert VRAM:       9.84 GiB

Technically, this was impressive.

Operationally, however, it was not optimal for me.

Most real OpenClaw conversations were in the 100K–250K range. Keeping the engine configured for a possible 1M-token request permanently reduced the space available for expert residency.

That led me back to the model’s native 262K context.

Why 262K Was Better for My Strata Workload

At:

--max-context 262144

the pinned KV allocation dropped from approximately:

12.38 GiB

to:

3.09 GiB

That immediately changed the memory balance.

The model no longer needed YaRN extension, the working context matched its native 262K training context, and considerably more expert capacity became available.

Strata 0.1.39 also contains a dedicated optimization for attention top-k in the 262K–524K-cell range. The release notes report a 26% improvement on a 243K-token prompt on an RTX 3060, making the 262K region particularly relevant for real long-context use.

For an agent system, I found that a large context window and persistent conversation state solve different problems.

A 1M context helps one enormous conversation.

A conversation cache helps many large conversations.

My workload needed the second much more often.

Understanding Strata Expert Cache Behavior

One of the more interesting parts of this experiment was tracking Strata’s GPU expert cache.

At one point I remembered seeing more than 6,500 resident experts, while later runs showed fewer than 5,900.

I initially suspected a Strata regression, YaRN, KV configuration, or the multi-conversation cache.

The actual explanation was much simpler.

It was vision.

How Much VRAM Does Strata GPU Vision Cost?

I performed a controlled A/B test, changing only Strata vision from GPU to CPU.

With GPU vision enabled:

strata-vision VRAM: ~1,248 MiB

Expert cache:
5,896 slots
11.20 GiB

With CPU vision:

strata-vision VRAM: 0

Expert cache:
6,542 slots
12.42 GiB

The difference matched almost perfectly:

Expert-cache difference:
12.42 - 11.20 ≈ 1.22 GiB

Vision GPU allocation:
~1.22 GiB

Expert difference:
6,542 - 5,896 = 646

In other words, Strata was not losing experts. I had simply chosen to spend roughly 1.2GB of VRAM on fast vision instead.

This illustrates one of the most important aspects of Strata tuning: VRAM is a shared budget.

Memory allocated to the image encoder cannot simultaneously hold MoE experts.

I ultimately kept GPU vision because my OpenClaw workload frequently handles screenshots, browser images, and product imagery. The additional expert residency from CPU vision was useful, but the much slower image path was not worth it for this system.

Strata 0.1.39 Parallel Inference

The biggest practical change came with Strata 0.1.39.

Parallel request handling is enabled through:

"parallel": 2

The Strata server converts this into:

--batch 2

and maintains two conversation slots.

This feature is not intended to double single-GPU throughput. Its primary purpose is to reduce waiting time when several conversations compete for the same engine.

Strata’s own release notes describe the trade-off clearly: additional slots consume session memory that would otherwise be available to the expert cache, and on smaller cards parallel execution may reduce single-request performance even while improving queue latency.

That is exactly what I observed.

With GPU vision and no parallel execution:

~5,896 experts

With GPU vision and:

"parallel": 2

the cache dropped to roughly:

4,877 experts

That is about a 17% reduction in resident experts.

Surprisingly, real decode performance did not drop nearly as much.

Strata Parallel=2 Was Actually Working

The logs made it clear that this was genuine concurrent execution rather than simple queued requests.

For example:

slot 1 takes 139730 tokens
slot 0 takes 117766 tokens

Batch windows frequently reported:

avg 1.99 rows
avg 2.00 rows
avg 1.93 rows

and long prompts could yield while another slot continued decoding:

the prompt was read in 2 parts,
the slots decoding 2397 ms between them

This behavior matters enormously in a multi-agent system.

Without parallel execution, one long OpenClaw request can block every other session behind it.

With parallel=2, another active agent can continue making progress while a large prompt is being ingested.

Strata 0.1.39 was explicitly designed to allow a long prompt to yield at chunk boundaries while another slot continues.

Parallel Inference Does Not Mean Twice the Tokens per Second

This distinction is important.

If a single conversation produces 50 tokens/s, two parallel conversations do not automatically produce:

50 + 50 = 100 tokens/s

They still share one GPU.

The value of parallel=2 is responsiveness under contention.

Instead of:

Session A: running
Session B: waiting

the engine can operate more like:

Session A: running
Session B: running

Each may be slower while both are active, but the second request no longer has to wait until the first completes.

For interactive agent workloads, that can be more valuable than maximizing the benchmark speed of one request.

Where Strata Spends Time During Parallel Decode

The 0.1.39 batch profiler made the trade-off visible.

Typical two-slot windows showed approximately:

Total run:       38–46 ms
GPU-reach wait:  15–16 ms
CPU experts:     21–29 ms
PCIe:            1–2 ms

The interesting part is that PCIe was not the dominant bottleneck.

The larger cost came from experts that were no longer resident in GPU memory and therefore had to execute on the CPU.

That makes sense: parallel=2 reduces expert residency to create room for another conversation slot.

Even so, the user-visible slowdown was smaller than the expert-count reduction suggested.

This is a good reminder that expert count alone is not a performance metric. Routing distribution and cache hit rate matter much more.

Strata KV Streaming at Long Context

KV streaming also performed well.

Across long-running sessions, I commonly saw:

96%–99.9%

of KV block reads hit GPU-resident data.

Examples included:

96.09%
98.32%
98.86%
99.49%
99.88%

Some log entries showed more than 14GB read from RAM, but those values represent cumulative KV traffic, not a single 14GB transfer.

The meaningful metric is the GPU hit rate.

For conversations well above 100K tokens, consistently staying in the high-90% range is a strong result.

Strata Conversation Cache Was the Biggest Win

The most valuable optimization in the entire setup may actually be Strata’s multi-conversation cache.

My current configuration is:

--conversation-cache-mib 131072
--conversation-cache-slots 64

That gives Strata 128GB of host RAM for parked conversation state.

After sustained OpenClaw use, the monitor reported:

472 / 500 requests reused part of their prompt

50,772,040 tokens reused

88% of all prompt tokens reused

The most recent request reused:

82,242 of 97,270 prompt tokens

The cache itself was nearly full:

62 / 64 parked conversations
125.8 / 128.0 GB RAM used

and had accumulated:

438 evictions

At first glance, hundreds of evictions might look unhealthy.

They are not.

The more important statistic is:

88% prompt-token reuse

That means the cache is still keeping the active working set hot.

Instead of repeatedly pre-filling hundreds of thousands of tokens when switching between agents, Strata restores the parked conversation state and processes only the newly appended portion.

More than 50 million prompt tokens had already been avoided.

That benefit is far larger than chasing a few additional decode tokens per second.

Why I Have Not Increased the Strata Conversation Cache Further

The system has almost 376GB of RAM, so increasing the cache beyond 128GB would be possible.

But cache occupancy is not itself a problem.

Cache misses are.

At 88% token reuse, the current cache is still extremely effective.

The rest of the machine also needs memory for the expert arena, PLE, KV state, Linux file cache, OpenClaw, and normal system activity.

For now, 128GB and 64 slots appear to be a good operating point.

I would only increase it if prompt reuse started falling materially.

My Current Strata Configuration

The resulting setup looks approximately like this:

{
  "parallel": 2,
  "vision": {
    "gpu": true,
    "max_tokens": 1024
  }
}

with the main engine arguments:

--max-context 262144

--kv int8
--kv-resident 32768

--conversation-cache-mib 131072
--conversation-cache-slots 64

--ple-io ram

--vram-reserve-mib 700
--pcie-frac 0.20
--spec-min-p 0.70

This produces a Strata configuration with:

Native 262K context
GPU vision
Two active inference slots
~4,877 GPU-resident experts
128GB multi-conversation cache
64 parked conversation slots

For my workload, this has been substantially more useful than maximizing any single resource.

One Startup Issue I Hit with Strata

I also encountered one operational issue after upgrading.

A generated launcher used:

python serve/server.py ... --open

and sometimes exited unexpectedly.

The stable form was:

.venv/bin/python -u -m serve.server \
  --engine strata \
  --config strata-iq3_s.json \
  --host 10.91.133.47 \
  --port 8085

Since this Strata instance is used as an API backend, automatically opening the browser serves no purpose anyway.

Removing --open produced a more predictable server-style launcher.

One Problem That Was Not Strata

During the same testing period, OpenClaw image generation occasionally appeared to stall after a successful image generation.

The relevant message was:

Image generation completion wake failed after successful generation

The image already existed.

The failure happened when the asynchronous image task attempted to wake the original agent conversation.

Strata itself remained available and continued serving requests.

That distinction matters because agent orchestration failures can easily look like model-server failures when several layers are involved.

For image-heavy iterative workflows, I am now considering a synchronous ComfyUI tool so the agent can wait for image generation and continue within the same workflow, rather than relying on a later asynchronous wake.

What I Learned from Tuning Strata

The most useful lesson was that optimizing Strata is not about maximizing one number.

A 1M context is not automatically better than 262K.

6,500 resident experts are not automatically more useful than 4,877.

A single request running at maximum speed is not automatically better than two responsive conversations.

And a conversation cache that is almost completely full can still be working extremely well.

For a real multi-agent workload, the current balance is:

Strata at native 262K context, GPU vision, parallel=2, KV streaming, and a large host-side conversation cache.

The system sacrifices some GPU expert residency, but gains concurrent responsiveness and avoids enormous amounts of repeated prompt processing.

The most impressive number in the final setup is not the context length or even the decode speed.

It is this:

88% of all prompt tokens were reused.

For agent workloads, eliminating redundant computation may ultimately matter more than making each individual token a few percent faster.



Posted

in

,

by

Tags:

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *