VibeSys Builds a Qwen3.5-397B Engine on MI300A, 2.3× Tuned SGLang

Vic Shihang Li, Keisuke Kamahori, Simon Peter, Baris Kasikci

New ResearchInference & ServingAgents

October 8, 2026

TL;DR: We used VibeSys, our multi-agent system for designing systems, to build a serving engine for Qwen3.5-397B-A17B, a hybrid mixture-of-experts model, on four AMD MI300As. No human wrote engine code. In 105 hours, the agents took a baseline PyTorch engine from 12.8 to 2,242 tok/s of goodput (throughput of responses that meet the latency SLO) on a multi-turn chat workload, 2.33× SGLang on the same hardware. Stock SGLang does not run this model on MI300A, so we first used VibeSys to port and tune it.

Goodput of the agent-built engine across rounds.

Figure 1. Best goodput so far (blue) across optimization rounds on 4× MI300A. Gray dots are individual measured runs, including regressions and rejected candidates. The dashed line is tuned SGLang (961 tok/s).

The VibeServe paper argues that coding agents make per-deployment specialization of serving engines affordable, mostly with small models on one GPU. This deployment tests the claim at scale. Over 105 hours, our agents iteratively built model, hardware, and workload aware optimizations, resulting in a bespoke serving engine that reaches 2.33x SGLang in goodput.

The target

Model. Qwen3.5-397B-A17B activates about 17B of its 397B parameters per token, across 512 routed experts with MXFP4 weights. Of its 60 layers, 15 use full attention with a KV cache and 45 are Gated DeltaNet layers, a form of linear attention with a fixed-size recurrent state updated in place.

Hardware. Each MI300A puts 24 CPU cores and GPU compute on one package with a shared 128 GB pool of HBM3, so weights, GPU caches, the OS page cache and host processes all compete for the same memory. Four of them give the node 512 GB.

Workload. A deterministic multi-turn chat benchmark ramps from 1 to 192 concurrent sessions. Each session has 3–6 turns of 80–300 generated tokens, with 1–8 s of think time, and each turn resends the full conversation. A run fails if any response is dropped, empty or truncated. The SLO is p95 time per output token (TPOT) ≤ 250 ms and p95 time to first token (TTFT) ≤ 10 s. We report each system’s peak goodput over the ramp.

How VibeSys ran

We drive the optimizations with VibeSys, a multi-agent framework for building systems created from the VibeServe prototype.

The agents received the model weights, the Hugging Face Transformers reference implementation, the benchmark, and a natural-language description of the hardware and the deliverable, an HTTP server. The implementation agents were allowed to use PyTorch, Triton and individual kernels from AMD’s AITER library, but no serving code from SGLang, vLLM or TensorRT-LLM. A review agent checked each change against this rule. As a result, the agents wrote the scheduler, caches, parallelism, graph capture and several kernels themselves.

An orchestrator agent proposes optimization hypotheses and dispatches each to an implementation agent, which calls benchmarking and profiling agents. VibeSys ties every result to a commit, so the orchestrator sees the trajectory across hypotheses. Each round integrates the accepted hypotheses. Independent hypotheses run in parallel: the 105-hour run used 392.8 hours of agent time. Every agent was Claude Code with Opus 5.

A hypothesis is accepted if it improves goodput on a 20% held-out set and passes the correctness checks: at least 1,074 of 1,088 teacher-forced top-1 tokens must match the reference, and 5 of 5 greedy reference completions must be reproduced. Nodes varied by about ±10%, more than most single optimizations, so acceptance also required at least two same-node pairs, with control and candidate run back to back on one node.

The baseline: using VibeSys to port and tune SGLang

SGLang v0.5.18 does not serve this model on MI300A, for three reasons:

  • No MoE kernel. Every MXFP4 kernel in the AITER version SGLang pins is restricted to gfx950 (MI350-class), so on the MI300A’s gfx942 the server aborts during warmup. The generic Triton fallback has no tuned configuration for 512 experts on this GPU and reaches 1–3% of HBM bandwidth.
  • Memory sizing. SGLang sizes its KV and recurrent-state pools from the free memory it observes after loading weights. On shared memory, loading through the page cache leaves the four ranks with uneven free memory, so boots fail with out-of-memory errors, a negative state budget, or a stalled node.
  • Boot time. Loading weights takes about 60 minutes. AITER then compiles an attention kernel on the first request (about 105 s), and SGLang’s 20 s health check kills the server during the compile.

VibeSys supports optimizing well-optimized serving engines like SGLang and vLLM, in addition to starting from the Hugging Face implementation. To make this model run on MI300A, we first use VibeSys to port SGLang. Agents wrote a fused HIP MoE kernel that dequantizes MXFP4 weights in registers, pinned the KV pool size, replaced memory-mapped loading with bounded reads of pre-sharded per-rank weights (load time fell to about 1 minute), and shipped a prebuilt AITER kernel cache.

VibeSys agents then tuned the running server. At 48 concurrent sessions, the largest wins were speculative decoding with the model’s built-in multi-token-prediction (MTP) head and 4 draft tokens (median TPOT 70.9 → 38.3 ms) and GEMM tuning tables for every captured decode shape (38.1 → 22.9 ms). Smaller wins came from three MoE kernel refinements, padded and tuned prefill GEMMs, and disabling overlap scheduling. Tuning raised goodput by 42%, to 961 tok/s at 48 sessions. In total, the agents added about 3,500 lines to SGLang, about 1,900 of them for the MoE kernel, a small-batch GEMM kernel and their wrappers.

Results

The agent-built engine reaches 2,242 tok/s at 96 concurrent sessions. This is 2.33× the peak throughput of the SGLang baseline ported and tuned by VibeSys. Both engines use MTP. Tuned SGLang peaks at 48 sessions. In a same-node run at 96 sessions, its p95 TTFT reached 15.5 s.

The agent-built engine has about 19,400 lines of code (excluding comments and tests), including 520 lines of HIP. Unlike the SGLang changes, it includes its own scheduler, caches, parallelism and MTP.

The agents optimized peak goodput, which rewards throughput at high load and leaves low-load latency unoptimized. At low concurrency, tuned SGLang is still faster, and at 16 sessions MTP raises the engine’s later-turn p95 TTFT by 11–15%. An objective that includes low-load latency, such as turn-2+ TTFT, would direct optimization there.

Where the speedup came from

The trajectory in Figure 1 has four phases. Gains in the last phase are measured against the previous build on the same node.

Prefix caching (12.8 → 55.6 tok/s)

The first engine prefilled the full conversation on every turn, and serialized prefill starved decode. A DeltaNet layer can resume only from a saved snapshot of its recurrent state, so the agents built a prefix cache that saves both attention KV and DeltaNet snapshots. It gave 4.3×.

Kernels and graph capture (55.6 → 659 tok/s)

Each decode step launched many small kernels, and the GPU waited on the CPU. The agents captured decode, prefill and mixed decode-plus-prefill steps as HIP graphs, one launch per step. They also wrote a fused MXFP4 MoE kernel, chunked DeltaNet prefill and split-K decode attention, and added an overlap scheduler that plans the next step on the CPU while the GPU runs the current one. This phase gave the largest gain, 11.9×.

Communication and concurrency (659 → 1,055 tok/s)

Splitting the model across four ranks costs two collectives per layer, about 120 per decode step. The agents replaced the stock all-reduce with a one-shot version fused with the residual add and RMSNorm. Larger decode graphs, up to batch 128, let the engine run 96 concurrent sessions, where it passed SGLang’s 961 tok/s.

Prefill, caching and speculation (1,055 → 2,242 tok/s)

Snapshot eviction (+27%). Each DeltaNet snapshot costs about 47 MiB, so the prefix cache holds far fewer conversations than it would for an attention-only model. The cache is a tree of past conversations, and it stored a snapshot at every branch point. Once a later turn extends a conversation past a branch point, no request resumes from that snapshot. These stale snapshots filled the shared memory budget, so the cache evicted active conversations, whose next turns then recomputed the full history. Evicting stale interior snapshots first, while keeping the snapshot at the end of each active conversation, cut prefilled tokens by about 40%.

MTP speculative decoding (+16–20%). The main model verifies the MTP head’s drafted tokens in one pass. With DeltaNet layers, verification must run the recurrence and causal convolution over the drafts and roll back the recurrent state on rejection, all under graph capture. The agents first closed MTP as noncompetitive because the captured version hit an illegal memory access on its first replay. A later round found the cause: the graph referenced tensors owned only by a Python closure, which were freed after capture. With 2 draft tokens, each sequence emits 2.68 tokens per verify step, and 84% of drafts are accepted.

Cheaper prefill and collectives. Accumulating waiting prefills into chunks of up to 1,536 tokens gave +9.3%, packing prefill sequences without padding +5.6%, and 32-row tiles in the MoE kernel +4.8%. Restructuring the all-reduce as reduce-scatter, RMSNorm and all-gather gave +8–9%.

What did not help

A BF16 cache of frequently used experts covered 80–86% of expert assignments but ran 5–7× slower than the MXFP4 kernel and used up to 20 GiB per rank. A custom dense GEMM in place of tuned hipBLASLt, a persistent MoE kernel, and parallel routing preparation were each neutral or slower. The agents discarded all four.

VibeSys is available at github.com/uw-syfi/vibesys, and the paper describes the full design. The agent-built engine and the tuned SGLang are public.