Inference & PerformanceAI Coding

100k Tokens In, 500 Tokens Out: Three Steps Modal Took to Speed Up Kimi K2.6 Coding Agent Inference

When programming with coding agents, we often run into an awkward dynamic: ask it to fix a single bug, and it has to read through the entire project codebase, including terminal error logs and prior modifications. It swallows over a hundred thousand tokens, only to spit out a patch of just a few lines. Such requests—massively long inputs, tiny outputs, stacked across multi-turn interactions—are brutally unfriendly to conventional model-serving architectures.

On September 23, 2026 (2026-09-23), five Modal engineers—Janelle Cai, Charles Frye, James Liu, Timothy Feng, and Richard Gong—published an engineering write-up. Tackling Moonshot AI’s trillion-parameter model Kimi K2.6 on an NVIDIA Blackwell B200 cluster, they boosted single-user interactive generation speed to 2.8x its original baseline and pushed aggregate multi-user replica throughput to 5.6x.

Retrospectives like this easily read like a laundry list of buzzwords. In reality, every single change was forced by this extreme workload shape. Understanding why the team chose these specific levers, what tradeoffs they accepted, and under what conditions those speedup multiples actually hold is the key to understanding how this engineering battle was won.

Swallowing 100k Tokens and Emitting 500 per Turn: A Workload That Breaks Default Serving

Using a coding agent to fix a bug in daily practice routinely looks like this: the patch itself might only be a few lines of code, but the agent has to ingest the entire project directory tree, error logs, and previous modifications. Profiling their production traffic, Modal found that a typical request averages around 100k input tokens while generating ~500 output tokens—an input-to-output ratio of 200:1.

Moreover, real-world coding tasks are rarely resolved in a single shot; they routinely span dozens of interaction turns. By turn T, the model has to ingest the entire conversational history and codebase state from turns 1 through T-1 all over again, along with the latest prompt. Those few lines of background context the user typed at the very beginning end up being repeatedly hauled into VRAM and recomputed across dozens of subsequent calls.

This demand presents a brutal challenge to serving architectures. Running an unoptimized open-source solution off the shelf, once single-node load exceeds 6 users, both overall machine throughput and per-user generation speed plummet to rock bottom. Long before hardware compute is fully utilized, frontend interactive responsiveness already grinds to a halt.

These numbers conceal three tightly interconnected bottlenecks. For every token generated, the hardware has to sweep gigabytes of weights and historical state from VRAM into the compute cores. When memory transfers cannot keep pace with compute arithmetic, single-request latency runs straight into a memory bandwidth wall. To avoid redundant recomputation, the system must preserve preceding input states in a KV cache; but multi-turn dialogue causes that cache to snowball, rapidly exhausting VRAM capacity and capping single-node concurrency. And because single-node concurrency cannot be scaled up, cluster compute capacity remains starved, leaving aggregate throughput languishing at a low level.

Faced with this chain reaction, the team chose to tackle single-request interactive speed first before driving up concurrent throughput. Users writing code are acutely sensitive to token streaming speed; boosting single-request generation efficiency provides a foundation where subsequent system scaling can compound returns. The 2.8x single-user speedup and 5.6x concurrent throughput they ultimately achieved trace their origins directly to this decision. Guided by this thinking, the first bottleneck the team set out to resolve was the memory bandwidth constraint incurred with every data transfer during the decoding phase.

A coding agent request averages 100k input tokens and only 500 output tokens, with performance pressure cascading across bandwidth, capacity, and throughput.

Emitting More Tokens per Step: Fixing Single-Request Latency First

During the decoding phase, generating each token requires sweeping gigabytes of weights and cache entries out of VRAM into the compute cores—a transfer that takes far longer than the actual arithmetic. Pure parallelization can create more memory bandwidth channels, but it cannot overcome the strict token-by-token serial dependency. Modal took a different tack: since memory transfers are unavoidable, make each transfer yield more tokens. Have a blazingly fast small draft model guess several tokens ahead, then let the target model compare and verify them in parallel during a single forward pass. As long as guesses are accurate, reading weights once yields multiple valid tokens in succession. This guess-first, verify-later technique is speculative decoding.

Within the SGLang ecosystem, the team adopted DFlash. This approach is a close cousin of Multi-Token Prediction (MTP): both leverage intermediate states computed mid-flight by the target model to guess upcoming tokens. But while native MTP modules built into recent models still emit speculative tokens sequentially one by one, DFlash lets a lightweight draft model generate an entire block of candidate tokens in parallel, which is far better suited to GPUs. It takes the target model’s previously computed intermediate features stored in the KV cache as speculative inputs. The final output is still determined entirely by the target model’s probability distribution, ensuring it never alters the original model’s output distribution.

Getting the small speculative draft model to guess accurately hinges entirely on feeding it the right training data. The team’s approach was to pretrain on general corpora first, then fine-tune on traces generated by the target model while executing real-world coding tasks. Once in production, new trajectories from the live environment continuously feed back, providing an ongoing stream of fresh training samples.

The test of whether speculative decoding works comes down to how many tokens the target model accepts per step on average—the acceptance length. In benchmark evaluation tasks, the untuned base speculative model achieved an average acceptance length of 5.00 tokens; after fine-tuning on programming trajectories, the acceptance length climbed to 5.84 tokens, delivering a full 20% extra speedup to the decoding phase.

Different workloads respond to speculative decoding in vastly different ways. Comparing workloads under the identical mechanism, because code exhibits highly standardized syntax and repetitive patterns, coding tasks yielded an average acceptance length roughly 2x that of standard prose tests. In the team’s earlier explorations on Qwen 3.5/3.6 models, speculative decoding had delivered multi-fold integer speedups.

During engineering implementation, however, the team discovered a tokenization conversion pitfall. Many system logs record human-readable text strings, but converting tokens to text and back to tokens is not an invertible mapping; as a result, speculative context prediction could introduce subtle errors. Modal submitted an upstream fix to SGLang that directly exposes raw token IDs in the underlying extension, merged in PR #34488.

Single-request latency had indeed improved. But if a single machine can only accommodate a handful of users, concurrent capacity remains bottlenecked. The next challenge—how to carve out cache headroom for more users within limited VRAM—became the core problem of hardware partitioning and resource management.

Reclaiming VRAM for the Cache: Carving Out Concurrency on a Single Node

To run a model of this magnitude, the first architectural fork in the road was how to partition the weights across multiple accelerator cards. Kimi K2.6 sits at around 1 trillion parameters; even in 4-bit microscopic scaling floating point (NVFP4), its weights still occupy 595 GB of space. Given that each B200 accelerator card has 180 GB of VRAM capacity, deployment is clearly impossible across 1-2 cards. The team therefore had to decide, within the constraints of single-node hardware, between choosing 4 cards or 8 cards—evaluating TP4/TP8.

The VRAM capacity gap between these two options is massive. A straightforward back-of-the-envelope calculation illustrates why: the model has 61 layers; in 16-bit half-precision BF16 format, assuming 576-dim per layer and 2 bytes per element, the historical state for a single token requires 72 KB of space. Under a 4-card setup, the entire machine offers 720 GB of VRAM; after subtracting the 595 GB model weight footprint, only 125 GB of VRAM remains. Divided across 4 cards, each card receives only about 31.3 GB of VRAM, meaning the aggregate KV cache across the whole instance can only hold ~450k tokens.

Opting for an 8-card setup, however, shoots total VRAM up to 1440 GB; deducting weights leaves 845 GB, giving each card 105.6 GB. Total cache capacity can surge to ~3M tokens, exactly 6x that of the 4-card layout. Looking purely from a VRAM capacity standpoint, the 8-card option appears distinctly superior.

Despite the tempting 6x cache headroom, the team ultimately chose 4 cards. The clinching factor came from head-to-head empirical benchmarks: running prompt prefill alone on the 8-card configuration barely matched the per-card throughput of the 4-card configuration running both prefill and decode combined. Research by SemiAnalysis in their InferenceX benchmarks on the same NVFP4 model reinforced this conclusion: 4-card or 8-card configurations extract maximum interactive speed while preserving per-card efficiency; attempting to fracture attention state across data parallelism simply to manufacture concurrency inflates latency and diminishes overall per-card yield.

Operational considerations similarly favored smaller clusters: an 8-card node can be managed by a single host, making on-demand procurement and scooping up cheap spot instances straightforward; a massive 72-card enclosure requires nine subsystems crammed into a single address space, plagued by tight supply and long contract commitments.

The 4-card layout preserved per-card efficiency, but at the cost of razor-thin cache headroom: a single instance could only sustain long sessions for roughly 4 users. Past that threshold, subsequent requests miss existing state; cache hit rates plummet off a cliff, and inputs spanning over a hundred thousand tokens must be recomputed from scratch. To carve out cache space against these tight walls, the team executed three layers of VRAM optimization.

The first layer cleared out VRAM redundancy in the codebase. In DFlash’s initial implementation, intermediate feature tensors used for speculative drafting were first collected into a list of pointers and only copied into contiguous memory after the forward pass completed, driving peak runtime VRAM usage to 2x actual requirements. The team refactored this to preallocate contiguous buffers and write into them directly during execution, shedding the secondary copy overhead and halving peak VRAM usage during this phase. The patch was merged in PR #28956.

The second layer reduced storage precision to slim down memory volume. The team compressed the KV cache from BF16 down to FP8, immediately doubling cache capacity; simultaneously, they quantized shared expert parameters invoked on every token from FP8 down to NVFP4. Benchmark data showed that output variance from both modifications fell entirely within the intrinsic run-to-run stochastic error of the model. The speculative draft model was also quantized to FP8, with patches merged in PR #28957. Operating under the safeguard of rigorous benchmarking before executing lossy optimizations is precisely where a bespoke hosted infrastructure gains its core edge over generic multi-tenant platforms.

The third layer built hierarchical storage as a safety net. No matter how much VRAM you carve out, super-long contexts will eventually overflow; when they do, the system offloads the spillover into host RAM and out to distributed storage. Reading from host memory is slower than VRAM, but orders of magnitude faster than recomputing hundreds of thousands of tokens from scratch. This multi-tiered caching is implemented natively via SGLang’s built-in HiCache: host RAM serves as L2 cushioning for sudden bursts, while distributed storage acts as an L3 backstop, turning what would have been a cliff-edge plunge in cache hit rate under concurrency into a smooth, gradual transition.

Stacking these optimizations together, compared against the unoptimized single-node baseline, single-user interactive generation speed reached 2.8x, and single-instance multi-user concurrent throughput advanced to 5.6x. Note that this is the cumulative outcome of multiple interlocking changes, achieved specifically on this distinct model, hardware configuration, and coding workload.

Routing Requests to the Right Cache: Don’t Let the Dispatcher Force Cold Recomputations

Once single-node performance was dialed in, the real test emerged with multi-node cluster scaling. As soon as the team expanded nodes horizontally, anomalous disruptions surfaced on the dashboards: request queues became volatile, latency from request dispatch to receiving the first token showed jitter, and TTFT (Time To First Token) deteriorated. Furthermore, aggregate cluster throughput failed to scale linearly with single-node throughput.

Upon analysis and debugging, the team traced the root cause back to cold starts triggered by scattered request routing. Concretely, when a client dispatched a request in its teens of turns, the input prefix history it provided had already been processed by the cluster in earlier turns. But due to the routing mechanism, this request was assigned to a machine lacking the corresponding historical cache, forcing that fresh machine to perform a full prefill encoding calculation from scratch. The code was not buggy; this was the inescapable mathematical fate of traditional stateless hash routing.

The system initially used consistent hashing on client-supplied session IDs to map each session to a fixed machine. Yet relying purely on random hash dispatching is like ushering guests randomly into a row of dining rooms: some rooms sit empty while others are packed beyond standing room. This is not bad luck; it is the mathematical essence of random allocation. When assigning 250 sessions across 50 machines, even with an expected average of 5 per machine, random distribution variance remains inevitable. Regardless of how many machines exist, at any given moment ~4% (roughly 4%) of the machines will receive only 1 or 0 sessions, idling hardware; meanwhile, about 0.5% of the machines will catch 12+ sessions, immediately hitting severe tail latencies. Neither percentage diminishes as the cluster expands. In production, reality is worse: long-session requests routinely span hundreds of thousands of tokens, taking longer to process and lingering on nodes far longer. Empirically, the degree to which node load deviated from the average frequently reached multiples of the mean, stretching queue times on certain machines into painful long tails.

Having laid bare these mechanics, the team instituted three structured routing rules. The first targeted high-concurrency sessions. An unavoidable operational reality exists here: session IDs are client-populated, and the system cannot prevent clients from firing multiple requests concurrently under the same ID. Coding agents frequently spawn parallel exploratory branches, generating bursts of concurrent long-context requests tagged with the same session ID. Rigidly honoring session affinity removes any ceiling on single-node concurrency, performing far worse than offloading excess traffic. Under the new policy, when concurrent requests exceed a safe high-water mark, the overflow is proactively shed to other nodes. The team ran the numbers: accepting the penalty of having another node redundantly rebuild part of the cache is far better than allowing an individual node to be crushed under runaway concurrency.

The second rule incorporated fine-grained real-time load awareness. The routing layer collects in-flight requests and VRAM cache utilization metrics across every node in real time, using these dynamic indicators to steer new sessions to whichever node is genuinely lightest at that moment; once state is established, subsequent requests generally preserve node affinity to protect cache hit rates.

The third rule avoids cache thrashing during cluster scale-outs. Traditional hash-based scaling forces roughly 1/N of existing sessions to drift during rebalancing, triggering widespread recomputations. The team optimized scale-out logic: when new nodes spin up, incoming sessions preferentially route to the fresh nodes (which have accumulated zero load); if the arrival rate of new sessions mismatches the fresh node capacity, session-level load awareness reassigns selected sessions to restore balance.

These three refinements distributed workload across nodes far more evenly than random hashing, slashing the magnitude of load deviation to less than half the mean. TTFT stabilized, and aggregate multi-instance throughput approached linear single-node scaling. Modal formalized this state- and load-aware request dispatching mechanism into the experimental configuration kv_aware_routing. While the write-up showcased pronounced improvements in load uniformity and operational stability, it did not quantify a standalone end-to-end speedup figure specifically attributable to the routing overhaul.

Pure random dispatching leaves some machines overloaded while others sit idle; load-aware dispatching levels out node workloads, cutting deviation by more than half.

The Limits of the Numbers

Stepping through these interlocking engineering implementations, systems engineers must remain circumspect. Every benchmark metric in inference optimization is a joint function of system configuration and workload characteristics. Change the workload, and the results shift dramatically—even if hardware and software remain untouched.

The write-up candidly highlighted three concrete examples. First, per-card token throughput is heavily dependent on the output length of the generation task. In a closed-loop request cycle, the time spent generating a single long sequence is long enough for the system to interleave multiple short-output requests, making aggregated throughput tightly coupled to output length ratios. Second, the speedup from speculative decoding hinges on the inherent regularity of the corpus. Applying the exact same technique to code versus prose, code achieved an average acceptance length roughly 2x that of prose. Shifting to more divergent, open-ended text generation erodes acceleration benefits significantly. Finally, cache hit rates are exquisitely vulnerable to messy client interaction patterns. If a client frequently reaches back to modify previous conversational instructions, cached prefixes spanning tens of thousands of tokens are instantly invalidated, reducing cache optimizations to zero on paper.

The 2.8x interactive speedup and 5.6x aggregate throughput disclosed in the retrospective represent the combined payoff of multiple compounding optimizations—speculative decoding, memory layout refactoring, reduced precision, and tiered caching—without isolated ablation of each measure. This pair of numbers is strictly bounded to the Kimi K2.6 model, Blackwell B200 accelerator cards, and this specific programming traffic profile; they cannot be unconditionally generalized.

From an industry perspective, as an infrastructure product providing compute services, Modal’s technical retrospective serves both as engineering proof-of-work and customer acquisition. Open-sourcing patch code cements engineering credibility, though dead ends and failed experiments naturally remain unpublicized. On the hardware evolution curve, B300 accelerator cards outfitted with larger VRAM will sustain higher concurrency; on model iteration, the team noted that this identical state management playbook has already been successfully migrated to Kimi K3, capturing the largest share of traffic for that model on public aggregation marketplaces.

The most valuable takeaway this retrospective offers is its diagnostic sequence: first characterize the workload profile and let bottlenecks reveal themselves; then systematically deconstruct the system across single-request latency, single-node cache capacity, and cluster-wide state routing in order. Numbers will inevitably date; this diagnostic sequence will not.