ScaleFlux has launched an AI-optimized SSD platform designed for NVIDIA CMX and different inference architectures that use SSDs as a shared KV-cache tier past GPU HBM and host DRAM. The platform combines high-endurance SSD {hardware}, Versatile Information Placement (FDP) assist, and workload telemetry meant to enhance information placement, scale back write amplification, and lengthen efficient endurance in high-churn inference environments.
The platform addresses three storage challenges with offloaded KV cache: characterizing real-world workload habits, separating information by lifecycle, and sustaining heavy write exercise with out utilizing extra flash capability to soak up writes.
Lengthy-context inference, shared-prefix reuse, agentic utility flows, and retained idle classes improve the amount of reusable runtime state saved exterior GPU reminiscence. Not like standard enterprise workloads, KV-cache blocks could also be written often, retained for various durations, reactivated after inactivity, and invalidated asynchronously throughout classes, staff, and tenants. These patterns improve rubbish assortment exercise and write amplification when information with incompatible lifecycles share the identical flash blocks.

ScaleFlux’s high-endurance structure is designed to ship 7 to greater than 10 efficient drive writes per day at 5 years for KV cache workloads, with the corporate noting that efficient endurance is dependent upon workload traits, FDP utilization, and system configuration. Increased efficient endurance reduces the uncooked flash capability operators should deploy purely to soak up write site visitors, leaving extra put in capability obtainable to carry energetic KV cache and different AI runtime state. ScaleFlux calls that overhead the “endurance tax,” and decreasing it’s the platform’s core financial argument.
The SSD platform helps over 200 FDP write streams per drive. This lets inference software program group information by lifecycle, session, tenant, shared-prefix classification, possession, or reuse habits earlier than inserting it on flash media. The objective is to jot down information with comparable invalidation patterns collectively, decreasing inner information motion throughout rubbish assortment and limiting interference throughout information lessons.
In preliminary managed testing, ScaleFlux measured greater than a twofold discount in write amplification utilizing lifecycle-aware FDP placement in contrast with a baseline placement configuration. The corporate notes that precise outcomes rely upon workload traits, lifecycle classification, software program integration, and system configuration.
“AI inference infrastructure wants SSDs that present greater than further capability,” stated Hao Zhong, CEO and co-founder of ScaleFlux. “Infrastructure groups want to grasp how KV workloads have an effect on the drive, separate information based on lifecycle, and maintain excessive write charges with out deploying extra capability merely to dilute writes. ScaleFlux brings workload intelligence, scalable FDP placement, and 7-10+ efficient DWPD collectively in a single AI-optimized SSD platform.”
On the software program and telemetry layer, ScaleFlux Context-Perception SSD exhibits how KV-cache insurance policies have an effect on SSD operation. The platform captures latency, queue depth, throughput, request-size distribution, information age, write-to-first-read intervals, learn reuse, NAND write quantity, rubbish assortment motion, and write amplification.
Context-Perception can function in SSD-only mode for preliminary workload evaluation with out modifications to higher software program layers. With deeper integration, it correlates SSD telemetry with utility metadata, together with session IDs, employee or tenant identifiers, shared-prefix IDs, KV-block possession, lifecycle state, and key-to-block mappings. This lets operators affiliate latency, endurance consumption, and write amplification with particular workload lessons as a substitute of treating the SSD as an opaque shared useful resource.
ScaleFlux positions the platform as a complement to NVIDIA’s just lately introduced CMX Context Reminiscence Storage Platform, which offers a shared, pod-level context tier for high-speed KV cache entry and reuse. The pitch is aimed toward AI manufacturing facility operators: CMX handles the context tier, whereas ScaleFlux addresses the endurance, information placement, and write amplification challenges particular to the underlying SSDs.
“As AI inference methods lengthen KV cache past GPU Reminiscence and DRAM, understanding the habits and necessities of the SSD tier turns into more and more necessary,” stated Jason Hardy, vp of storage know-how at NVIDIA. “Our engagement with ScaleFlux helps characterize how KV cache offload impacts storage necessities for latency, endurance, and write amplification, contributing to the broader storage ecosystem round NVIDIA CMX.”
The corporate is creating a trace-driven simulator that fashions KV-cache motion throughout GPU HBM, host reminiscence, and SSD tiers. The simulator generates replayable SSD traces to guage placement, eviction, and lifecycle-grouping insurance policies underneath managed circumstances.
ScaleFlux plans to showcase the platform at FMS, masking Context-Perception workload evaluation, KV metadata correlation, lifecycle-aware FDP placement, write amplification discount, and high-endurance operation for write-intensive KV cache workloads. The corporate has a considerable presence on the present: ScaleFlux stated in July that its specialists would lead seven displays there, together with a keynote co-delivered with NVIDIA on reminiscence options for scaling the AI information pipeline.
The platform announcement caps a busy stretch of silicon information. Two days earlier, ScaleFlux unveiled two PCIe Gen6 elements it is going to introduce on the similar present: the FC6116 NVMe SSD controller and the MC600 CXL 3.2 Kind 3 reminiscence controller. ScaleFlux charges the FC6116 at as much as 28 GB/s sequential learn and 25 GB/s sequential write, as much as 7 million 4K random learn IOPS and greater than 1 million sustained 4K random write IOPS, underneath 9W energetic controller energy, with assist for TLC, QLC, and SLC NAND as much as 256TB throughout E1.S/L, E3.S/L, and U.2/3. The MC600 attracts underneath 9W typical in a Gen6 x8 configuration and handles quad-channel DDR5 or dual-channel DDR4 with as much as 2TB of DDR5, a dual-generation functionality ScaleFlux positions as a strategy to carry present DDR4 right into a CXL deployment. Each start sampling with key prospects in This autumn 2026.
