icon-double-hexNOW AVAILABLE I BUILT ON icon-double-hex

Scality AI Inference Factory

Open-weight models on your own infrastructure, at a cost you set.

An open-code inference stack on Scality ADI (Al Data Infrastructure), validated, delivered and maintained by Scality. Keep your coding assistants, apps and agents, point them at one endpoint and keep every prompt inside your perimeter.

icon-double-hexCost, control and sovereignty

You control the model, the data and the cost.

Enterprises have found the Al workflows that pay off and now need them running at scale, every day. Three dependencies push the most critical ones onto infrastructure they own.

A budget that stays fixed.

A metered API bill grows with every token. On your own servers the budget is fixed, and the more your teams use it, the cheaper each token gets.

Data that stays yours.

Codebases, documents and tool outputs never leave your perimeter or your jurisdiction, and nobody outside your company can switch them off. The whole stack can run air-gapped.

The model you choose stays the model you run.

You choose from open-weight models Scality validates and delivers with the stack. Version, quantization and retirement change when you decide, not when a provider rebalances its fleet.

icon-double-hex-lightWHY A STORAGE COMPANY BUILT IT

The hardest part of serving a model is its memory.

Every returning conversation is a choice between recomputing its context on the GPU and reading it back. That context, the KV cache, outgrows GPU memory fast. In our lab, one GLM-5.2 session with a 438,764-token context produced 189 GB of it. Ten thousand sessions like that would need about 1.9 PB, which makes it a storage tier rather than a cache.

inference-storage-stack-diagram

icon-double-hexMEASURED IN THE SCALITY LAB

Context restored in 166 ms, not recomputed in seconds.

One eight-GPU server, five storage servers a few generations old and four 100 Gbit/s RDMA links, August 2026. The network was the ceiling, and the product runs on a 400 Gbit/s fabric. Your numbers depend on model, context length and traffic. See the full results ›

166 ms

time to first token (TTFT) from Scality ADI on a 14K-token context, 64 ms behind server DRAM

14x to 72x

faster to read the context back than to recompute it: 14x at 14K tokens, 72x at 439K tokens

8 TB

of KV cache held on Scality ADI, about 10x the server's GPU memory, first token flat as it grew

1.9s

to load Gemma-3 27B's 54.86 GiB of weights into one GPU, almost 10x faster than local drives

icon-double-hexWHAT YOU DEPLOY

One endpoint over the whole serving stack.

Scality packages, validates, delivers and maintains every layer as models and serving engines change, on standard servers from Dell, HPE, Lenovo and Supermicro. Every software component ships as open code: you receive the source and can audit how inference state is stored and moved.

inference-storage-endpoint-diagram

icon-double-hexWORKS WITH YOUR TOOLS

Keep your harness. Choose a validated model.

Scality packages, validates, delivers and maintains every layer as models and serving engines change, on standard servers from Dell, HPE, Lenovo and Supermicro. Every software component ships as open code: you receive the source and can audit how inference state is stored and moved.

use-case-inference-harness-icon

Harnesses, for example

use-case-inference-claude
use-case-inference-codex
use-case-inference-opensource

and the other tools your developers use

use-case-inference-harness-icon-1

APIs

use-case-inference-openai
use-case-inference-Anthropic

and the other tools your developers use

use-case-inference-validated-icon

Validated models

use-case-inference-gemma
use-case-inference-glm
use-case-inference-mistral
use-case-inference-qwen
use-case-inference-kimi
use-case-inference-servers

Standard servers

use-case-inference-dell
use-case-inference-hpe
use-case-inference-lenovo
use-case-inference-supermicro

icon-double-hexGO DEEPER

Read the story behind the launch.

icon-double-hexPress release

Scality launches AI Inference Factory for enterprise on-premises AI

The announcement, availability and the three dependencies driving inference on premises.

icon-double-hexProduct

Scality ADI

The AI Data Infrastructure underneath: one platform for AI, cyber resilience and sovereign control.

Column 3

Column 4

Bring your traffic. Well size it with you.

Talk to a Scality Al architect about your workload. Bring your traffic and a repository, and we will size the GPUs, the fabric and the cache against it.