USE CASE - AI
On-Premises Al Inference
Scality Al Inference Factory is an open-code inference stack built on Scality ADI (Al Data Infrastructure), validated, delivered and maintained by Scality on your own infrastructure. Your developers keep the tools they use today and point them at one endpoint. The KV cache moves to Scality ADI, where it has no size limit, and you control the model, the data and the cost.
THE PRESSURES
Pressures on Al inference at enterprise scale.
Coding agents, internal assistants and Al inside your own products have moved from pilots to daily work. What slows them down now is what a token costs, who controls the model, where the context goes and how much memory serving takes.
workloads
The bill grows fastest when Al works.
A metered API bill grows with every token, so it grows fastest when adoption is working. That makes it hard to forecast and impossible to cap.
model control
Cloud models change on the providers schedule.
Version, quantization, safety routing and retirement are the provider's decisions. A result you validated in March may not reproduce in September.
DATA
Every request carries your context.
An external API never receives only a question. It receives the codebase, the retrieved documents and the output of every tool the agent called, under another jurisdiction's rules.
Resilience
AI doesn't change the regulator's posture.
The board, the regulator, and the cyber-insurance carrier still expect immutability, recovery, and audit trail. AI workloads inherit that posture rather than escape it. The platform has to carry it without slowing the GPU.
THE SCALITY ANSWER
How Scality Al Inference Factory serves your models.
One endpoint.
Claude Code, Codex, OpenCode and the other tools your developers use connect through OpenAI- and Anthropic-compatible APIs, and keep working as before.
A control plane that governs every request.
Every request is authenticated, metered and throttled, routed to the GPU that already holds its context and scheduled against your service levels.
Prefill and decode, split.
Prompt processing and token generation run in separate GPU pools that scale independently, so a long prompt no longer stalls everyone else's answer.
Zero-copy movement.
KV cache and model weights move over a 400 Gbit/s RDMA fabric with the CPU off the data path.
A shared KV cache on Scality ADI.
Every conversation's context, readable by any GPU, in a tier that grows as you add storage servers.
MEASURED IN THE SCALITY LAB
Don't rebuild the context. Read it back.
We spent the past year wiring GPUs straight into Scality ADI over RDMA and measuring what came back. The lab was deliberately modest: one eight-GPU server, five storage servers a few generations old and four 100 Gbit/s links. The network was the ceiling, and the product runs on a 400 Gbit/s fabric.
166 ms
time to first token (TTFT) from Scality ADI on a 14K-token context, 64 ms behind server DRAM
14x to 72x
faster to read the context back than to recompute it: 14x at 14K tokens, 72x at 439K tokens
8 TB
of KV cache held on Scality ADI, about 10x the server's GPU memory, first token flat as it grew
1.9s
to load Gemma-3 27B's 54.86 GiB of weights into one GPU, almost 10x faster than local drives
Inference workloads we serve
Inference workloads that re-read the same context all day.
Agents and long conversations come back to the same repository, the same documents and the same instructions, turn after turn. A published study of agentic workloads found 84.6 to 99.5 percent of input tokens reused from cache between turns.
Coding agents for developers.
Coding agents for developers.
Coding agents for developers.
The fastest-growing line on many AI bills.
Every session loads the same repository, system prompt and tool definitions. A session left open through a meeting expects to pick up where it stopped, and the same study saw cache hit rates fall from about 98 percent to between 16 and 26 percent once too many agents crowded the cache.
Challenges.
A metered bill that grows with every developer you onboard. Prompts that carry proprietary code outside the perimeter. Context that falls out of GPU memory and has to be recomputed.
Benefits of Scality AI Inference Factory.
Developers keep Claude Code, Codex, OpenCode and the other tools they use, and change one setting. The shared KV cache on Scality ADI lets any GPU pick up a returning session. The budget is fixed, and the more your teams use it, the cheaper each token gets.
Assistants and agents over internal data.
Assistants and agents over internal data.
Assistants and agents over internal data.
Where the context is your data.
Thousands of employees ask questions of the same policies, contracts and knowledge bases, and agentic workflows chain tools across them. Long conversations and multi-step agent runs keep returning to context they have already built.
Challenges.
Documents, retrieved data and tool outputs sent to an external API. Retention and moderation policies set by the provider. A jurisdiction you do not choose.
Benefits of Scality ADI.
High-concurrency S3 access and GPU-Direct data paths keep GPUs fed at scale. Durable, immutable checkpointing built into the same platform that holds the corpus. Performance at scale, not benchmark peak.
AI inside your own products.
AI inside your own products.
AI inside your own products.
Where the model's behavior is part of your promise.
When AI is part of what you sell, a model that stays as you validated it, a known cost per request and a service level you control are product requirements. A provider's quiet model update becomes your customer's changed answer.
Challenges.
Model updates that change answers without notice. Per-token pricing that squeezes margin as usage grows. Latency targets a shared cloud queue does not promise.
Benefits of Scality AI Inference Factory.
You choose from open-weight models Scality validates and delivers with the stack, and version, quantization and retirement change only when you decide, so a result you validated in March reproduces in September. SLA-aware scheduling and split serving hold first-token and token-pace targets, and the control plane meters every request.
THE STACK UNDER ONE ENDPOINT
One endpoint, four planes, one shared memory.
Scality Al Inference Factory packages the whole serving stack behind one endpoint. Scality validates and delivers it and keeps every layer current as models and serving engines change, and every software component ships as open code.
ENDPOINT
WORKS WITH YOUR TOOLS
Keep your harness. Choose a validated model.
Scality packages, validates, delivers and maintains every layer as models and serving engines change, on standard servers from Dell, HPE, Lenovo and Supermicro. Every software component ships as open code: you receive the source and can audit how inference state is stored and moved.
Bring your traffic. Well size it with you.
Talk to a Scality Al architect about your workload. Bring your traffic and a repository, and we will size the GPUs, the fabric and the cache against it.














