Multi-Tenant GPUaaS Proven at Scale: DDN EXAScaler + Bridge by Armada
Inference has overtaken training as the dominant AI workload, consuming two-thirds of all AI compute in 2026. Enterprises now need to execute always-on inference and fine-tuning at large scale, and serving that demand requires a highly efficient AI services platform that can cater to hundreds of customers on a single cloud infrastructure.
DDN EXAScaler® and Bridge by Armada enable neoclouds, NVIDIA Cloud Partners (NCPs), telcos, and sovereign AI operators to sell AI services to hundreds of tenants at once with market-leading GPU efficiency. The solution is in production today with multiple customers.
Together, the two products form a GPUaaS platform that is dynamic, multi-tenant, and self-service from the ground up. DDN's EXAScaler supplies high-performance parallel storage with multi-tenancy built into the data plane. Bridge by Armada adds the GPUaaS management plane, automating on-demand, self-service, cloud-like consumption of GPUs — from infra to tokens — across compute, networking, and storage, with secure multi-tenancy, onboarding, expansion, observability, billing, and chargeback.
The result: faster rack-to-revenue per tenant, higher GPU utilization, and a clean path into the inference and fine-tuning market.
The Challenge: Delivering Dynamic Multi-Tenancy at Scale
Inference and fine-tuning workloads have a fundamentally different shape from the training contracts that built today's GPUaaS industry. They run for days or weeks rather than months, use smaller GPU footprints in the 1-to-100 range, and arrive in bursty, always-on patterns. To serve them profitably, operators need to share infrastructure across hundreds of customers rather than dedicate separate hardware per tenant.
That sharing model only works with real on-demand multi-tenancy, which requires strong dynamic isolation across compute, networking, and storage in tight integration. It means an automated tenant lifecycle and on-demand self-service consumption of GPUs covering onboarding, expansion, eviction, observability, billing, and chargeback — and it has to hold up consistently across multiple regions and data centers. It must deliver bare-metal-to-tokens for hundreds of tenants while maintaining the per-tenant performance SLAs and audit-grade isolation that sovereign and regulated customers require.
Today, operators have two paths to build this:
- Hard-partition the cluster per customer, which leaves utilization on the table and caps revenue growth.
- Build a custom multi-tenant management layer in-house, which requires a hyperscaler-sized engineering investment that few organizations can justify.
The Solution: A Turnkey Enabler for Multi-Tenant GPUaaS
DDN EXAScaler and Bridge by Armada offer a turnkey solution for multi-tenant GPUaaS, proven at production scale. The two products integrate as tightly coupled layers, replacing what operators would otherwise have to build themselves.
DDN EXAScaler is the high-performance storage foundation — a parallel file system purpose-built to keep AI clusters saturated, with native multi-tenancy primitives built directly into the data plane.
Bridge by Armada sits above it as the GPU-as-a-Service management plane that turns storage, GPU, and networking primitives into a self-service multi-tenant cloud, handling tenant lifecycle automation, the tenant portal, per-tenant observability, and billing and chargeback.
Bridge provides cloud services across three layers:
| Layer | Services |
|---|---|
| IaaS | Bare metal, VM, storage, VPC, security groups |
| PaaS | Managed Kubernetes |
| AIaaS | Models, Slurm, Jupyter Notebook, MLOps |
It also provides an ecosystem of AI application and platform software partners that includes NVIDIA AI Enterprise software, ClearML, Dataiku, Multiverse Computing, Primalabs, SecureIn, LinkerVision, and more.
The two layers integrate directly rather than as bolted-together products. Bridge drives EXAScaler through the EMF API to provision tenants, configure isolation, and enforce quotas. These outcomes translate directly into faster rack-to-revenue and lower cost-to-serve per tenant. Bridge plus EXAScaler is in production today at NCP customers, including deployments in Indonesia and North Africa.
Dynamic Multi-Tenancy at Scale
A real multi-tenant cloud must work the day an operator signs their hundredth tenant, not only the day they sign their first. EXAScaler's native multi-tenancy is validated at up to 1,000 tenants per cluster in internal testing, with 150 tenants in a current production deployment. That is proof the architecture holds up at production scale, not just in the lab.
Activation is a single configuration flag. There is no separate management stack to deploy and no offline maintenance window. Once active, tenants can be added, expanded, shrunk, or evicted online. Block and inode quotas change on the fly while a tenant is running, and new clients can be brought into a tenant without remounting.
Bridge wraps that lifecycle in a self-service portal, so tenants provision their own bare-metal GPU and storage environments through a UI or API rather than a ticket queue — keeping the operator's ops team off the critical path for onboarding and expansion.
This is what makes the inference economy workable for the operator. Inference customers are bursty by design, and capturing them profitably requires dynamic tenant onboarding, expansion, and chargeback. Multi-tenancy at this scale is what allows the operator to serve them, accelerating time-to-revenue per tenant while keeping the ops team scaling with the platform rather than linearly with the tenant count.
Dynamic Multi-Tenancy from a Storage Perspective
| Capability | Details |
|---|---|
| Maximum number of tenants | Up to 1,000 tenants (soft limit, from internal testing). Current largest production deployment uses 150 tenants. |
| Tenant expansion / shrink | Add or remove clients via CLI/API — removing clients requires client unmount or eviction to revoke data access. Increase or decrease tenant storage quota at any time. |
| Tenant security | Unique subdirectory mount accessible only by specific clients, by network address or segment. No limit on the number of clients per tenant. |
| Tenant authentication | Each tenant can support up to 100,000 unique users/groups. |
| Tenant administration | Only the Filesystem Administrator can create or delete tenants or modify tenant quota allocations. The Local Tenant Administrator can fully manage POSIX permissions, and will support tenant-local user/group/project quotas in future. |
| Tenant QoS | Tenant QoS management is under development, with per-tenant throughput limits. |
Hard Tenant Isolation
Multi-tenancy at scale only delivers value if the isolation between tenants holds. EXAScaler enforces that isolation at the storage data plane. Under EXAScaler's multi-tenancy protocol, each tenant's clients see only their own subtree of the filesystem; everything outside that file set is invisible, and traffic from clients not registered with a tenant is rejected at the protocol layer.
Resource isolation is enforced through Lustre Quota Aggregation (LQA), which applies block and inode limits across the tenant's entire ID range. This way, tenant local admins can freely manage user, group, and project quotas inside their tenant without ever expanding the operator's allocation to them. EXAScaler also provides full network isolation: storage management traffic rides on a dedicated VLAN segmented from tenant networks, so tenant clients have no path to administrative interfaces.
Hard multi-tenancy is enforced at every layer of the stack: EXAScaler at the storage data plane, and Bridge by Armada across compute (CPU, GPU), networking (Spectrum-X), InfiniBand, NVLink, and the firewall gateway.1 The combined architecture is NCP-validated end to end, which opens up revenue from sovereign governments, telcos, regulated industries, and security-conscious enterprises that would otherwise reject any shared-infrastructure offering.
Storage Multi-Tenancy

Each tenant's volume on the EXAScaler servers is reachable only over that tenant's own VLAN. A client attached to the Tenant 1 VLAN has no path to the Tenant 2 or Tenant 3 file sets — the separation is carried end to end, from the server-side file set through the storage network to the client.
AI-as-a-Service
To ensure that the data scientist or ML engineer spends their time generating AI insights rather than getting applications and infrastructure to work, Bridge by Armada has simplified AI-as-a-Service all the way from data ingest, model optimization, model security, inferencing, and fine-tuning, through to sending a checkpoint back to the cloud.
Bridge offers managed open-source projects that use DDN storage for Model-as-a-Service, Jupyter Notebooks, Slurm, and MLOps. In addition, Bridge offers a marketplace of commercial AI model, platform, and application vendors:
| Category | Vendors |
|---|---|
| Model optimization / compression | NVIDIA NIM, Multiverse Computing, Primalabs |
| Model security | SecureIn |
| AI platforms | NVIDIA Run:ai, ClearML, Dataiku |
| Vision language model | LinkerVision |
Conclusion
The shift from training contracts to inference and fine-tuning is reshaping how AI infrastructure providers do business. The opportunity is large, but capturing it depends on a different operating model: real multi-tenancy with hard isolation across compute, networking, and storage; an automated tenant lifecycle that scales with the platform rather than the headcount; and the audit-grade controls that sovereign and regulated customers demand.
DDN EXAScaler and Bridge by Armada deliver this as a turnkey enabler for multi-tenant GPUaaS, proven at scale. The two-layer solution combines high-performance parallel storage that has native multi-tenancy primitives with a GPU-as-a-Service management plane that automates the entire tenant lifecycle. The combined stack is in production today at NCP customers, validated at up to 1,000 tenants per cluster, and delivered under an NCP-validated reference architecture.
For NCPs, sovereign AI operators, and AI factory builders, it is the fastest path from infrastructure to revenue.

