5 Companies Building Rack-Scale AI Systems in 2026
A powerful accelerator does not automatically make a powerful AI system.At desktop scale, hardware buyers already understand this. A fast GPU can still be constrained by memory capacity, bandwidth, thermals, power…

A powerful accelerator does not automatically make a powerful AI system.
At desktop scale, hardware buyers already understand this. A fast GPU can still be constrained by memory capacity, bandwidth, thermals, power delivery, or the rest of the system. Hardware Times recently made a similar point about local AI hardware: CPU, GPU, memory, and workload requirements have to be considered together rather than judged from one specification.
Rack-scale AI takes that same problem and makes it much larger.
NVIDIA’s GB300 NVL72, for example, combines 72 Blackwell Ultra GPUs and 36 Grace CPUs into one liquid-cooled rack-scale architecture. NVIDIA lists 130 TB/s of NVLink bandwidth, while its enterprise reference architecture specifies integrated power shelves, NVSwitch trays, management switches, liquid-leak detection, and a full-rack power requirement of up to 142 kW.
At this scale, the accelerator is only one part of the engineering problem.
The facility has to support the rack. Cooling has to sustain it. The fabric has to move data quickly enough. Storage cannot starve the compute. Workloads need orchestration. Someone has to monitor the system after it becomes operational.
That is why a new group of AI infrastructure companies is increasingly building around the rack as the system, rather than treating individual accelerators as isolated components.
Here are five companies taking notably different approaches to that problem in 2026.
How I evaluated these companies
I focused less on headline accelerator counts and more on the engineering around the hardware.
The main questions were:
- Which current accelerator architectures are supported?
- Is the infrastructure designed at rack or cluster scale?
- How is cooling handled?
- What network fabric is available?
- How much control does the customer have over the physical system?
- Are Kubernetes, Slurm, or other orchestration tools supported?
- How is infrastructure health monitored?
- Who takes responsibility for deployment and ongoing operation?
- Is the architecture aimed at sustained production workloads?
For rack-scale infrastructure, architecture matters more than a simple hardware inventory. I gave greater weight to how each company handles networking, cooling, orchestration, deployment, monitoring, and the operational work required to keep high-density AI systems productive.
Quick comparison
1. Cambridge Nexus
CambridgeNexus (CNEX) is a Boston-based AI Factory operator focused on full NVIDIA GB300 NVL72 racks.
The company owns and operates the physical racks and customers lease them bare-metal, from a single full rack upward. CambridgeNexus structures each deployment around seven operating layers: power, cooling, networking, compute, orchestration, compliance, and customer workload planning.
That operating model is what makes CNEX particularly relevant to the current generation of rack-scale hardware.
Why CambridgeNexus stands out
It treats the NVL72 rack as one operating system
The GB300 NVL72 is not simply a collection of accelerators installed in the same cabinet.
NVIDIA describes it as a fully liquid-cooled rack-scale architecture in which 72 Blackwell Ultra GPUs operate within a high-bandwidth NVLink domain. Each accelerator also has access to ConnectX-8 networking for scale-out communication.
That tight coupling means the surrounding environment matters.
CambridgeNexus approaches the deployment as one integrated infrastructure system rather than separating the GPU purchase from networking, cooling, orchestration, and facility operations.
For a team running large training, reasoning, or production inference workloads, that reduces the number of infrastructure layers it has to coordinate independently.
Workload planning is part of the infrastructure decision
This is an area hardware comparisons often miss.
The fastest hardware on paper is not necessarily the best configuration for every workload.
Large-scale training prioritizes different things from latency-sensitive inference. Reasoning workloads can create different memory and utilization patterns again.
CNEX includes customer workload planning as one of its seven operating layers.
It also proposes the installation location according to workload, compliance, and latency requirements, with the selected site then established contractually.
That is a more useful way to approach infrastructure than choosing hardware first and trying to fit the workload around it afterward.
The deployment model is built around complete racks
CambridgeNexus works from one full GB300 NVL72 rack upward.
The racks are bare-metal and assigned at full-rack scale rather than being divided into smaller customer allocations.
For engineering organizations that already know they need an NVL72-scale environment, that provides a clear physical boundary around the infrastructure.
The deployment schedule is unusually concrete
Racks are pre-manufactured at CambridgeNexus’s factory in Taiwan and prepared for installation at the selected data-center location.
The deployment commitment is 60 days from contract to installation and acceptance, or faster depending on rack availability.
Typical industry lead times can run to quarters, so infrastructure teams planning a production launch need to consider physical deployment time long before current capacity is exhausted.
A useful metric is therefore not simply current utilization.
It is:
At the current workload growth rate, when must the next rack be installed and operational?
That question connects hardware planning to the actual product roadmap.
Best fit
CambridgeNexus is my first choice here for enterprises, AI labs, and other organizations that already require complete NVIDIA GB300 NVL72 racks and want the surrounding physical and operational layers handled under one model.
It is particularly relevant where full-rack isolation, deployment planning, orchestration, networking, cooling, and compliance all need to be treated as part of the same infrastructure project.
2. Sesterce
Sesterce takes a European AI Factory approach with a particularly detailed infrastructure stack.
Its current architecture combines direct-to-chip liquid cooling, NVIDIA accelerators, high-speed InfiniBand, storage, and its own orchestration layer. The company’s published hardware lineup includes GB300, B200, and H200 systems.
What stands out technically
Sesterce describes its AI Factory infrastructure as supporting rack densities of 150 kW and uses NVIDIA Quantum InfiniBand at 800 Gb/s.
Those two figures tell you quite a lot about the intended market.
This is not conventional server infrastructure with a few accelerators installed.
It is designed for high-density AI systems where the cooling loop, electrical design, and network fabric are all first-class parts of the architecture.
Sesterce OS adds another layer.
The platform supports both Slurm and Kubernetes and provides rack-level telemetry. That makes the infrastructure relevant to teams that may run large training jobs under Slurm while also operating production inference through Kubernetes.
Why the networking matters
Once training moves across large numbers of accelerators, the speed at which those accelerators exchange data can become a bottleneck.
Buying faster compute without matching it with an appropriate network fabric can leave expensive hardware waiting.
Sesterce’s emphasis on 800 Gb/s InfiniBand and RDMA is therefore more important than the accelerator list alone.
Best fit
I would shortlist Sesterce for European organizations that want NVIDIA-based high-density infrastructure with integrated orchestration and a strong emphasis on fabric design.
It is also interesting for teams that need both Slurm and Kubernetes rather than standardizing the entire environment on one scheduler.
3. TensorWave
TensorWave is the outlier in this list because its infrastructure is built around AMD Instinct accelerators rather than NVIDIA.
That makes it useful for hardware teams that want to evaluate whether AMD’s current memory capacity, ROCm software stack, and Ethernet-based scale-out architecture fit their workload.
Its current flagship infrastructure is based on the AMD Instinct MI355X.
What stands out technically
The MI355X provides 288 GB of HBM3E per accelerator with up to 8 TB/s of memory bandwidth.
TensorWave’s standard system combines eight MI355X accelerators in a liquid-cooled server with dual AMD Turin CPUs, 3 TiB of DDR5 memory, local NVMe storage, and a 3.2 Tb/s interconnect configuration.
That makes memory one of the most interesting reasons to evaluate it.
Large models can become memory-bound before they become compute-bound.
More accelerator memory can allow larger models or batches to remain closer to the compute instead of relying as heavily on techniques that shift data across the system.
Ethernet rather than InfiniBand
TensorWave also takes a different networking path.
Its inference infrastructure uses RoCEv2 and 400 Gb/s networking for distributed workloads.
That is worth watching because the AI networking market is no longer only a question of how much InfiniBand a cluster can deploy.
Ethernet-based AI fabrics are becoming more technically interesting as vendors push for higher bandwidth, lower latency, and easier integration with existing data-center networking.
Orchestration options
TensorWave provides bare-metal access with optional managed Kubernetes and Slurm.
That gives infrastructure teams several choices.
A research organization can keep a Slurm-centered workflow, while a production platform can use Kubernetes for inference or service-oriented deployments.
Best fit
TensorWave is the provider I would evaluate if the engineering team specifically wants to test AMD Instinct infrastructure against its NVIDIA-based alternatives.
The most compelling workloads will be those where accelerator memory capacity, open software tooling, or Ethernet scale-out is important enough to justify testing a different hardware ecosystem.
4. Firmus
Firmus approaches AI infrastructure from an unusually physical direction.
Rather than starting with the accelerator, the company designs the facility, cooling, electrical systems, compute, and control software as one system.
Its core building block is called HyperCube.
Firmus describes each HyperCube as a physical AI Factory module designed to accommodate 32 NVL racks in a primarily liquid-cooled configuration.
What stands out technically
Infrastructure is designed from the chip to the grid
Firmus describes itself as vertically integrated across AI infrastructure.
Its AI FactoryOS combines GPU telemetry, cooling information, power systems, and orchestration within one control layer.
That creates an interesting hardware-management model.
Instead of viewing the workload scheduler and electrical system as unrelated domains, Firmus is working toward coordinating them.
In March 2026, the company announced grid-integrated AI Factory software based around NVIDIA’s DSX Blueprint. Its Model-to-Grid approach is designed to connect workload behavior with grid conditions in real time.
For hardware engineers, that is one of the more interesting developments in the AI Factory space.
At sufficient scale, electricity is no longer just a utility bill. It becomes an input that can affect workload scheduling.
HyperCube makes the physical design repeatable
The other interesting part is modularity.
HyperCube integrates electrical and cooling infrastructure into a repeatable factory unit rather than designing each high-density environment entirely from scratch.
Firmus says the system is primarily liquid cooled and is designed around high-density NVL rack deployments.
That is one possible answer to a broader industry challenge: AI hardware generations are changing faster than traditional data centers were designed to change.
A more modular physical system can potentially make future infrastructure upgrades easier to plan.
Best fit
Firmus makes the most sense for organizations thinking at AI-campus scale rather than simply acquiring a rack or small cluster.
Its differentiation is the integration of compute infrastructure with energy, cooling, modular facility design, and control software.
5. Penguin Solutions
Penguin Solutions takes another path.
It is less about leasing capacity and more about helping organizations design, integrate, deploy, and operate their own AI infrastructure.
The company breaks its lifecycle into four stages: Design, Build, Deploy, and Manage.
For Hardware Times readers, the Build phase is probably the most interesting.
What stands out technically
Penguin Solutions performs rack-level integration and network configuration before hardware reaches the final data-center environment.
Its build process includes validation and burn-in testing of the integrated system prior to deployment.
This addresses a real problem with large AI systems.
Installing the servers does not prove that the cluster is production-ready.
Network fabrics must be validated. Storage needs to behave as expected. Thermal systems have to support sustained load. Accelerators need health monitoring. Firmware and software versions have to work together.
Penguin tries to move more of that validation earlier in the process.
ClusterWareAI handles the operational layer
Once the infrastructure is running, Penguin’s ClusterWareAI software provides cluster monitoring, automation, governance, and remediation.
The company’s June 2026 update added an AI Factory Operations Agent, automated accelerator remediation for Kubernetes workloads, and expanded system-health monitoring.
That is an interesting sign of where infrastructure management is heading.
The same AI techniques consuming accelerator infrastructure are increasingly being used to help operate that infrastructure.
Multi-vendor hardware support
Penguin’s OriginAI infrastructure supports both AMD and NVIDIA accelerator technologies, alongside different server, storage, and networking configurations.
That makes Penguin useful for organizations that do not want the infrastructure strategy tied entirely to one hardware vendor.
Best fit
Penguin Solutions is the strongest option on this list for organizations that want to own or control more of the AI infrastructure themselves but need specialist engineering around architecture, integration, commissioning, and long-term operation.
What rack-scale AI changes about hardware selection
The biggest change is that comparing accelerators is no longer enough.
With a desktop GPU, you can still make a reasonably useful comparison from board power, memory, bandwidth, benchmark results, and price.
Rack-scale AI introduces several additional systems that can each determine real performance.
1. Cooling becomes part of the compute architecture
The GB300 NVL72 is liquid cooled by design.
NVIDIA’s enterprise architecture includes tray-level and rack-level liquid leak detection, cooling manifolds, and integrated facility requirements.
That means cooling is no longer merely something the building provides.
It is part of the system.
Infrastructure buyers should therefore ask how the cooling loop is monitored, where responsibilities change between the hardware operator and facility operator, and what happens when the hardware moves to sustained load.
2. Fabric design matters as much as accelerator count
The GB300 NVL72 uses fifth-generation NVLink inside the rack.
Scale-out networking can then use NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet. NVIDIA specifies 800 Gb/s connectivity per GPU through ConnectX-8.
TensorWave takes a different route with AMD hardware and RoCEv2 Ethernet.
Sesterce emphasizes high-bandwidth InfiniBand.
Those are not minor implementation details.
Training efficiency can collapse if accelerators spend too much time waiting for data from other parts of the cluster.
3. Storage needs to feed the accelerators
The accelerator can only work on data it receives.
Large training sets, checkpoints, retrieval systems, and model weights all create storage traffic.
A serious infrastructure comparison should therefore include:
- Local NVMe capacity
- Network storage throughput
- Checkpoint behavior
- Read/write patterns
- Dataset size
- Network path between storage and compute
A rack with impressive theoretical FLOPS can still perform poorly if storage repeatedly stalls the workload.
4. Orchestration affects utilization
Hardware utilization is partly a software problem.
Slurm, Kubernetes, queueing policies, workload placement, and failure recovery determine how much useful work expensive hardware actually performs.
That is why CambridgeNexus includes orchestration and workload planning in its operating model, Sesterce supports both Slurm and Kubernetes, TensorWave provides managed orchestration options, Firmus connects orchestration to facility telemetry, and Penguin Solutions builds operational management into ClusterWareAI.
The hardware and software layers are becoming harder to separate.
5. Deployment time belongs in the architecture discussion
Engineers naturally focus on performance after installation.
Procurement teams also have to think about how long it takes to reach that point.
If a product roadmap assumes new capacity will be available in a particular quarter, physical infrastructure delays can become product delays.
CambridgeNexus’s 60-day contract-to-installation-and-acceptance model is one example of treating deployment time as an explicit infrastructure parameter rather than leaving it as an open-ended project variable.
Firmus approaches the same problem differently through prefabricated HyperCube modules.
Penguin Solutions pushes integration and burn-in into its factory process before on-site deployment.
Three different architectures, but the same underlying lesson: time to operational capacity is a hardware metric too.
Which company fits which type of deployment?
Full NVIDIA GB300 racks: CambridgeNexus
CambridgeNexus is my strongest overall recommendation for organizations that have already reached complete GB300 NVL72 rack requirements.
The differentiator is not simply access to Blackwell Ultra hardware. It is the decision to operate power, cooling, networking, compute, orchestration, compliance, and workload planning together.
For full-rack buyers, that is the most complete operating model in this group.
European NVIDIA infrastructure: Sesterce
Sesterce is compelling for high-density European deployments where modern NVIDIA hardware, 800 Gb/s fabric, and Slurm/Kubernetes orchestration all matter.
AMD Instinct infrastructure: TensorWave
TensorWave is the obvious alternative architecture to benchmark if the workload could benefit from the MI355X’s 288 GB of HBM3E, AMD ROCm, or a high-bandwidth Ethernet fabric.
Facility-scale AI engineering: Firmus
Firmus is the option to watch when the problem extends beyond racks into electrical systems, modular facility construction, energy-aware orchestration, and large AI Factory campuses.
Customer-controlled AI factories: Penguin Solutions
Penguin Solutions fits buyers that want to retain control of their infrastructure while bringing in specialist engineering for design, factory integration, deployment, validation, and operations.
The hardware question has moved up one level
For years, AI infrastructure discussions centered on the accelerator.
Which GPU?
How much HBM?
How many FLOPS?
What precision?
Those questions still matter.
But rack-scale systems are forcing hardware buyers to move one level higher.
The more useful questions now are:
How does the entire rack communicate?
How is it cooled?
How is power delivered?
How fast can storage feed it?
How are workloads scheduled?
How is infrastructure health measured?
Who operates the system when something degrades?
How quickly can additional capacity become operational?
NVIDIA’s GB300 NVL72 makes this shift especially obvious because the architecture itself is explicitly built at rack scale rather than around an individual server.
The companies that solve these surrounding problems well will matter just as much as the companies designing the accelerators.
FAQ
What does rack-scale AI mean?
Rack-scale AI treats the complete rack as an integrated computing system rather than a collection of independent servers.
The architecture can combine accelerators, CPUs, high-speed interconnects, networking, power delivery, cooling, storage, and management into one coordinated system.
How many GPUs are in an NVIDIA GB300 NVL72 rack?
NVIDIA’s GB300 NVL72 contains 72 Blackwell Ultra GPUs and 36 Grace CPUs. The GPUs operate within a fifth-generation NVLink domain.
Why does liquid cooling matter for AI infrastructure?
Rack-scale AI systems concentrate large amounts of compute into a small physical space.
Liquid cooling removes heat closer to the processors and makes much higher rack densities practical than conventional air cooling alone.
Is InfiniBand required for rack-scale AI?
No.
InfiniBand remains important for high-performance AI clusters, but Ethernet-based fabrics are increasingly being used for scale-out AI as well.
The right choice depends on workload communication patterns, scale, operational expertise, and the hardware architecture.
What should buyers benchmark before selecting an AI infrastructure provider?
Use the real workload where possible.
Measure training throughput or inference throughput, accelerator utilization, communication overhead, storage performance, failure recovery, and performance over sustained runs.
The best infrastructure is not the system with the most impressive isolated specification.
It is the system that keeps the expensive hardware doing useful work.
The post 5 Companies Building Rack-Scale AI Systems in 2026 appeared first on Hardware Times.