AI as an internal service
One pool of GPU capacity for many departments, each with its own quotas and access rules.
Tell us the models and the users. We return a configuration sheet and a quote.
Spec a system →Model serving, access control and monitoring on hardware you own.
See the platform →
Thirty minutes with an engineer: your workflow, your data, and whether custom AI fits.
Book a scoping call →
cstmAI™ Cluster joins two or more GPU servers with a high-bandwidth fabric, shared storage and a scheduler, for organizations that need the biggest open-weight models, thousands of users, or regular training runs on their own hardware.
A cluster is a different project from a server. The network fabric, storage throughput, power density and cooling decide whether the GPUs are busy or waiting, and each one is a design decision. We design all of it, and plan the facility work with your team or your colocation provider before anything ships.
On the software side, the cstmAI Platform schedules inference and training across nodes, so capacity is shared instead of stranded on one box. You grow by adding nodes of the same design.
| Form factor | One or more full racks (42U–48U) |
|---|---|
| Nodes | 2 or more GPU servers, typically 8 GPUs each |
| GPU memory | 80–192 GB per GPU |
| Fabric | 200–400 Gb/s per GPU; InfiniBand or RoCE Ethernet |
| Storage | Shared high-throughput storage, tens to hundreds of TB |
| Typical models | The largest open-weight models at full precision; many models at once |
| Concurrent users | Hundreds to thousands |
| Power & cooling | Often 20–40 kW or more per rack; liquid cooling at the high end |
| Setting | Data center or colocation |
Typical ranges. The quote names every part.No list prices
A representative elevation. The chassis we specify varies by vendor and generation; the callouts don't.
Fabric switches
GPU nodes, 8 GPUs each
Rear power distribution (hidden)
Shared storage
One pool of GPU capacity for many departments, each with its own quotas and access rules.
The biggest open-weight models at full precision, for the tasks smaller models can't handle.
Full fine-tunes and continued training on large private datasets, without renting GPUs by the hour.
Processing millions of documents, records or images in days instead of months.

Users, models, training cadence and growth over three years, turned into node counts.
Power, cooling, rack layout, network topology and storage, agreed with your facilities team.
Nodes and fabric tested together, not just one at a time.
Installed, then accepted against a test plan you sign off on.
Longer than a single server. GPU and networking lead times plus facility work set the pace, so we give a dated plan after the site survey. If you need production capacity sooner, start with a Rack designed to join the cluster later.
Both work. InfiniBand is the established choice for large training jobs; RoCE Ethernet suits teams that want one network technology across the data center. We recommend one based on workload, your staff's skills and supply.
Yes, and it is often the right path. We design the first Rack with the cluster's network and power in mind, so it joins later as a node instead of becoming a stranded box.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.