cstm.ai · Custom AI hardware, software & developmentVendor-neutral · Quote onlyBuilt to spec
CSTM-R · REV AcstmAI™ · Hardware

On-premises AI hardware for production work.

cstmAI™ Rack is a 4 to 8 GPU server for the point where a pilot becomes a service the whole company relies on. We specify it to your models and traffic, burn it in, rack it in your server room or colocation cage, and run the cstmAI Platform on it.

Fig. 01 · Integrating an electronics box, NASA Goddard2021

Production means many people at once, uptime that someone answers for, and models large enough for hard tasks. One well-specified 8-GPU server can serve a large open-weight model to hundreds of users, or several smaller models side by side, with room left for fine-tuning jobs overnight.

We size from the workload backward: which models, how many requests a day, how much document data, what response time is acceptable. Then we choose GPUs, memory, networking, storage and power from across vendors and write down why. You get a configuration sheet, not a catalog number.

Typical configuration rangesCSTM-R · REV A
Form factor2U–8U rackmount, standard 19-inch rack
GPUs4–8 data-center GPUs
GPU memory48–192 GB per GPU
System memory512 GB–2 TB
Storage8–60 TB NVMe; optional shared storage
Networking25–100 GbE to your network; 200–400 Gb/s fabric if clustered later
Typical modelsUp to 70B at full precision; large mixture-of-experts models on 8-GPU builds
Concurrent usersAbout 25 to 500+, depending on workload
Power & coolingRoughly 3–11 kW per server on 208–240 V circuits
SettingServer room, data center or colocation

Typical ranges. The quote names every part.No list prices

Front elevation

The drawing before the build.

A representative elevation. The chassis we specify varies by vendor and generation; the callouts don't.

Front elevation line drawing of a 4U rackmount GPU server19 IN RACK · 482.6 MM4U · 177.8 MM1234
Fig. 1 CSTM-R · REV A · Front elevationNot to scale
01

GPU sleds, 4–8 (hidden)

02

Hot-swap NVMe bays

03

Rack ears on 19 in rails

04

Front-to-back airflow

What runs on it

One server, several jobs, all of them yours.

01CSTM-R

Company-wide assistant

Chat and document search for every employee, signed in with the accounts they already have.

02CSTM-R

Models side by side

A large general model, a small fast one, and the embedding and reranking models that make search work.

03CSTM-R

Agents with real access

Workflow agents that read and write in your ERP or CRM through scoped connectors, with approvals.

04CSTM-R

Training after hours

Fine-tuning and evaluation runs scheduled around daytime serving load.

Wide view down a factory aisle with an overhead crane rail, worker reading a clipboard beside rows of machine tools
Reel 02 · Overhead crane above the shop floor, Paterson, NJ1994
How it's delivered

Four steps, each one signed off.

01

Size the workload

Measure your task on candidate models and size for peak load with headroom.

02

Configuration & quote

A written configuration sheet and a quote. No surprises in the bill of materials.

03

Build & burn-in

Assembled and run under sustained load before it leaves the bench.

04

Rack & connect

Installed, networked, joined to your identity provider and monitoring.

FAQ

Questions we hear first.

How do you choose between 4 and 8 GPUs?

Model size and concurrency. We measure response speed per user for your task on candidate models, then size for peak load with headroom. If four GPUs meet the target with margin, we won't quote eight.

Do we need a data center?

Not necessarily, but you need power and cooling an ordinary office lacks: dedicated 208–240 V circuits and continuous heat output comparable to several space heaters. We survey the site first and recommend colocation when the room can't take it.

Can it run without an internet connection?

Yes. Everything it needs runs locally, so it can operate fully disconnected. Software and model updates are applied by your team, on your schedule.

Get a quote

Spec your system.

Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.

Form CSTM-Q · Quote only