Company-wide assistant
Chat and document search for every employee, signed in with the accounts they already have.
Tell us the models and the users. We return a configuration sheet and a quote.
Spec a system →Model serving, access control and monitoring on hardware you own.
See the platform →
Thirty minutes with an engineer: your workflow, your data, and whether custom AI fits.
Book a scoping call →
cstmAI™ Rack is a 4 to 8 GPU server for the point where a pilot becomes a service the whole company relies on. We specify it to your models and traffic, burn it in, rack it in your server room or colocation cage, and run the cstmAI Platform on it.
Production means many people at once, uptime that someone answers for, and models large enough for hard tasks. One well-specified 8-GPU server can serve a large open-weight model to hundreds of users, or several smaller models side by side, with room left for fine-tuning jobs overnight.
We size from the workload backward: which models, how many requests a day, how much document data, what response time is acceptable. Then we choose GPUs, memory, networking, storage and power from across vendors and write down why. You get a configuration sheet, not a catalog number.
| Form factor | 2U–8U rackmount, standard 19-inch rack |
|---|---|
| GPUs | 4–8 data-center GPUs |
| GPU memory | 48–192 GB per GPU |
| System memory | 512 GB–2 TB |
| Storage | 8–60 TB NVMe; optional shared storage |
| Networking | 25–100 GbE to your network; 200–400 Gb/s fabric if clustered later |
| Typical models | Up to 70B at full precision; large mixture-of-experts models on 8-GPU builds |
| Concurrent users | About 25 to 500+, depending on workload |
| Power & cooling | Roughly 3–11 kW per server on 208–240 V circuits |
| Setting | Server room, data center or colocation |
Typical ranges. The quote names every part.No list prices
A representative elevation. The chassis we specify varies by vendor and generation; the callouts don't.
GPU sleds, 4–8 (hidden)
Hot-swap NVMe bays
Rack ears on 19 in rails
Front-to-back airflow
Chat and document search for every employee, signed in with the accounts they already have.
A large general model, a small fast one, and the embedding and reranking models that make search work.
Workflow agents that read and write in your ERP or CRM through scoped connectors, with approvals.
Fine-tuning and evaluation runs scheduled around daytime serving load.

Measure your task on candidate models and size for peak load with headroom.
A written configuration sheet and a quote. No surprises in the bill of materials.
Assembled and run under sustained load before it leaves the bench.
Installed, networked, joined to your identity provider and monitoring.
Model size and concurrency. We measure response speed per user for your task on candidate models, then size for peak load with headroom. If four GPUs meet the target with margin, we won't quote eight.
Not necessarily, but you need power and cooling an ordinary office lacks: dedicated 208–240 V circuits and continuous heat output comparable to several space heaters. We survey the site first and recommend colocation when the room can't take it.
Yes. Everything it needs runs locally, so it can operate fully disconnected. Software and model updates are applied by your team, on your schedule.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.