Memory capacity
The first filter. The model's weights, plus working memory for every active conversation, must fit. Quantization (storing weights at lower precision) shrinks the first part; it doesn't remove the second.
Tell us the models and the users. We return a configuration sheet and a quote.
Spec a system →Model serving, access control and monitoring on hardware you own.
See the platform →
Thirty minutes with an engineer: your workflow, your data, and whether custom AI fits.
Book a scoping call →
We spec from multiple vendors, with no lock-in. Every GPU, server, switch and drive on a cstmAI configuration sheet is there because it suits your workload, can be delivered, and can be supported.
Each configuration sheet answers these in writing, so you can check our reasoning and reuse it for your next purchase.
| Workload fit | Measured on your task: quality, response speed and concurrency on candidate hardware. |
|---|---|
| Memory | Enough GPU memory for the model, its context and the number of simultaneous users. |
| Availability | Real lead times from distribution, not list availability. A part you can't get is the wrong part. |
| Power & cooling | What your room or colocation cage can actually supply and remove, continuously. |
| Support | Warranty terms, parts availability and the manufacturer's repair process. |
| Software support | Mature drivers and runtime support for the models and serving software you'll use. |
| Exit cost | Nothing proprietary that would tie your next purchase to one vendor, or to us. |
Model names and launch benchmarks get the attention. These six factors decide whether a GPU suits your workload.
The first filter. The model's weights, plus working memory for every active conversation, must fit. Quantization (storing weights at lower precision) shrinks the first part; it doesn't remove the second.
How fast the GPU can read those weights. For serving language models it often matters more than raw compute, because each generated word reads the model again.
When a model spans several GPUs, the links between them set the ceiling. High-bandwidth GPU-to-GPU links matter for large models and training; less so for many small ones.
Newer GPUs run lower-precision number formats natively, which can raise throughput substantially. We check the model and serving software support them before counting on it.
Data-center GPUs draw several hundred watts each, some approaching a kilowatt. That decides circuits, cooling and whether a room works at all.
Lead times swing from days to months by part and quarter. We price alternatives so a shortage changes the schedule, not the project.

Our recommendations aren't tied to any manufacturer. If a commercial relationship ever bears on a quote, we'll tell you in the quote.
Whichever fits. NVIDIA has the broadest software support today; AMD's data-center GPUs offer large memory and are well supported by mainstream serving software. We test your model on the software stack we'd ship before recommending either.
Rarely. A new generation takes months to reach volume supply and mature software. If your workload is ready now, a system sized for it now usually pays back sooner; we plan the refresh at the same time.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.