cstm.ai · Custom AI hardware, software & developmentVendor-neutral · Quote onlyBuilt to spec
CSTM-SRC · REV AcstmAI™ · Hardware · Sourcing

Vendor-neutral AI hardware, sourced on the merits.

We spec from multiple vendors, with no lock-in. Every GPU, server, switch and drive on a cstmAI configuration sheet is there because it suits your workload, can be delivered, and can be supported.

Fig. 01 · Pattern storage above the shop floor, Paterson, NJ1994
Selection criteria

Seven questions for every part.

Each configuration sheet answers these in writing, so you can check our reasoning and reuse it for your next purchase.

CriteriaCSTM-SRC · REV A
Workload fitMeasured on your task: quality, response speed and concurrency on candidate hardware.
MemoryEnough GPU memory for the model, its context and the number of simultaneous users.
AvailabilityReal lead times from distribution, not list availability. A part you can't get is the wrong part.
Power & coolingWhat your room or colocation cage can actually supply and remove, continuously.
SupportWarranty terms, parts availability and the manufacturer's repair process.
Software supportMature drivers and runtime support for the models and serving software you'll use.
Exit costNothing proprietary that would tie your next purchase to one vendor, or to us.
Choosing GPUs

What actually decides the GPU.

Model names and launch benchmarks get the attention. These six factors decide whether a GPU suits your workload.

MEM

Memory capacity

The first filter. The model's weights, plus working memory for every active conversation, must fit. Quantization (storing weights at lower precision) shrinks the first part; it doesn't remove the second.

BW

Memory bandwidth

How fast the GPU can read those weights. For serving language models it often matters more than raw compute, because each generated word reads the model again.

LINK

Interconnect

When a model spans several GPUs, the links between them set the ceiling. High-bandwidth GPU-to-GPU links matter for large models and training; less so for many small ones.

FMT

Precision support

Newer GPUs run lower-precision number formats natively, which can raise throughput substantially. We check the model and serving software support them before counting on it.

PWR

Power and heat

Data-center GPUs draw several hundred watts each, some approaching a kilowatt. That decides circuits, cooling and whether a room works at all.

SUP

Supply

Lead times swing from days to months by part and quarter. We price alternatives so a shortage changes the schedule, not the project.

Families we evaluate: NVIDIA data-center and professional GPUs, AMD Instinct accelerators, and embedded AI modules for edge systems.
Archival photo of a two-story machine shop floor filled with belt-driven line-shaft machinery, flywheels, and gears laid out on the floor
Reel 02 · Machine shop interior, Paterson, NJca. 1890
FAQ

Straight answers on sourcing.

Do you have partnerships with hardware vendors?

Our recommendations aren't tied to any manufacturer. If a commercial relationship ever bears on a quote, we'll tell you in the quote.

NVIDIA or AMD?

Whichever fits. NVIDIA has the broadest software support today; AMD's data-center GPUs offer large memory and are well supported by mainstream serving software. We test your model on the software stack we'd ship before recommending either.

Should we wait for the next GPU generation?

Rarely. A new generation takes months to reach volume supply and mature software. If your workload is ready now, a system sized for it now usually pays back sooner; we plan the refresh at the same time.

Get a quote

Spec your system.

Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.

Form CSTM-Q · Quote only