
AI server configuration, compared.
Four sizes side by side, in the ranges we specify within. Use it to find your starting point; the quote turns each range into a specific part list for your workload.
Four sizes, one table.
Typical ranges, not product specifications. Prices are never listed: every configuration is quoted from a written sheet.
| Spec | CSTM-D · REV AcstmAI™ Desk | CSTM-R · REV AcstmAI™ Rack | CSTM-C · REV AcstmAI™ Cluster | CSTM-E · REV AcstmAI™ Edge |
|---|---|---|---|---|
| Elevation | ||||
| GPU count | 1–2 | 4–8 | 16+ (2 or more nodes) | Embedded module or 1–2 GPUs |
| GPU memory class | 24–96 GB per GPU (workstation) | 48–192 GB per GPU (data center) | 80–192 GB per GPU (data center) | 16–64 GB |
| Typical model size | 7B–70B, quantized at the top end | Up to 70B at full precision; large MoE on 8 GPUs | Largest open-weight models, several at once | 1B–14B, quantized |
| Concurrent users | ≈ 1–25 | ≈ 25–500+ | Hundreds to thousands | One site, or a machine feed |
| Training | Small adapter fine-tunes | Adapter and moderate full fine-tunes | Large full fine-tunes | Inference only, as a rule |
| Power | ≈ 0.8–2 kW | ≈ 3–11 kW per server | Often 20–40 kW+ per rack | ≈ 60 W–1 kW |
| Rack units / form | Tower, or 4U rackmount | 2U–8U | Full racks, 42U–48U | Fanless box, short-depth 1U–2U, rugged |
| Cooling | Air | Air; data-center airflow | Air or direct liquid cooling | Passive or sealed |
| Network | 1–10 GbE | 25–100 GbE | 200–400 Gb/s fabric per GPU | LAN, cellular or satellite |
| Deployment setting | Office, lab, closet | Server room, data center, colo | Data center or colo | Plants, vessels, field sites |
| Best for | Pilots and single teams | Production for a company | Org-wide AI service and training | Offline and remote work |
| Next step |
Figures are typical ranges for planning. Actual capacity depends on model, context length and usage pattern.
Sizing, without guesswork.
How do you decide which size we need?
From measurements, not a rule of thumb. In discovery we run your task on candidate models, record response speed and quality, then size for your peak number of simultaneous users with headroom.
Why are these ranges so wide?
Because the same box serves very different loads. Ten people asking long questions about large documents can use more GPU than two hundred asking short ones. The quote narrows each range to one number for your workload.
Can we mix sizes?
Yes, and many organizations do: a Rack in the data center, Edge units at remote sites, and a Desk for development. They all run the same software and are managed the same way.
Spec your system.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.


