Pilot on real data
Prove search, extraction or an agent on your own documents before you commit to production hardware.
Tell us the models and the users. We return a configuration sheet and a quote.
Spec a system →Model serving, access control and monitoring on hardware you own.
See the platform →
Thirty minutes with an engineer: your workflow, your data, and whether custom AI fits.
Book a scoping call →
cstmAI™ Desk is a desk-side system with one or two professional GPUs, sized for a proof of concept, a department, or a lab that wants models running next to the work. It arrives burned in, with the cstmAI Platform installed and your first use case loaded.
Most custom AI projects should start small and close to the people who will judge them. A workstation-class machine runs a capable open-weight model for a team of a dozen or two, needs no server room, and at the lower end of the range plugs into an ordinary office circuit.
We choose the chassis, GPUs, memory and storage for the model you plan to run and the number of people who will use it. If the pilot earns a production budget, the same software moves to a cstmAI™ Rack without rework: models, prompts, evaluation sets and connectors carry over.
| Form factor | Desk-side tower; 4U rackmount on request |
|---|---|
| GPUs | 1–2 professional workstation GPUs |
| GPU memory | 24–96 GB per GPU |
| System memory | 128–512 GB |
| Storage | 2–16 TB NVMe |
| Typical models | 7B–70B parameters (quantized at the upper end) |
| Concurrent users | About 1–25, depending on workload |
| Power | Roughly 0.8–2 kW; a dedicated circuit at the top of the range |
| Setting | Office, lab or small server closet |
Typical ranges. The quote names every part.No list prices
A representative elevation. The chassis we specify varies by vendor and generation; the callouts don't.
GPU cards, 1–2 (hidden)
Front I/O and power
Intake grille
Prove search, extraction or an agent on your own documents before you commit to production hardware.
A shared model and document search for a single team: legal, engineering, finance or operations.
Where engineers (yours or ours) iterate on prompts, evaluation sets and fine-tunes without touching production.
A machine that stays in one room, for data that should not travel even inside your own network.

A workload interview: which model, how many people, what data. We write the configuration sheet.
Assembled, stress-tested under sustained GPU load, imaged with the cstmAI software.
Delivered, networked, connected to your single sign-on, first use case loaded.
Runbook, admin walkthrough and a named contact for support coordination.
For many workloads, yes. Two 48 GB-class GPUs hold a 70B-parameter model at 4-bit quantization, and 8B–32B models run comfortably on one. We test your actual task on candidate models during discovery and show you the scores before you choose hardware.
The software is the same on every cstmAI configuration, so moving to a Rack is a hardware change, not a rebuild. Many teams keep the Desk afterward as their development and evaluation machine.
Whichever fit the workload and are available when we order. We spec across vendors and generations; for workstations that usually means current professional cards with 48 GB or more of memory. The quote explains the trade-offs.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.