Model serving
Several models at once, each request routed to the model that suits the task.
Tell us the models and the users. We return a configuration sheet and a quote.
Spec a system →Model serving, access control and monitoring on hardware you own.
See the platform →
Thirty minutes with an engineer: your workflow, your data, and whether custom AI fits.
Book a scoping call →
cstmAI™ Platform is the base layer of every cstmAI system. It serves open-weight models, controls who may use them, and shows what they are doing. Search, agents and fine-tuning all run on top of it.
Running a model is the easy part. Running it for an organization means single sign-on, roles, quotas, logs, health checks, model versions and an upgrade path. The Platform packages those, so your IT team operates one system instead of a pile of separate tools held together by a contractor.
It runs entirely on your hardware and depends on no outside model service. Your prompts, documents and any weights we train stay on your disks.
| Serves | Open-weight language, embedding, reranking, speech and vision models |
|---|---|
| Interfaces | Web chat for staff; an HTTP API for applications |
| Identity | Single sign-on through your identity provider; role-based access |
| Monitoring | Usage, response time and GPU health, with alerts |
| Runs on | Every cstmAI hardware configuration, Desk to Cluster |
| Network | Operates with no outbound internet connection |
Every module runs on the same platform and hardware, under the same governance.
Several models at once, each request routed to the model that suits the task.
Sign-in with your existing accounts; roles and per-team limits on models and data.
Who uses what, how fast it answers, and whether the GPUs are healthy.
Versioned models, staged rollouts and rollback when a new version scores worse.

Loaded and tested during hardware burn-in, before delivery.
Joined to your identity provider; roles mapped to your groups.
The models chosen in discovery, with their evaluation scores recorded.
Runbooks and admin training for your IT team.
Open-weight models such as Llama, Qwen, Mistral, DeepSeek, Gemma and gpt-oss, plus embedding and speech models. We pick and test models for your tasks. License terms differ by model, and we review them with you.
That is the aim. Handover includes runbooks and admin training. If you'd rather we keep it running, managed operations is available.
No. It runs entirely on your hardware, and updates are applied on your schedule.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.