
Custom AI software that ships on the hardware.
Six modules, one stack, installed and tested on every system before it leaves our bench. The platform serves the models; the modules on top search, act, learn, govern and connect.
Layers you can point at.
Your applications sit on top. The hardware sits underneath. Governance and connectors run through everything in between.
Every module, in a sentence.
cstmAI™ Platform
The on-premises AI stack: model serving, access control, monitoring
cstmAI™ Search
Answers grounded in your own documents, with citations
cstmAI™ Agents
Workflow agents that act inside your systems, with human approval steps
cstmAI™ Studio
Fine-tuning and evaluation on your data
cstmAI™ Govern
Audit logs, permissions, data redaction and usage policy
cstmAI™ Connect
Connectors to ERP, CRM, document stores and data warehouses
Open-weight models we run.
| Llama | Meta |
|---|---|
| Qwen | Alibaba |
| Mistral | Mistral AI |
| DeepSeek | DeepSeek |
| Gemma | |
| gpt-oss | OpenAI |
Plus embedding, reranking, speech and vision models. Chosen per task by evaluation score.
Open-weight means the model files themselves run on your hardware: no per-request fee, no outside service in the loop, and no surprise when a provider retires a version. We pick candidates for each task and let your evaluation set decide.

Software questions.
Can we use the software on hardware we already have?
We assess existing hardware in discovery. If it fits the workload we can deploy on it; if not, we'll say what's missing.
Which model will we end up using?
The one that scores best on your evaluation set at a speed and hardware cost you accept. Often it's two or three: a large model for hard questions, a small one for routing, and an embedding model for search.
Are open-weight model licenses a problem for business use?
Usually not, but terms differ: some are permissive, some carry use restrictions or attribution requirements. We review the license of every model we propose with you before it goes into production.
Spec your system.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.


