Evaluation sets
Real examples with expected answers, versioned like code.
Tell us the models and the users. We return a configuration sheet and a quote.
Spec a system →Model serving, access control and monitoring on hardware you own.
See the platform →
Thirty minutes with an engineer: your workflow, your data, and whether custom AI fits.
Book a scoping call →
cstmAI™ Studio is the workbench for deciding which model to trust with a task, and for making it better at that task: evaluation sets, side-by-side comparisons, and fine-tuning runs that never leave your systems.
Every claim about a model should come with a score on your own work. Studio holds the evaluation sets (real inputs with expected outputs, written with your experts) and runs candidate models, prompts and fine-tunes against them, so a change ships only when the numbers support it.
When a better prompt isn't enough, Studio fine-tunes open-weight models on your GPUs, often overnight on the same system that serves by day, and keeps the resulting weights in your storage.
| Evaluation | Task test sets with automatic and human-graded scoring |
|---|---|
| Comparisons | Models, prompts and versions side by side |
| Fine-tuning | Adapter methods; full fine-tuning where hardware allows |
| Data | Stays on your systems throughout |
| Outputs | Versioned weights and score reports that you own |
Every module runs on the same platform and hardware, under the same governance.
Real examples with expected answers, versioned like code.
Candidate models scored on the same set, with cost and speed beside quality.
Lightweight fine-tunes that teach a model your formats and vocabulary.
Every new version re-scored before it replaces the old one.

Real inputs and the outputs your experts would accept.
Score current models and prompts to see the gap.
Prompt changes first, fine-tuning if the gap remains.
Ship the new version only if it wins on the set.
Often not. Retrieval and good prompts handle most tasks. We fine-tune when the evaluation set shows a gap prompts can't close: a strict house format, specialized vocabulary, or a small model that must match a large one.
Your domain experts and our engineers, together, from real examples. It is the most valuable artifact of the project, and it belongs to you.
Yes, scheduled around serving load. Large full fine-tunes may call for a Cluster.
Tell us the models you want to run, how many people will use them and where the hardware should live. An engineer replies with a first configuration and the questions that decide the quote.