A model server is a GPU machine in your own server room that runs open-weight models for everyone in the company. Nobody counts your tokens, task content never leaves the network, and a busy month costs the same as a quiet one. Four builds, from a single-GPU box for a team of ten up to an eight-GPU rack.
Hardware, assembly and model bring-up in one price
Model serverOne to eight GPUs in one chassis
Worth saying first
A Harness licence and a model server are two different purchases, and most companies start with only the first. The licence runs the AI worker on a computer you own, pointed at whichever model you like — a hosted one is fine. A model server is what you buy when you want the model itself inside the building. Harness licence pricing.
Why own one
Four things a metered model cannot give you
The bill stops growing with use
Per-task pricing charges you most for the thing you bought the product to do. A routine that reads two hundred invoices costs two hundred times one that reads a single invoice, so the month you finally automate something properly is the month the invoice jumps. On your own machine that number does not move: a saturated GPU and an idle one cost the same.
The content never leaves the network
Whatever a routine reads, the model sees — contracts, payroll, customer records, source code. On a model server that text travels between two machines in your own building, and the answer to an auditor asking where the data went is a rack number instead of a sub-processor list.
The model does not change under you
Hosted models get retired, retuned and rate-limited on somebody else’s schedule, and a routine approved last quarter can quietly start behaving differently this one. Weights you host stay byte-for-byte the same until you decide to replace them, and you pick the day.
At the end of it you own something
Four years of per-seat subscriptions leave you with cancelled invoices. Four years of a model server leave you with a machine that still runs, weights that still work without anyone’s permission, and hardware you can resell or put to another job.
The arithmetic
Where the line crosses
Round numbers, not a quote. They are here because the question comes up in the first five minutes of every conversation, and because the honest answer depends on how hard you run the thing.
€25
per person, per month
What a business plan on a hosted AI costs per seat. Ten people is €3,000 a year — before a single routine touches a metered API on top of it.
€10–12k
once, for up to ten people
The Studio build, delivered with the models installed and serving. After that it is roughly €400 a year of electricity, whether ten people hammer it all day or nobody signs in.
3–4 years
until the seats cost more
That is the payback on ten quiet seats. Routines running all day against a metered API turn years into months — that workload is where a model server stops being a preference and starts being the cheaper option.
We would rather lose the sale than sell the wrong machine. If your load is genuinely light, a hosted model is the cheaper answer, and that is what we will tell you.
Four builds
From a box on a shelf to a rack in a datacentre
The configurations below are indicative: they exist so you can see roughly what your headcount costs before anyone picks up the phone. The exact build follows a conversation about the models you need and the load you will put on them.
gpt-oss-120b at MXFP4, Qwen3-32B FP8 at full context, or Llama 3.3 70B at INT4 — with an embedding model alongside for search.
The one that starts the conversation: an ordinary wall socket, a cupboard with airflow, no rack and no server room. If ten people is where you are, this is the whole answer.
DeepSeek V3.x and R1 at INT4, GLM-4.6 FP8, Qwen3-Coder-480B FP8.
Needs a rack with matching cooling. There is no NVLink, so tensor parallelism runs over PCIe 5: fine for inference, slow for fine-tuning.
Frontier
8× H200 141 GB HGX — 1128 GB of VRAM over NVLink
≈ €300–400k
Users1000+ people
Power~10–12 kW under load
Platform
2× Xeon or EPYC, 2 TB RAM, 400 GbE or InfiniBand; 8× B200 is the alternative
Models
DeepSeek V3.x FP8, Kimi K2 at INT4 (or FP8 with a limited KV cache), and full fine-tuning of 70B-class models.
Datacentre or colocation only, and lead times run into months.
What the price covers
In the price, and on your side
In the price
The hardware, assembled and tested under load before it ships.
Open-weight models installed, quantised and served behind one endpoint on your network.
Harness pointed at it, with your routines run against the model you chose.
A written handover: what runs where, how to restart it, how to add a model later.
On your side
Power, cooling and the space — a socket and a shelf for Studio, a rack with matching airflow for the rest.
Network access from the computers that will use it.
Day-to-day operation. Whoever keeps your other servers running keeps this one too.
There is no SLA and no remote administration in the price. Ask if you want either and we will quote it separately.
How it goes
Four steps, and only the first one needs you
1
Tell us the load
How many people, what the routines actually do, and whether anything has to stay on a segregated network. That is the whole brief — there is no discovery phase to pay for.
2
Configuration and a fixed price
We pick the GPUs for the models you need rather than the other way round, and you get one number with the build in it. If a smaller machine does the job, the quote is for the smaller machine.
3
Built and measured
The machine is assembled, the models are installed and served, and we run it at your kind of load before it leaves: tokens per second on the models you will actually use, not a benchmark sheet.
4
Installed and connected
It goes into your room, Harness is pointed at it, and your routines run against it end to end before we call the job done.
Tell us how many people and what they run
You get a configuration and one number back. No discovery workshop, no per-seat calculator, and no call to sit through before you see a price.