Model servers

Rent a model forever, or buy the machine once.

A model server is a GPU machine in your own server room that runs open-weight models for everyone in the company. Nobody counts your tokens, task content never leaves the network, and a busy month costs the same as a quiet one. Four builds, from a single-GPU box for a team of ten up to an eight-GPU rack.

  • No per-task fee, ever
  • Task content stays on your network
  • Hardware, assembly and model bring-up in one price

Model serverOne to eight GPUs in one chassis

Worth saying first

A Harness licence and a model server are two different purchases, and most companies start with only the first. The licence runs the AI worker on a computer you own, pointed at whichever model you like — a hosted one is fine. A model server is what you buy when you want the model itself inside the building. Harness licence pricing.

Why own one

Four things a metered model cannot give you

The bill stops growing with use

Per-task pricing charges you most for the thing you bought the product to do. A routine that reads two hundred invoices costs two hundred times one that reads a single invoice, so the month you finally automate something properly is the month the invoice jumps. On your own machine that number does not move: a saturated GPU and an idle one cost the same.

The content never leaves the network

Whatever a routine reads, the model sees — contracts, payroll, customer records, source code. On a model server that text travels between two machines in your own building, and the answer to an auditor asking where the data went is a rack number instead of a sub-processor list.

The model does not change under you

Hosted models get retired, retuned and rate-limited on somebody else’s schedule, and a routine approved last quarter can quietly start behaving differently this one. Weights you host stay byte-for-byte the same until you decide to replace them, and you pick the day.

At the end of it you own something

Four years of per-seat subscriptions leave you with cancelled invoices. Four years of a model server leave you with a machine that still runs, weights that still work without anyone’s permission, and hardware you can resell or put to another job.

The arithmetic

Where the line crosses

Round numbers, not a quote. They are here because the question comes up in the first five minutes of every conversation, and because the honest answer depends on how hard you run the thing.

€25

per person, per month

What a business plan on a hosted AI costs per seat. Ten people is €3,000 a year — before a single routine touches a metered API on top of it.

€10–12k

once, for up to ten people

The Studio build, delivered with the models installed and serving. After that it is roughly €400 a year of electricity, whether ten people hammer it all day or nobody signs in.

3–4 years

until the seats cost more

That is the payback on ten quiet seats. Routines running all day against a metered API turn years into months — that workload is where a model server stops being a preference and starts being the cheaper option.

We would rather lose the sale than sell the wrong machine. If your load is genuinely light, a hosted model is the cheaper answer, and that is what we will tell you.

Four builds

From a box on a shelf to a rack in a datacentre

The configurations below are indicative: they exist so you can see roughly what your headcount costs before anyone picks up the phone. The exact build follows a conversation about the models you need and the load you will put on them.

Start here

Studio

1× RTX PRO 6000 Blackwell Max-Q — 96 GB of VRAM

≈ €10–12k

  • UsersUp to 10 people
  • Power~0.6 kW under load

Platform

Tower chassis, 16-core Ryzen or Threadripper, 128 GB RAM, 2 TB NVMe, 10 GbE

Models

gpt-oss-120b at MXFP4, Qwen3-32B FP8 at full context, or Llama 3.3 70B at INT4 — with an embedding model alongside for search.

The one that starts the conversation: an ordinary wall socket, a cupboard with airflow, no rack and no server room. If ten people is where you are, this is the whole answer.

Business

4× RTX PRO 6000 Max-Q — 384 GB of VRAM

≈ €45–60k

  • Users50–150 people
  • Power~2 kW under load

Platform

2U/4U chassis, 32-core EPYC, 512 GB RAM, 4 TB NVMe, 25 GbE

Models

Qwen3-235B-A22B FP8, Llama 3.3 70B at full context, or several smaller models side by side — chat, embeddings and a coder at the same time.

Still fits an ordinary server room: no datacentre, no special cooling.

Enterprise

8× RTX PRO 6000 Server Edition — 768 GB of VRAM

≈ €100–130k

  • Users200–500 people
  • Power~4–5 kW under load

Platform

4U eight-GPU platform (Supermicro, Gigabyte), 2× EPYC, 1 TB RAM, 8 TB NVMe, 100 GbE

Models

DeepSeek V3.x and R1 at INT4, GLM-4.6 FP8, Qwen3-Coder-480B FP8.

Needs a rack with matching cooling. There is no NVLink, so tensor parallelism runs over PCIe 5: fine for inference, slow for fine-tuning.

Frontier

8× H200 141 GB HGX — 1128 GB of VRAM over NVLink

≈ €300–400k

  • Users1000+ people
  • Power~10–12 kW under load

Platform

2× Xeon or EPYC, 2 TB RAM, 400 GbE or InfiniBand; 8× B200 is the alternative

Models

DeepSeek V3.x FP8, Kimi K2 at INT4 (or FP8 with a limited KV cache), and full fine-tuning of 70B-class models.

Datacentre or colocation only, and lead times run into months.

What the price covers

In the price, and on your side

In the price

  • The hardware, assembled and tested under load before it ships.
  • Open-weight models installed, quantised and served behind one endpoint on your network.
  • Harness pointed at it, with your routines run against the model you chose.
  • A written handover: what runs where, how to restart it, how to add a model later.

On your side

  • Power, cooling and the space — a socket and a shelf for Studio, a rack with matching airflow for the rest.
  • Network access from the computers that will use it.
  • Day-to-day operation. Whoever keeps your other servers running keeps this one too.
  • There is no SLA and no remote administration in the price. Ask if you want either and we will quote it separately.

How it goes

Four steps, and only the first one needs you

  1. Tell us the load

    How many people, what the routines actually do, and whether anything has to stay on a segregated network. That is the whole brief — there is no discovery phase to pay for.

  2. Configuration and a fixed price

    We pick the GPUs for the models you need rather than the other way round, and you get one number with the build in it. If a smaller machine does the job, the quote is for the smaller machine.

  3. Built and measured

    The machine is assembled, the models are installed and served, and we run it at your kind of load before it leaves: tokens per second on the models you will actually use, not a benchmark sheet.

  4. Installed and connected

    It goes into your room, Harness is pointed at it, and your routines run against it end to end before we call the job done.

Tell us how many people and what they run

You get a configuration and one number back. No discovery workshop, no per-seat calculator, and no call to sit through before you see a price.