03 · sovereign

From €7,500 for a starter deployment: one fine-tuned open model, running on hardware you own or a dedicated EU server registered to you — never a shared platform. Four to six weeks. Typical full deployments run €15,000–25,000 as node count, document sources and fine-tuning depth scale past the starter, and the data never crosses that line at either size.

What's included

  • One open model (Gemma or Mistral class), fine-tuned to your domain voice and format where the task needs it
  • A single node — your rack, or a dedicated EU server (Hetzner, OVH) registered to you, not to me
  • Retrieval over one document source, with the same acceptance-test discipline as the pilot
  • A standard open stack — llama.cpp or vLLM, pgvector or Qdrant — nothing proprietary between you and your system
  • Deployment without your production data: the build, the fine-tune and the acceptance test run on synthetic or anonymised material you prepare, and your own staff load the live corpus following the runbook
  • Full documentation, runbook, handover pack, and a 30-day post-launch warranty
  • Compliance paperwork support: DPIA input and a processing-records template for your DPO

The safe first step

Not sure self-hosting is even right for you? Start with the €1,400 readiness assessment — one week, a written architecture recommendation, credited in full against a deployment. If an EU-hosted API with a DPA genuinely covers your case, the assessment says exactly that, and you’ve just saved €6,100.

monthly

Sovereign Care €1,500–3,000/month

  • Monitoring, security patches, and model updates as the open-weights frontier moves
  • Compliance upkeep: the paperwork stays current as regulation shifts
  • Response within one business day in EU hours

Cancel any month. Everything is documented and standard — your IT team can take over whenever you choose. That is the design, not a concession. Where support needs access to your live system, that access is named, time-boxed, logged on your side and revoked when the ticket closes — agreed in writing before it exists, never as a standing key.

How engagements work

  • Fixed price, milestone payments — typically 30/40/30 — your prepayment never exceeds one milestone.
  • Payment by bank transfer against an invoice — the account details and the currency are on the invoice. Invoices carry no VAT: EU and UK clients self-account for it under the reverse charge, other jurisdictions owe none.
  • Contract and NDA before any work beyond the audit — yours or mine, with an English-law option.
  • Registered sole proprietorship in Armenia since February 2026. GMT+4: my afternoon is your morning, every working day.

how the work is split

Deployment and support are two separate contracts

They are papered separately because they are not the same work, and because the second one can need something the first one never asks for: access.

Deployment

The training and inference pipeline is delivered as code and runs on your infrastructure. The model is trained on your side and the weights stay there — nothing is exported to me. I build and tune against synthetic or de-identified material that you approve, and I do not ask for accounts in your production systems. Day-to-day operation is handed to your own staff, with a written runbook.

Support

A separate agreement, taken or declined on its own merits, because this is where access becomes a question again. Two ways to run it: your operator executes the runbook and I never touch production data; or you open time-boxed break-glass access on request — granted per ticket, logged into your own infrastructure, and covered by its own set of documents.

One more thing a hosting line usually hides: where the server is rented in your name, the contract with the provider is yours and I am not a party to it. In neither arrangement is your workload placed on a platform shared with my other clients.

Questions

What people ask

  • Do we have to buy servers?

    Usually not. In practice, self-hosting for a European SME means a machine you control: your existing rack if you already have one, or a dedicated EU server (Hetzner, OVH) rented in YOUR name for a few hundred euros a month. What matters for compliance is that the hardware, the model and the data are yours, in a jurisdiction you chose — not whose basement the machine happens to sit in. Proof that this runs on modest hardware: the assistant on this very site, self-hosted on a 4-core ARM box.

  • Who patches it once it is running?

    You can, or I can — and it is worth keeping the two apart, because they are two different agreements. The deployment ends with your team holding every key: full documentation, a runbook and a handover pack exist precisely so your own IT is never locked out, and I need no access to your live system to finish the job. If you would rather not carry the upkeep yourself, the optional Sovereign Care retainer (€1,500–3,000/month) covers monitoring and security patches — and that is where access gets defined in writing before it exists: named, time-boxed, logged on your side, revoked when the ticket closes. Cancel any month: the whole point of the handover pack is that leaving is always possible.

  • What if the model goes stale — will a small open model keep up?

    The stack is standard and swappable by design — llama.cpp or vLLM, nothing proprietary — so moving to a stronger open model later is a deployment task, not a rebuild. For narrow-domain work — answering from your documents, in your formats, in your terminology — a fine-tuned small model already performs at a level you used to pay GPT-4 prices for; when a task genuinely needs frontier reasoning, the honest answer is a hybrid: sensitive work stays local, while that one task goes to an EU-hosted API. The Sovereign Care retainer tracks open-weight releases and applies upgrades once they clear your acceptance test; without it, the system keeps running exactly as delivered, just without the upgrades.

  • What happens when you are not available?

    You get a written response within one business day, in EU hours (the engineer works from GMT+4) — the honest promise one engineer can actually keep, not a nonstop-availability guarantee that no studio of this size could actually keep. It matters less than it sounds: the deployment is documented, standard and self-hosted on infrastructure you own, so your own IT can operate and even patch it without me. That is a design requirement, not a courtesy — the same rule the RAG pilot is built to.

  • Why not just use Azure OpenAI with a DPA? Our lawyers accept it.

    For many workloads you should — and if the readiness assessment says an EU-hosted API covers your case, that is what the report will say. The honest reasons to self-host: the US CLOUD Act reaches US providers regardless of the DPA; hosted APIs carry abuse-monitoring and telemetry caveats; per-token pricing turns success into a variable cost; and a hosted model can be deprecated out from under your workflows. Self-hosting trades convenience for jurisdiction, fixed cost, and a model nobody can take away.

Ready when you are

Write two sentences about what you're trying to do. You'll get a straight answer — including "you don't need me for this" when that's the truth.

Contract & NDA before any work beyond the audit · milestone payments, typically 30/40/30 · 30-day warranty · registered business, Armenia · GMT+4, EU hours · Trust & Process