Skip to content

Local models

Use the models
without handing
any data over.

souveraen.ai runs its language models entirely on your own hardware: open weights, pinned versions, no external model API. No prompt and no document leaves your building.

  • Open-weight models
  • No external model API
  • A fixed monthly price
Open-weight models sit locked inside an appliance in your own building; no path leads outside

How a model reaches the appliance

From the shortlist
to the first working day.

No model is fetched from the internet. Every one takes the same path first.

Which models those are
  1. 1

    Selection

    Only models with open weights and a licence that allows commercial use.

  2. 2

    Qualification

    Quality, speed and memory use are measured on the reference appliance.

  3. 3

    Pinning

    Version and revision are recorded. Exactly the reviewed build is what runs.

  4. 4

    Delivery

    Model builds arrive on the appliance as reviewed packages, including with no internet access.

  5. 5

    Operation

    The routing layer assigns every task to the matching model and logs the call.

A new model version replaces the old one only after it has taken the same path. Nothing is fetched automatically at runtime.

Why local models

What you gain when the model sits in your building.

The weights are open, the versions are pinned, and the bill does not follow how many questions you ask.

Compare operating models and prices

Data sovereignty

Every model runs on the appliance in your building. With no external model API, no prompt and no document leaves your network.

Under your control

Every model version is pinned and delivered after review. What answers today answers the same way tomorrow, until you approve an update.

Costs you can plan

A fixed monthly price instead of billing per token. Heavy use does not make the platform more expensive.

Operation without the internet

In air-gap operation the platform works with no internet connection at all. Updates arrive as reviewed offline packages.

One way in, several models

Your applications do not pick a model.

Applications talk to one stable interface. Behind it the routing layer assigns every task to the pinned model. Chat, search, reranking and the safety check run separately, but on the same appliance.

  • A stable interface
  • Assignment is configured
  • Logged on every request
Chat, search and the safety check run through one shared routing layer onto several pinned models inside a sealed appliance

Models on the appliance

Every task has its own fixed model.

Each model has a clearly bounded task and a licence that allows commercial use. Every model sits on the appliance in your building.

ModelTaskLicence
Qwen3-8BChat and reasoning in the Appliance/Standard profile; also rephrasing and classification.Apache-2.0
Qwen3-32B (FP8)Chat in the quality profile: higher answer quality with fewer concurrent users. We approve the profile together for your installation.Apache-2.0
BGE-M3Embeddings for vector search in the hybrid index.MIT
BGE Reranker v2-M3Local reranking of the passages before an answer is composed.MIT
Qwen3-0.6B (GGUF)Safety checks and short classification as a separate CPU process beside the chat model.Apache-2.0
Further open modelsEuropean models for tenders, for example: we review the licence and the quality and add the model to your profile.per model
  • Every model is qualified on the reference appliance, measured and delivered as a reviewed package.
  • Which profile is active in your installation we approve together; the assignment of task to model stays documented.
  • Licence, quality and effort for further open models we confirm in the offer.

The exact model set depends on the profile you choose. Which models sit on your appliance we record in the offer.

Profiles and task classes

Which profile for which task.

The assignment of task to model is configured and can be read back on every request. Which profile runs in your installation decides how many people can have answers composed at once.

ProfileModel for chatWhat it is built for
Appliance/Standard profileQwen3-8BChat and reasoning for many concurrent users. This profile runs on the reference appliance.
Quality profileQwen3-32B in FP8 quantisationHigher answer quality for fewer concurrent users. Which profile fits we settle in the offer.

What works on your appliance

Answers with sources

In the Appliance/Standard profile Qwen3-8B handles chat and reasoning, with sources from your documents.

  • Qwen3-8B
  • Chat
  • Reasoning

Search and reranking

BGE-M3 and the BGE reranker carry vector search and the reranking of passages in the hybrid index.

  • BGE-M3
  • BGE Reranker v2-M3

A separate safety check

A separate CPU process checks prompts apart from the chat model, without tying up time on the GPU.

  • Qwen3-0.6B
  • Separate CPU process

Task classes kept apart

Chat, reasoning, classification, rephrasing and the safety check are kept apart. Model and version are logged on every request.

  • Configured
  • Logged on every request

Open weights, reviewed licence

Further open-weight models we add to your profile after a licence and quality review, including European models for tenders.

  • Licence review
  • European models

Every task in the building

Tasks such as transcription and translation also run on the same appliance. There is no path to an external service.

  • Transcription
  • Translation

A matched set of models carries the answer, the search, the reranking and the safety check. Which models sit on your appliance we record in the offer.

The hardware

NVIDIA DGX Spark, the compute unit in the building.

Private Spark delivers the platform on a compact compute unit that moves into your building. You do not need a data centre of your own for that.

Private Spark in detail
NVIDIA DGX Spark, the compute unit of Private Spark
NVIDIA DGX Spark – the compute unit of Private Spark.
  • NVIDIA GB10 (Grace Blackwell) with 128 GB of unified memory for CPU and GPU, enough for the chat model, embeddings and reranking at the same time.
  • A compact device for an office or a server room. It is connected on your network, behind your firewall.
  • For higher demand or fully separated networks we size Air-Gap Enterprise in the offer.

Common questions

What organisations most often get stuck on before they buy.

Which model answers our questions?

In the Appliance/Standard profile it is Qwen3-8B, an open-weight model under the Apache-2.0 licence. The quality profile uses Qwen3-32B in FP8 quantisation. Which models sit on your appliance depends on the profile and is recorded in the offer.

Are our data used to train the models?

No. The models run as pinned versions and are not further trained on your content. What the platform learns from your documents sits as sourced knowledge in your knowledge layer, not in model weights; that knowledge can be checked and deleted.

Do the models need an internet connection?

No. Every model sits entirely on the appliance. In air-gap operation the platform works with no internet connection at all; updates arrive as reviewed offline packages.

What does a single request cost us?

Nothing extra. There is no billing per token: you pay a fixed monthly price for the appliance, regardless of how heavily your team uses it.

How do new model versions reach the appliance?

As reviewed packages. A new version takes the same qualification as the first model, is pinned and only then replaces the old build after delivery. Nothing is fetched automatically from the internet.

Can we connect an external model as well?

On Private Spark and Air-Gap Enterprise, deliberately not: there is no path to an external model API. Anyone who prefers a hosted offering has European On Demand as a separate operating model; the two stay strictly apart.

Local models

How many people will work with it at once?

From that number and from what you require of the data, it follows which model profile belongs on your appliance. Tell us both and we will work out the profile and the cost.