Skip to content
souveraen.ai

Models

Every model in one platform.

Open models such as Qwen, Mistral and gpt-oss run on your appliance, fully offline if you wish. Frontier models from OpenAI and Anthropic are yours EU hosted when your plan includes them.

  • Local or fully offline
  • Open and frontier models
  • One interface for every task

Model catalogue

Mistral AI8
  • Mistral Large 3

    Context window
    256,000 tokens
    Max output
    –
    ChatEU hosted
  • Mistral Medium 3.5

    Context window
    256,000 tokens
    Max output
    –
    ChatEU hosted
  • Mistral Small 4

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable offline
  • Ministral 3 8B

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable offline
  • Devstral Small 2

    Context window
    256,000 tokens
    Max output
    –
    CodeAvailable offline
  • Magistral Small 1.2

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable offline
  • Codestral

    Context window
    256,000 tokens
    Max output
    –
    CodeEU hosted
  • Mistral Embed

    Context window
    –
    Max output
    –
    EmbeddingEU hosted
Qwen (Alibaba)7
  • Qwen3.6-35B-A3B

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable offline
  • Qwen3.6-27B

    Context window
    256,000 tokens
    Max output
    –
    ChatAvailable offline
  • Qwen3-Coder-30B-A3B

    Context window
    262,144 tokens
    Max output
    –
    CodeAvailable offline
  • Qwen3-30B-A3B-2507

    Context window
    262,144 tokens
    Max output
    –
    ChatAvailable offline
  • Qwen3-14B

    Context window
    32,768 tokens
    Max output
    –
    ChatAvailable offline
  • Qwen3-8B

    Context window
    32,768 tokens
    Max output
    –
    ChatAvailable offline
  • Qwen3 Embedding 8B

    Context window
    –
    Max output
    –
    EmbeddingAvailable offline
DeepSeek2
  • DeepSeek V4-Pro

    Context window
    1,000,000 tokens
    Max output
    –
    ChatGlobal
  • DeepSeek V4-Flash

    Context window
    1,000,000 tokens
    Max output
    –
    ChatGlobal
Anthropic7
  • Claude Fable 5.1

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • Claude Fable 5

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • Claude Opus 5.5

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • Claude Opus 5

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • Claude Opus 4.8

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • Claude Sonnet 5

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • Claude Haiku 4.5

    Context window
    200,000 tokens
    Max output
    64,000 tokens
    ChatEU hosted
Google6
  • Gemini 3.1 Pro

    Context window
    1,000,000 tokens
    Max output
    64,000 tokens
    ChatGlobal
  • Gemini 3.5 Flash

    Context window
    1,000,000 tokens
    Max output
    64,000 tokens
    ChatGlobal
  • Gemini 3.1 Flash-Lite

    Context window
    1,000,000 tokens
    Max output
    64,000 tokens
    ChatGlobal
  • Gemini Embedding 2

    Context window
    –
    Max output
    –
    EmbeddingGlobal
  • Gemma 4

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable offline
  • Gemma 3 27B

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable offline
OpenAI11
  • GPT-5.6 Sol

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT-5.6 Terra

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT-5.6 Luna

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT-5.5

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT-5.4

    Context window
    1,050,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT-5.4 mini

    Context window
    400,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT-5.1

    Context window
    400,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT-5

    Context window
    400,000 tokens
    Max output
    128,000 tokens
    ChatEU hosted
  • GPT Image 2

    Context window
    –
    Max output
    –
    ImageEU hosted
  • gpt-oss-120b

    Context window
    131,072 tokens
    Max output
    131,072 tokens
    ChatAvailable offline
  • gpt-oss-20b

    Context window
    131,072 tokens
    Max output
    131,072 tokens
    ChatAvailable offline
Meta3
  • Llama 4 Scout

    Context window
    10,000,000 tokens
    Max output
    –
    ChatAvailable offline
  • Llama 3.3 70B

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable offline
  • Llama 3.1 8B

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable offline
Microsoft2
  • Phi-4

    Context window
    16,384 tokens
    Max output
    –
    ChatAvailable offline
  • Phi-4-reasoning-plus

    Context window
    32,768 tokens
    Max output
    –
    ChatAvailable offline
NVIDIA2
  • Llama 3.3 Nemotron Super 49B

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable offline
  • Nemotron Nano 9B v2

    Context window
    128,000 tokens
    Max output
    –
    ChatAvailable offline
Z.ai (GLM)1
  • GLM-5.2

    Context window
    1,000,000 tokens
    Max output
    128,000 tokens
    ChatGlobal
Moonshot AI (Kimi)1
  • Kimi K2.6

    Context window
    256,000 tokens
    Max output
    –
    ChatGlobal
MiniMax1
  • MiniMax-M3

    Context window
    1,000,000 tokens
    Max output
    –
    ChatGlobal

51 of 51 models

Models marked available offline also run on your appliance, without internet. OpenAI, Anthropic and Mistral AI compute in EU tenants; global models only once approved per tenant. Which models you can use is set by your plan and your offer.

Missing a model?

We add further open models to your profile after a licence and quality review. Tell us which model you need.

Request a model

The right model per task

Chat, code, embedding and image: for each task you pick the model that fits – open or frontier, from twelve providers.

Secure and sovereign

Run on your appliance, fully offline if you wish, or in European data centres.

Easy to integrate

Your applications talk to one stable interface. Which model answers is configured and logged on every request.

Transparent

Context window, maximum output and operating route are out in the open. An open model only changes its build once you approve an update.

Operating routes

First the data path, then the model.

A model name says nothing about where your request is processed. So you choose the operating route first; the models that fit follow from it.

Local on the appliance

Open models such as Qwen3.6, Mistral Small 4, Gemma 4, Llama and gpt-oss run on the DGX appliance in your building. No external model interface.

  • Works without the internet
  • No token costs
  • A fixed monthly price
See Private Spark

European On Demand

The same platform in European data centres, with no hardware of your own. Plus models from OpenAI, Anthropic and Mistral AI in EU tenants.

  • GDPR data processing agreement
  • A separate tenant
  • Start for free
Plans and pricing

Global models

Models that run neither offline nor in the EU, such as Gemini, DeepSeek, GLM, Kimi and MiniMax. Once your administrator has enabled them, your users decide for themselves whether to use them.

  • Enabled by your administrator
  • Users decide for themselves
  • Revocable at any time
Discuss an exception

Common questions

What people ask most before they choose a model.

Which model answers our questions?

You decide. On the appliance an open model such as Qwen3.6 or gpt-oss-120b does the work; which one depends on your tasks and how many people work at once, and is recorded in the offer. In European On Demand you can add frontier models.

Are our data used to train the models?

No. The models run as pinned versions and are not further trained on your content. What the platform learns from your documents sits as sourced knowledge in your knowledge layer, not in model weights.

Can we add a model that is not in the catalogue?

Yes, if it has open weights and a licence that allows commercial use. We review licence, quality and memory needs on the reference appliance and then add it to your profile.

Can we use GPT, Claude or Gemini models?

In European On Demand, yes, depending on the plan. OpenAI and Anthropic models run in the vendors' EU tenants; what goes there is the single request with selected excerpts, not your collection. Gemini computes globally; once your administrator has enabled it, users decide for themselves whether to use it. On the appliance, frontier models are an approved exception.

How do new model versions reach the appliance?

As reviewed packages. A new version goes through the same qualification as the first model and only replaces the old build once you approve it. Nothing is fetched automatically.

The right profile

Which models belong on your appliance?

Tell us your tasks and how many people work at once. We propose the model profile and record it in the offer.