Alpha
All entriesContact
Model

Soofi

Publisher
Soofi Consortium (Sovereign Open Source Foundation Models) (DE)
Family
Soofi-S
Openness
open_weights
Licence
Permissive open-source committed (OSAID 1.0) but not yet finalised - SPDX pending, LICENSE access-restricted
Context
up to 256k

The documentation and openness intent here are exceptional: a genuine pretraining report with full per-source data accounting and a stated OSAID 1.0 commitment to release weights, checkpoints, code and data permissively, and the German-sovereign infrastructure is a real data-control draw. This is preview-stage: the weights are 'coming soon' rather than downloadable, and the licence is unconfirmed (page inconsistent, LICENSE file 401'd, no final SPDX), so use/modify and data-control cannot yet be fully exercised. Treat the openness as promised and well-documented rather than delivered.

Do you really own it?
Partial
none·limited·partial·substantial·full
Analytical input: AOI B · 73.2/100
The four ownership factors

Floor-weighted, not averaged. Nothing is weak, but use & modify and data control are only moderate, so it misses the bar for substantial - strong on both use & modify and data control - and lands at partial.

1

Use and modify freelyCan you run, modify and adapt it with no gate and no field-of-use trap?

Moderate

The OSAID-compliant permissive licence is a stated future commitment, not yet in force: the licence is unconfirmed/pending (page inconsistent, LICENSE file 401'd, no final SPDX) and the model is a preview whose weights are 'coming soon' rather than generally downloadable - so unconditional commercial use, fine-tuning and continued pre-training cannot yet be fully exercised.

How this scores (AOI sub-dimensions)
Openness3/5how much is released - weights, data, code, licence - and how freelyOpenness here is documented and committed but not yet delivered, so it earns the open-weights ceiling (3) instead of the fully-open 5.
Legal4/5how permissive and clean the licence is for real commercial useCommitted permissive OSAID-compliant licensing with unconditional commercial use, an EU/Germany jurisdiction, well under the systemic-risk compute threshold, and a published per-source data inventory that makes the EU AI Act training-content summary directly satisfiable.
2

TransparencyDo you know what it is: weights, training, behaviour, and legible terms?

Strong

Openness rests on a real, detailed pretraining technical report with full per-source data accounting plus a stated OSAID 1.0 commitment - you can already see what it is and how it is built - even though the full weights/checkpoints/data release is a future intention, not yet delivered.

How this scores (AOI sub-dimensions)
Provenance4/5how well we can trace and verify what went into the modelDistributed from the verified Soofi-Project org on Hugging Face in bf16 safetensors (plus first-party GGUF/3-bit quants) with checksums, and because the data accounting, code and intermediate checkpoints are public the pipeline is highly auditable (checklist 6/8).
Governance4/5how accountable and well-documented the publisher isAccountable, well-identified publisher: a named government-funded consortium coordinated by the KI Bundesverband with reputable research members, a detailed pretraining technical report and open GitHub repositories.
3

ReliabilityIs it reliable and good enough for the job?

Moderate

Pre-release: the weights are not yet downloadable, so runtime dependability is unproven and the documented-incomplete safety cannot yet be exercised. The reported ~3B-active-param strength counts as a capability point in the AOI score rather than an ownership factor.

How this scores (AOI sub-dimensions)
Operational4/5how practical it is to run, serve and maintain in productionRuns on mainstream stacks - HF transformers (trust_remote_code), vLLM, llama.cpp/llama-server and Ollama - with first-party GGUF and 3-bit quantizations and a near-constant long-context cache.
Safety3/5whether misuse risks are evaluated and guardrails are providedA post-trained Instruct-Preview variant exists and the L3S Research Centre contributes safety/evaluation frameworks, but the Instruct model card documents several privacy and safety features as still incomplete, this is an early preview, and there is no companion guard model or broad independent red-team battery.
4

Doesn't extract your dataDoes running it keep your knowledge and data yours?

Moderate

EU/Germany-domiciled and trained on sovereign German infrastructure, so once self-hosted your data would stay yours - but with weights not yet downloadable and the licence unconfirmed, that self-hosting cannot yet actually be exercised from a settled grant.

How this scores
Not a scored AOI dimension. For a self-hosted model, data-control is a structural property of running the weights yourself, strong by default unless the model phones home or the licence claws back rights. For a hosted API this factor is the retention + train-on-inputs + residency read, scored from the binding terms.

How the AOI score is computed

The seven dimensions above, each scored 0 to 5, weighted and summed to the 0 to 100 headline. The score is the analytical input behind the ownership verdict, not the verdict itself.

DimensionScoreWeightPoints
Openness3/50.1810.8
Provenance4/50.1612.8
Legal4/50.1612.8
Safety3/50.169.6
Performance4/50.1411.2
Operational4/50.129.6
Governance4/50.086.4
HeadlineB · 73.2/100
Dossier coverageAssess 87%Implement 91%Use 72%Support 50%How complete our four-domain documentation is, a measure of our coverage, not of the model. Each domain links to its page.

Sources

Every rating traces to a primary document. Read means the text was verified; unverified means it is known to exist but has not yet been read.

DocumentWhat it grounds
Technical_reportread2026-07-25
Real arXiv report ("A Sovereign, Open-Source Foundation Model for German and English" / PDF "Soofi S Pretraining Report v1.0"): a ~30B (30B-A3B) hybrid Mamba-2/Transformer MoE trained on ~27T primarily German+English tokens on Deutsche Telekom's Industrial AI Cloud in Munich, with full per-source data accounting (public identifier + raw/effective token counts; the ~1.3% Genios corpus reported in aggregate).
Licenceunverified2026-07-25
Soofi-S-Base exists and is hosted on the verified Soofi-Project org on Hugging Face in bf16 safetensors (custom hybrid Mamba-2/MoE modelling code via trust_remote_code=True), but its licence is unconfirmed/pending: the page characterises it inconsistently ("Other"/TODO vs "closed-beta" and "will be released under a permissive license, without gated access"), terms are described as "not yet finalized", the LICENSE file returned HTTP 401, and no final SPDX has been published.
Model cardunverified2026-07-25
The Soofi-Project Hugging Face org publishes the base plus Instruct/Isar/Rhine preview variants with first-party GGUF and 3-bit quantizations, embedded Jinja chat/tool templates, and model cards noting safety features as incomplete.
Documentationread2026-07-25
The public soofi-project/Soofi-Pretraining GitHub repo (based on the Nvidia Nemotron 3 Nano architecture) hosts the training/evaluation code and per-source data-construction scripts, but its README states "Open model weights coming soon" - the model is not yet generally downloadable (Hugging Face currently hosts a preview/internal checkpoint).
Vendor announcementunverified2026-07-25
A consortium press release presents Soofi as a German public-private consortium project for sovereign industrial AI, trained on Deutsche Telekom's Industrial AI Cloud in Munich, coordinated by the KI Bundesverband and funded by the BMWE.
Third-party analysisunverified2026-07-25
CAIRNE's launch page describes the SOOFI consortium, its members (including CAIRNE Gold member L3S Research Centre) and its sovereign open-source foundation-model mission.
Third-party analysisunverified2026-07-25
Independent technology coverage reports Soofi-S as a leading fully-open model that meets the Open Source AI Definition, publishing weights, intermediate checkpoints, training/eval code and a detailed data inventory.
Third-party analysisunverified2026-07-25
Independent coverage reports Soofi-S is the strongest fully-open model on aggregate English and German benchmarks, ahead of OLMo 3 32B, Apertus 70B, EuroLLM 22B and Alia 40B, and winning code aggregates among open base models.