most-embed-de: State-of-the-art German Embeddings for RAG and Semantic Search

Community Article
Published September 25, 2026

most-embed-de, a German embedding model for retrieval and customer-support RAG

most-embed-de is a 1.1B-parameter German embedding model for RAG, semantic search, and customer-support retrieval. It is the state-of-the-art among models up to 1.5B parameters on the German RTEB subset, and first among entries with results on all four MTEB German retrieval tasks in the snapshot below.

The model is optimized for retrieval application in German company documentation: help-center search, FAQ retrieval, and the retrieval stage of a support assistant. It scored strong results on various industries and domains, like legal, medical, e-commerce, and banking. The weights are freely available on Hugging Face, with Sentence Transformers support, 2,048-dimensional embeddings, and a context window of up to 32,768 tokens.

Get the model →

A German embedding model for the way customers search

Consider three ways someone might look for the same help-center article:

  • Keyword search: "retoure ohne drucker"
  • Question: "Wie kann ich etwas zurückschicken, wenn ich keinen Drucker habe?"
  • Problem description: "Ich kann das Rücksendeetikett nicht ausdrucken."

The relevant passage might explain how to use a mobile return code at a parcel shop. Matching these requests to the right instructions is the job of the embedding model. It represents the query and each document as vectors, allowing a search system to retrieve passages with related meanings. In a RAG application, those passages then become context for the large language model (LLM) that writes the answer. Sentence Transformers explains this query-to-document search setup here.

For German customer support, the practical question is how well that matching works across your terminology, document types, and customers' phrasing. That is the focus of most-embed-de, a fine-tune of NVIDIA's Nemotron-3-Embed-1B-BF16.

SupportIR-DE: evaluating German customer-support retrieval

Public benchmarks give us useful external comparisons - but they are well align with real-world use case. Therefore, we measure the deployment most-embed-de was built for: searching one company’s support documentation.

For that, we developed SupportIR-DE, an benchmark using German help-center content sourced from public web pages through Common Crawl. It covers telecommunications, banking and payments, insurance, and e-commerce, with questions, keyword searches, and problem descriptions.

SupportIR-DE customer-support retrieval benchmark: most-embed-de scores 0.848 overall across e-commerce, insurance, banking, and telecommunications

Company-scoped nDCG@10, aggregated over five company folds. Original model-card figure; its vertical axis is truncated to make differences visible.

Each query is searched against its own company’s documents. The company restriction comes from the evaluation setup, just as an application should enforce which knowledge base a user can access. The model’s task is to find the right passage within that collection.

The release comparison reports 0.848 for most-embed-de, versus 0.835 for its Nemotron-1B base and 0.838 for Qwen3-Embedding-8B. Nemotron-8B leads at 0.869. most-embed-de has the highest overall score among the evaluated models up to 1.5B parameters. Benchmark design and results are documented in the model card.

MTEB German retrieval: #1 with 60.86 across the four official tasks

On MTEB(deu, v1), most-embed-de scores 60.86 in the Retrieval category. It ranks first among the 110 entries with complete results for that category in the 7 September snapshot.

Model Total parameters MTEB German Retrieval
most-embed-de 1.14B 60.86
F2LLM-v2-14B 13.99B 59.52
NVIDIA Nemotron-3-Embed-1B 1.14B 59.10
F2LLM-v2-8B 7.57B 59.09
inf-retriever-v1 7.07B 58.93
Snowflake arctic-embed-l-v2.0 0.57B 57.09

This category averages GermanQuAD-Retrieval, GermanDPR, XMarket in German, and the full GerDaLIR task. Scores are multiplied by 100; GermanQuAD uses MRR@5 and the other three use nDCG@10. The result is a retrieval-category ranking, not an overall ranking across MTEB German’s 19 tasks. Models without all four retrieval scores are excluded. Source: MTEB German per-task and category results.

RTEB German results: first in the 1B size class

The RTEB German benchmark covers legal documents, healthcare retrieval, and business dialogue. RTEB relies on private test set, reducing the risk of overfitting. The published results now include most-embed-de on all four tasks.

German embedding model comparison on RTEB: most-embed-de scores 82.38, ahead of its 1B base at 80.65 and Qwen3-Embedding-8B at 80.30

Four-task mean nDCG@10 × 100, calculated from the published per-task scores. Selected comparators; higher is better. Parameter counts are total parameters reported by the current leaderboard.

Model Total parameters RTEB German four-task mean
NVIDIA Nemotron-3-Embed-8B 7.95B 85.33
Voyage 4 Large, 2,048 dimensions Undisclosed 84.99
most-embed-de 1.14B 82.38
Octen-Embedding-8B 7.57B 82.25
NVIDIA Nemotron-3-Embed-1B 1.14B 80.65
Qwen3-Embedding-8B 7.57B 80.30
Snowflake arctic-embed-l-v2.0 0.57B 76.77
multilingual-e5-large-instruct 0.56B 71.34
BGE-M3 0.57B 69.04

The full comparison contains 79 entries with all four task scores. Among entries with disclosed total parameter counts of up to 1.5B, most-embed-de has the highest average. It improves on its Nemotron-1B base by 1.73 points and exceeds Qwen3-Embedding-8B by 2.08 points, with roughly 6.6 times fewer parameters. Larger models still lead the overall comparison but at significantly more computational costs. Source: the RTEB German leaderboard’s per-task data.

The gains vary by task. Relative to the base, most-embed-de improves healthcare retrieval from 78.10 to 82.96 and LegalQuAD from 73.15 to 76.17. The average captures a useful improvement, while the individual results show where specialization helps and where it costs something.

How this average is calculated: each model must have scores for LegalQuAD, German1Retrieval, GermanHealthcare1Retrieval, and GermanLegal1Retrieval. The scores are aggregated as the unweighted mean. 82.38 is the four-task average calculated here from published measurements, including the three private tasks; it is not the value currently displayed in that summary column.

Use most-embed-de with Sentence Transformers

Here is a small German semantic-search example. Install a current Sentence Transformers release with support for the Nemotron base architecture:

pip install -U sentence-transformers transformers
import torch
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "malteos/most-embed-de",
    device="cuda",
    model_kwargs={"dtype": torch.bfloat16},
)
model.max_seq_length = 4096

documents = [
    "Für eine Rücksendung ohne Drucker nutzen Sie den mobilen Retourencode. "
    "Zeigen Sie den QR-Code im Paketshop vor; dort wird das Etikett gedruckt.",
    "Ihre Lieferadresse können Sie vor dem Versand im Kundenkonto ändern.",
    "Die Rechnung steht nach dem Versand als PDF im Kundenkonto bereit.",
]
query = "Ich möchte etwas zurückschicken, habe aber keinen Drucker."

document_embeddings = model.encode_document(
    documents, normalize_embeddings=True, convert_to_tensor=True,
)
query_embeddings = model.encode_query(
    [query], normalize_embeddings=True, convert_to_tensor=True,
)

scores = model.similarity(query_embeddings, document_embeddings)[0]
for index in scores.argsort(descending=True).tolist():
    print(f"{scores[index].item():.3f}  {documents[index]}")

encode_query() and encode_document() apply the model’s saved query: and passage: prompts automatically. Avoid adding those prefixes a second time. This example assumes a CUDA GPU with BF16 support.

For a larger collection, precompute document embeddings and store the 2,048-dimensional vectors in your vector index. At query time, encode the question and retrieve relevant passages.

Is most-embed-de the right German embedding model for your RAG system?

If you are comparing German embedding models for RAG, most-embed-de offers a concrete starting point: strong German retrieval results at 1.1B parameters, a support-focused evaluation, and downloadable weights for local deployment.

For FAQ search or a customer-support assistant, test it against your current retriever using actual queries and the documents users are allowed to search. Include short searches, complete questions, and descriptions of problems. Measure retrieval quality alongside latency and memory at realistic document lengths.

Try most-embed-de on Hugging Face

Community

Sign up or log in to comment