Mixedbread Quality Evals
How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
How we measure
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
Wholembed V3, served through the Stores API, against other embedding models on public retrieval evals.
Miracl-Vision
Open MIRACL on HuggingFaceMultilingual visual document retrieval across 18 languages.
Per-language breakdown
| Model | Arabic | Bengali | Chinese | English | Farsi | Finnish | French | German | Hindi | Indonesian | Japanese | Korean | Russian | Spanish | Swahili | Telugu | Thai | Yoruba |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mixedbread | 90.7 | 91.1 | 79.1 | 79.5 | 81.4 | 92.0 | 87.9 | 80.1 | 82.0 | 73.7 | 89.5 | 78.6 | 88.6 | 78.1 | 81.3 | 92.7 | 90.4 | 92.0 |
| Gemini Embedding 2 | 83.2 | 75.8 | 68.0 | 71.7 | 72.7 | 87.2 | 79.2 | 71.2 | 70.7 | 62.0 | 84.3 | 70.4 | 77.3 | 70.7 | 77.0 | 50.0 | 79.7 | 87.5 |
| Voyage Multimodal 3.5 | 78.1 | 53.9 | 69.7 | 72.6 | 57.6 | 88.0 | 78.8 | 75.5 | 42.5 | 64.9 | 82.9 | 66.9 | 81.8 | 72.5 | 74.0 | 15.1 | 58.0 | 84.0 |
| Qwen3-VL Embedding 8B | 61.6 | 55.2 | 62.8 | 71.5 | 42.5 | 78.5 | 72.6 | 65.5 | 49.3 | 55.9 | 70.5 | 51.5 | 75.1 | 63.5 | 59.9 | 17.6 | 47.5 | 75.3 |
| GME Qwen2-VL 2B | 48.9 | 37.5 | 63.1 | 67.8 | 30.9 | 68.6 | 68.5 | 63.5 | 31.3 | 54.2 | 73.0 | 62.0 | 72.0 | 62.8 | 53.5 | 8.9 | 35.6 | 48.8 |
NDCG@10 · higher is betterTested March 2026MIRACL on HuggingFace
Dataset and methodology
Miracl-Vision is a multilingual visual document retrieval eval covering 18 languages, including several low-resource ones. We report the published score for gme-Qwen2-VL-2B-Instruct as the prior paper SOTA baseline.
We evaluate retrieval quality using publicly available eval datasets. Each model encodes the corpus and queries using its standard inference pipeline, then we compute the reported metric over the full test split. We report results as published by the eval or as measured on identical splits with default evaluation scripts.