Mixedbread Quality Evals

How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.

How we measure

Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.

Wholembed V3, served through the Stores API, against other embedding models on public retrieval evals.

ViDoRe V3 (Markdown)

Open Introducing ViDoRe V3

Real-world, domain-specific document retrieval: ViDoRe V3 spans specialized industry corpora, evaluated on documents pre-parsed to markdown (English, French).

Per-domain breakdown

ModelCS (EN) Open CS (EN) datasetFinance (EN) Open Finance (EN) datasetPharma (EN) Open Pharma (EN) datasetHR (EN) Open HR (EN) datasetIndustrial (EN) Open Industrial (EN) datasetEnergy (FR) Open Energy (FR) datasetFinance (FR) Open Finance (FR) datasetPhysics (FR) Open Physics (FR) dataset
Mixedbread77.767.966.164.153.368.353.247.4
+ mxbai-rerank-v3.1-listwise86.673.474.673.663.975.459.252.7
Voyage 4 Large75.263.867.763.552.865.747.349.0
Cohere Embed 471.161.965.055.848.061.445.144.9
Qwen3 Embedding 8B73.756.661.953.644.760.237.145.9
Gemini Embedding 266.960.060.754.242.750.440.544.3
OpenAI text-embedding-3-large66.856.562.351.838.955.633.444.6
BM2564.749.956.949.645.657.435.939.8

NDCG@10 · higher is betterTested September 2026Introducing ViDoRe V3

Dataset and methodology

ViDoRe V3 covers 8 specialized industry corpora (computer science, finance, pharma, HR, industrial, energy, physics) in English and French, evaluated on documents pre-parsed to markdown.

We evaluate retrieval quality using publicly available eval datasets. Each model encodes the corpus and queries using its standard inference pipeline, then we compute the reported metric over the full test split. We report results as published by the eval or as measured on identical splits with default evaluation scripts.