Mixedbread Quality Evals
How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
How we measure
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
Retrieval
Wholembed V3, served through the Stores API, against other embedding models on public retrieval evals. Evaluation code is reproducible and available here.
- ViDoRe V3 (Markdown)Domain-specific document retrieval over eight industry corpora, in markdown (English, French).62.25NDCG@10 · #1 of 7
- ViDoRe V3 Crosslingual (Markdown)The same corpora in markdown, queried cross-lingually (de, en, es, fr, it, pt).60.14NDCG@10 · #1 of 6
- ViDoRe V3 (Images)The same corpora retrieved from page images instead of parsed text.64.24NDCG@10 · #1 of 5
- Miracl-VisionMultilingual visual document retrieval across 18 languages.85.22NDCG@10 · #1 of 5
- AudioAudio retrieval with graded relevance judgements.71.17NDCG@10 · #1 of 6
- VideoText-to-video retrieval: general clips, localized moments and instructional videos.66.97NDCG@10 · #1 of 4
Agents on Mixedbread
Third-party agents with Mixedbread Search as their retrieval tool, against other retrievers. Higher accuracy with fewer tool calls is better. Read more in Closing the Oracle Gap.
- BrowseComp-PlusA deep-research agent over a 100k-document web corpus, with Mixedbread Search vs. other retrievers.90.48%Accuracy · #1 of 4
- MADQAQuestion answering over 800 multimodal PDFs, with Mixedbread Search vs. Google File Search and BM25.93.4%Accuracy · #1 of 3
- OfficeQA-ProCodex answering enterprise finance questions over 89,000 PDF pages, with Mixedbread Search vs. the raw corpus.64.42%Accuracy · #1 of 3
Toast 1
Toast 1, our search agent, against frontier models running the same search loop. Higher NDCG@10 at lower cost and latency per query is better.
- BrowseComp-PlusToast 1 vs. GPT-5.6, Claude, Kimi, Qwen, GLM and DeepSeek as standalone search agents over 100k web documents.0.86NDCG@10 · #2 of 11
- OfficeQA-ProToast 1 vs. frontier models as standalone search agents over 89,000 pages of financial PDFs.0.53NDCG@10 · #3 of 11
- LongSealToast 1 vs. GPT-5.6, Claude, Kimi, Qwen, GLM and DeepSeek as standalone search agents on LongSeal (SealQA).0.71NDCG@10 · #2 of 11
- OfficeQA-Pro V2GPT-5.6 Sol in Codex with Toast 1 as its search sub-agent, vs. every system in Databricks' OfficeQA Pro V2 release.70%Answer correctness · #1 of 13
- Harvey LAB Firm KnowledgeThe same legal agent with filesystem search, Mixedbread Search, and Toast 1: identical answers, far fewer tokens and turns.11.2 turnsTurns per task · #1 of 3