Mixedbread Quality Evals
How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
How we measure
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
Third-party agents with Mixedbread Search as their retrieval tool, against other retrievers. Higher accuracy with fewer tool calls is better. Read more in Closing the Oracle Gap.
BrowseComp-Plus
Open BrowseComp-Plus paperA deep-research agent answers multi-hop questions over a 100k-document web corpus. Mixedbread Search as the retriever vs. Reason-ModernColBERT and Qwen3-Embed-8B, in two agent scaffolds. Higher accuracy and fewer tool calls are better.
- 1MIXEDBREAD (GET_DOCUMENT)11.53 calls90.48%
- 2REASON-MODERNCOLBERT (GET_DOCUMENT)13.27 calls87.59%
- 3MIXEDBREAD (STANDARD)16.24 calls80%
- 4REASON-MODERNCOLBERT (STANDARD)19.31 calls79.52%
- 5QWEN3-EMBED-8B (STANDARD)21.74 calls71.69%
Accuracy vs. tool calls · higher is betterTested March 2026BrowseComp-Plus paperLeaderboard
Dataset and methodology
BrowseComp-Plus is a deep-research eval with multi-hop questions over ~100k web documents, designed to isolate retrieval quality from model capability. We report results in both the default (standardized) scaffold and the stronger get_document scaffold.
We evaluate retrieval pipelines paired with the BrowseComp-Plus harness, reporting accuracy and tool-call counts as published on the eval leaderboard.