Mixedbread Quality Evals

How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.

How we measure

Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.

Third-party agents with Mixedbread Search as their retrieval tool, against other retrievers. Higher accuracy with fewer tool calls is better. Read more in Closing the Oracle Gap.

A deep-research agent answers multi-hop questions over a 100k-document web corpus. Mixedbread Search as the retriever vs. Reason-ModernColBERT and Qwen3-Embed-8B, in two agent scaffolds. Higher accuracy and fewer tool calls are better.

  1. 1MIXEDBREAD (GET_DOCUMENT)11.53 calls90.48%
  2. 2REASON-MODERNCOLBERT (GET_DOCUMENT)13.27 calls87.59%
  3. 3MIXEDBREAD (STANDARD)16.24 calls80%
  4. 4REASON-MODERNCOLBERT (STANDARD)19.31 calls79.52%
  5. 5QWEN3-EMBED-8B (STANDARD)21.74 calls71.69%

Accuracy vs. tool calls · higher is betterTested March 2026BrowseComp-Plus paperLeaderboard

Dataset and methodology

BrowseComp-Plus is a deep-research eval with multi-hop questions over ~100k web documents, designed to isolate retrieval quality from model capability. We report results in both the default (standardized) scaffold and the stronger get_document scaffold.

We evaluate retrieval pipelines paired with the BrowseComp-Plus harness, reporting accuracy and tool-call counts as published on the eval leaderboard.