Mixedbread Quality Evals

How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.

How we measure

Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.

Third-party agents with Mixedbread Search as their retrieval tool, against other retrievers. Higher accuracy with fewer tool calls is better. Read more in Closing the Oracle Gap.

OpenAI Codex answers enterprise finance questions over 89,000 pages of PDFs (Databricks benchmark). Codex with Mixedbread Search vs. Codex with the raw corpus, plus Databricks' GPT 5.4 semantic-search agent. Higher correctness and fewer tool calls are better.

  1. 1CODEX + MIXEDBREAD (THINKING HIGH)17.35 calls64.42%
  2. 2CODEX + CORPUS (THINKING HIGH)34.5 calls56.39%
  3. 3GPT 5.4 + SEMANTIC SEARCH86.4 calls51.9%

Accuracy vs. tool calls · higher is betterTested March 2026OfficeQA-Pro paper

Dataset and methodology

OfficeQA-Pro contains 89,000 pages of complex financial documents (U.S. Treasury Bulletins, dense tables, scanned PDFs) with questions that require multi-document reasoning. We hold the harness (OpenAI Codex, thinking high) constant and vary only the retrieval tooling.

All Codex configurations use the same model and reasoning budget; we vary retrieval (Mixedbread / Corpus baseline). The GPT 5.4 + Semantic Search row is reported from the Databricks paper.