Mixedbread Quality Evals
How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
How we measure
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
Third-party agents with Mixedbread Search as their retrieval tool, against other retrievers. Higher accuracy with fewer tool calls is better. Read more in Closing the Oracle Gap.
OfficeQA-Pro
Open OfficeQA-Pro paperOpenAI Codex answers enterprise finance questions over 89,000 pages of PDFs (Databricks benchmark). Codex with Mixedbread Search vs. Codex with the raw corpus, plus Databricks' GPT 5.4 semantic-search agent. Higher correctness and fewer tool calls are better.
- 1CODEX + MIXEDBREAD (THINKING HIGH)17.35 calls64.42%
- 2CODEX + CORPUS (THINKING HIGH)34.5 calls56.39%
- 3GPT 5.4 + SEMANTIC SEARCH86.4 calls51.9%
Accuracy vs. tool calls · higher is betterTested March 2026OfficeQA-Pro paper
Dataset and methodology
OfficeQA-Pro contains 89,000 pages of complex financial documents (U.S. Treasury Bulletins, dense tables, scanned PDFs) with questions that require multi-document reasoning. We hold the harness (OpenAI Codex, thinking high) constant and vary only the retrieval tooling.
All Codex configurations use the same model and reasoning budget; we vary retrieval (Mixedbread / Corpus baseline). The GPT 5.4 + Semantic Search row is reported from the Databricks paper.