Mixedbread Quality Evals

How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.

How we measure

Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.

Third-party agents with Mixedbread Search as their retrieval tool, against other retrievers. Higher accuracy with fewer tool calls is better. Read more in Closing the Oracle Gap.

Agents answer questions over 800 multimodal PDFs (Snowflake leaderboard). Gemini and Button with Mixedbread Search as the retriever vs. the same or comparable agents with Google File Search or BM25. Higher accuracy and fewer tool calls are better.

  1. 1GEMINI 3.5 FLASH + MIXEDBREAD AGENTIC4.7 calls93.4%
  2. 2BUTTON + MIXEDBREAD12.8 calls91.7%
  3. 3GEMINI 3 PRO + MIXEDBREAD1 call88.2%
  4. 4CLAUDE SONNET 4.5 + BM2535.1 calls80.6%
  5. 5GEMINI 3 PRO + FILE SEARCH1 call78.6%

Accuracy vs. tool calls · higher is betterTested May 2026MADQA paperLeaderboard

Dataset and methodology

MADQA tests whether agents can navigate >18,000 pages from 800 heterogeneous PDFs to answer 500 human-authored questions. The eval is multimodal (page screenshots) and runs in one-shot or agentic (up to 10 turns) modes.

We report accuracy and tool-call counts for each model + retriever combination. For Mixedbread Agentic, tool calls count the internal retrieval rounds executed inside the Mixedbread API invocation, not only the single API call visible to the outer agent.