Mixedbread Quality Evals
How Mixedbread performs on public retrieval and agent benchmarks, measured against the strongest competing models.
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
How we measure
Every result comes from a public benchmark on its standard split and evaluation script. Retrieval numbers are Wholembed V3 served through the Stores API; agent numbers come from Toast 1 or third-party agents running on Mixedbread Search, with cost and latency measured per query at list prices. Each benchmark lists its dataset, exact methodology and testing dates.
Toast 1, our search agent, against frontier models running the same search loop. Higher NDCG@10 at lower cost and latency per query is better.
Harvey LAB Firm Knowledge
Open Harvey LAB Law Firm KnowledgeToken efficiency: the same GPT-5.6 Sol legal agent with filesystem search, with Mixedbread Search, and with Toast 1 as its search sub-agent. Answer quality is identical in all three; fewer tokens and turns per task are better.
- 1+ TOAST 1 SUBAGENT23M tokens11.2 turns
- 2+ MIXEDBREAD SEARCH47M tokens14.6 turns
- 3VANILLA AGENT80.6M tokens21.7 turns
Turns per task vs. total tokens · lower is betterTested August 2026Harvey LAB Law Firm KnowledgeIntroducing Toast 1
Dataset and methodology
Harvey LAB's Law Firm Knowledge benchmark evaluates how well an agent can search and use institutional legal knowledge at realistic scale. We evaluate a randomly selected subset of 33 tasks to make repeated comparative runs tractable.
GPT-5.6 Sol on identical tasks and evaluation setup; only the retrieval stack changes: the vanilla agent's filesystem search, Mixedbread Search, and Toast 1 as a dedicated search sub-agent on Mixedbread Search. Tokens are totals across the 33 tasks; turns are agent-loop iterations per task.