01
A capable harness beat a raw request.
The largest effect appeared when the task required current information. Yet harnesses also beat Raw on historical-control cases with search disabled, so retrieval alone does not explain the gain.
CoffeeBench V1 · Research report
Across 20 matched coffee-research cases, every agent harness was preferred to a raw model request. We did not find a clear winner among the harnesses, and an always-visible catalog tool added work without improving answer quality.
Abstract
CoffeeBench V1 asked how much the system around a fixed model changes its ability to answer real coffee-industry research questions.
We ran DeepSeek V4 Flash as a raw request and inside three agent harnesses. Harnessed systems were preferred over Raw in 70.3–74.5% of overall pairwise judgments and 88.3–90.8% on live-web cases. Purveyors Search ranked first, but its uncertainty overlapped the other harnesses. Adding a frozen Parchment catalog did not improve pairwise quality. Trace analysis suggests a useful design principle for the next round: expose specialized capabilities only when their evidence is relevant.
Results
01
The largest effect appeared when the task required current information. Yet harnesses also beat Raw on historical-control cases with search disabled, so retrieval alone does not explain the gain.
02
Purveyors Search received 55.5% against Pi Search overall, but the quality intervals overlap and the live-web head-to-head narrowed to 53.3–46.7. V1 does not establish a harness winner.
03
The Parchment arm called the catalog 28 times even though no case asked the model to choose a coffee for sale. Fifteen calls returned nothing, and four answers carried an empty result into unrelated market reasoning.
Figure 1
Purveyors Search ranked first on Bradley–Terry preference. Every interval overlaps, so V1 does not establish a winner.
Purveyors Search
0.582 · 0.547–0.619
Purveyors, Parchment, and Search
0.561 · 0.520–0.609
Pi Search
0.537 · 0.487–0.592
Raw
0.327 · 0.249–0.400
Figure 2
Preference share against Raw, with ties split evenly. Search was disabled on historical control cases and enabled on live-web cases.
Purveyors Search
Purveyors, Parchment, and Search
Pi Search
A 50% share is even. The harnesses retained an advantage over Raw on historical cases, where web retrieval was disabled.
Figure 3
Reported input/output tokens only; reasoning-token usage was unavailable for every treatment. Successful-task median latency and jury-marked unacceptable answers are also shown. Bar lengths are scaled within each metric.
Purveyors Search
Purveyors, Parchment, and Search
Pi Search
Raw
Experiment
Every treatment used the same DeepSeek V4 Flash 0731 FP8 endpoint, case input, evidence, temperature, output limit, and provider route. “Purveyors Search” and Pi used the same Brave search backend. Their comparison changes the agent runtime and system prompt; Purveyors versus Parchment changes only whether the catalog tool is exposed.
Historical-control cases disabled public-web tools. Live-web cases enabled the same search and fetch tools for all three harnesses. The Parchment treatment retained its frozen catalog snapshot in both cohorts.
Fixed model settings
One direct model request, without a system prompt or agent loop.
Tools: No tools in either cohort.
Pi agent loop with a general research prompt and up to five steps.
Tools: Brave search and page fetch on live-web cases only.
Purveyors AI SDK loop with a green-coffee decision prompt and up to five steps.
Tools: The same Brave search and page fetch contract as Pi.
The same Purveyors loop, prompt, model, and step budget.
Tools: The shared web tools plus a frozen Parchment catalog snapshot.
Judge-model voting
OpenAI, Google, and Anthropic judge agents each voted on every response pair. The bars below show their preference shares with ties split evenly; the headline row combines all three families. Switch cohorts to inspect all 1,800 pairwise ballots.
107 wins · 78 ties · 115 losses
48.7% · 51.3%
OpenAI
46% · 54%
47% · 53%
Anthropic
53% · 47%
95 wins · 77 ties · 128 losses
44.5% · 55.5%
OpenAI
42% · 58%
45.5% · 54.5%
Anthropic
46% · 54%
182 wins · 58 ties · 60 losses
70.3% · 29.7%
OpenAI
69.5% · 30.5%
74.5% · 25.5%
Anthropic
67% · 33%
102 wins · 77 ties · 121 losses
46.8% · 53.2%
OpenAI
52% · 48%
44.5% · 55.5%
Anthropic
44% · 56%
201 wins · 45 ties · 54 losses
74.5% · 25.5%
OpenAI
74% · 26%
77.5% · 22.5%
Anthropic
72% · 28%
190 wins · 53 ties · 57 losses
72.2% · 27.8%
OpenAI
71.5% · 28.5%
78% · 22%
Anthropic
67% · 33%
Discussion
Model capability, orchestration, and evidence access interacted to shape answer quality. The raw model answered quickly while omitting important content. The catalog tool supplied legitimate data, yet the harness exposed it when that data was unrelated to the decision. The next system should make relevant capability easy to reach and irrelevant capability easy to ignore.
The next Parchment arm will expose catalog search for inventory, supplier, price, and purchase decisions, then keep it out of unrelated research tasks.
Cherry should face the current model inside the same prompt, tools, evidence, and output contract. That is the clean way to measure specialist-model lift.
We want Cherry to lead with the decision, evidence, and uncertainty, while making deeper structured detail available when the calling harness needs it.
Methods & data
CoffeeBench used 12 historical-control cases and eight live-web cases, with five trials per treatment and case. Three judge-model families produced 1,200 absolute rubric evaluations and 1,800 pairwise ballots. We report Bradley–Terry preference, absolute rubric outcomes, and operational measurements separately rather than combining them into one score.
Model-call cost was captured for every treatment. The published cost field, however, requires a complete model-plus-tool total. Brave search and Parchment tool calls did not have pinned marginal prices, so tool-using trials cannot support a complete end-to-end cost. Raw used no tools and is therefore the only fully priced treatment. Reported input/output token usage, excluding unavailable reasoning tokens, and latency remain comparable in Figure 3; future runs will capture tool-inclusive cost directly.
V1 studies one model on 20 cases with an agent jury and no human calibration. It identifies useful system patterns and concrete follow-up experiments; it does not establish a universal model or harness winner, isolate retrieval as the sole cause of the Raw gap, or prove that Parchment exposure caused the observed quality result.
| Treatment | Quality · 95% interval | Strict pass | Must-miss | Critical | Unacceptable | Reported input/output tokens / task | Median latency | Complete cost / task |
|---|---|---|---|---|---|---|---|---|
| Purveyors Search | #1 · 0.582 · 0.547–0.619 | 69% | 17% | 17% | 30% | 6,785 | 9.8 s | Unavailable |
| Purveyors, Parchment, and Search | #2 · 0.561 · 0.520–0.609 | 68% | 22% | 14% | 30% | 7,624 | 9.3 s | Unavailable |
| Pi Search | #3 · 0.537 · 0.487–0.592 | 64% | 24% | 18% | 32% | 7,045 | 8.8 s | Unavailable |
| Raw | #4 · 0.327 · 0.249–0.400 | 59% | 41% | 8% | 41% | 1,801 | 5.8 s | $0.0001419654 |
Coffee intelligence platform. Daily-normalized data from 40+ US specialty importers, turned into procurement-ready analytics for green coffee buyers and roasting teams.
© 2026 Purveyors.