host_process) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.
How it works
A DeepSearchQA run has two stages — inference and judging — where the judging stage applies different criteria based on the task’s answer form.Inference and judging
- Inference. The model under test acts as a search agent and, driven by the harness (default
naive_search_agent), completes multi-turn tool loops such as search / visit per task, producing a natural-language answer. - Judging. The judge model (
judge_model) receives “question + ground truth + answer form + answer under test” and grades it with the official rubric template. The judge and the model under test are two separate endpoints;judge_modelmust be specified explicitly.
How the two answer forms are judged
The judge applies different criteria based on each task’sanswer_type:
- Single Answer (316 tasks): the answer under test is judged correct if it semantically hits the ground truth; verbatim matching is not required.
- Set Answer (584 tasks): the ground truth is a set of items, and the answer under test must ** hit every item ; the judge also checks whether the answer includes ** excessive answers beyond the ground truth.
Correctness Details (a per-item boolean dictionary of hits), Excessive Answers (a list of extra answers), and Explanation (the grading rationale). A task is judged correct ** if and only if ** all expected items are hit ** and ** no excessive answers exist; any missing item or any excessive answer counts as incorrect.
Parameters
Pass a JSON object via--benchmark-params '{...}', or a benchmark.params block in the YAML given to --config; the CLI wins on shared keys. See the Benchmark overview for merge precedence.
Parameter reference
| Parameter | Type | Default | Choices / values | Description |
|---|---|---|---|---|
judge_model | dict | null | id, base_url, api_key, api_protocol, params | Judge model spec, required (see Judge model spec). It decides grading, and is not the CLI —model-*. |
category | string / list | ”all" | "all”, a single category name, or a list of category names (17 listed below) | Filter tasks by category; “all” = no filter. A list takes the union. |
answer_type | string | ”all” | all / Single Answer / Set Answer | Filter tasks by answer form; all = no filter. Case and full name must match exactly. |
sample_ids follow Benchmark Parameters. Configure repeated attempts with --k and --attempt-strategy; see Metrics and Aggregation.
All 17 category values (click to expand)
All 17 category values (click to expand)
Politics & Government (148), Finance & Economics (132), Geography (95), Education (94), Health (92), Science (90), Other (65), History (44), Travel (36), Media & Entertainment (29), Arts (26), Technology (22), Sports (20), Current Events (3), Biology (2), Linguistics (1), Arts & Entertainment (1). Numbers in parentheses are the task count per category (900 total).Judge model spec
judge_model is passed as a dict with the fields id, base_url, api_key, api_protocol, and params, pointing to the judge model’s own endpoint, with inference parameters under params.
We recommend fixing a single judge across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. The judge need not be especially strong — DeepSearchQA’s criteria (semantic hit + excess check) are relatively objective, so a mid-sized model suffices. AgentCompass recommends Qwen3.6-35B-A3B.
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. These examples use deepsearchqa, naive_search_agent, and $MODEL_NAME, with host_process as the Environment.
Set the following environment variables in your terminal before running the examples:
- Model under test:
MODEL_NAME,MODEL_BASE_URL, andMODEL_API_KEY; see Model connection details. - Judge Model:
JUDGE_MODEL_NAME,JUDGE_MODEL_BASE_URL, andJUDGE_MODEL_API_KEY. Use a fixed, independent judge configuration. - Search tools:
SERPER_API_KEYforsearchandJINA_API_KEYforvisit.
- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Use
sample_ids to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
DeepSearchQA’s primary metric is binarycorrect: under the answer-form rules, it is true only when all expected items are matched with no excessive answers. Item-level judgments explain the result but do not earn partial credit. An empty answer is graded incorrect.
With the default configuration, each task has one attempt and the overall score is accuracy over tasks with valid scores, ranging from 0 to 1; higher is better. Single-answer and set-answer tasks each count as one task, regardless of the number of items in the answer set. Scores cover only the tasks selected by category, answer_type, or sample_ids.
See Metrics and Aggregation for repeated attempts, category aggregation, and scoring failure rules.
Task Results and Scoring Evidence
Each attempt’s scoring record is stored inscoring under meta.benchmark. A successful judge response produces:
An empty answer skips the judge and records
correct=false with reason=empty_model_response, without item-level judgments. For judge-call or response-parsing failures, inspect error in the same record; parsing failures may also preserve a truncated raw_response.