> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# BrowseComp-ZH

BrowseComp-ZH ([arxiv](https://arxiv.org/abs/2504.19314)) evaluates a browsing agent's web browsing ability in Chinese: given a question whose answer is a hard-to-find fact on the Chinese-language web — one that a single search cannot answer and that requires persistent, multi-step browsing to pin down — the agent produces a short final answer, which an **LLM judge** then grades as correct or not against the ground truth.

Like DeepSearchQA, BrowseComp-ZH uses single-sided judging. The judge only compares the agent-under-test's answer against the ground truth, without comparing to any baseline. Both inference and judging run in the local process (`host_process`) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.

## How it works

A BrowseComp-ZH run has two stages — inference and judging.

### Inference and judging

* **Inference.** The model under test acts as a search agent and, driven by the harness (default [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent)), completes multi-turn tool loops such as search / visit per task, producing a short natural-language final answer.
* **Judging.** The judge model (`judge_model`) receives "question + ground truth + answer under test" and grades it with the built-in A/B/C protocol. The judge compares only the final answer, ignoring reasoning and formatting differences; equivalent expressions are accepted. The judge and the model under test are two separate endpoints; `judge_model` must be specified explicitly.

### The A/B/C verdict

The judge returns exactly one verdict, and only **A** counts as correct:

* **A — CORRECT:** the answer semantically matches the ground truth (equivalent expressions and formatting allowed).
* **B — INCORRECT:** any deviation from the ground truth.
* **C — INCOMPLETE / REPETITIVE / REFUSAL:** an invalid answer (cut off mid-sentence, looping repetition, or an explicit refusal).

## Parameters

Pass a JSON object via `--benchmark-params '{...}'`, or a `benchmark.params` block in the YAML given to `--config`; the CLI wins on shared keys. See the [Benchmark overview](/en/user_guide/modules/benchmarks/overview) for merge precedence.

### Parameter reference

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'1040px', width:'100%'}}>
    <colgroup>
      <col width="18%" />

      <col width="16%" />

      <col width="15%" />

      <col width="20%" />

      <col width="31%" />
    </colgroup>

    <thead>
      <tr><th style={{whiteSpace:'nowrap'}}>Parameter</th><th style={{whiteSpace:'nowrap'}}>Type</th><th style={{whiteSpace:'nowrap'}}>Default</th><th>Choices / values</th><th>Description</th></tr>
    </thead>

    <tbody>
      <tr><td style={{whiteSpace:'nowrap'}}><code>judge\_model</code></td><td style={{whiteSpace:'nowrap'}}>dict</td><td style={{whiteSpace:'nowrap'}}><code>null</code></td><td><code>id</code>, <code>base\_url</code>, <code>api\_key</code>, <code>api\_protocol</code>, <code>params</code></td><td>Judge model spec, <strong>required</strong> (see <a href="#judge-model-spec">Judge model spec</a>). It decides grading, and is not the CLI <code>--model-\*</code>.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>category</code></td><td style={{whiteSpace:'nowrap'}}>string / list</td><td style={{whiteSpace:'nowrap'}}><code>"all"</code></td><td><code>"all"</code>, <code>影视</code>, <code>艺术</code>, <code>地理</code>, <code>音乐</code>, <code>历史</code>, <code>医学</code>, <code>电子游戏</code>, <code>科技</code>, <code>体育</code>, <code>政策法规</code>, <code>学术论文</code></td><td>Filter tasks by category; <code>"all"</code> = no filter, a list takes the union. Task counts by category — 影视 (45), 艺术 (40), 地理 (37), 音乐 (32), 历史 (29), 医学 (26), 电子游戏 (23), 科技 (22), 体育 (18), 政策法规 (10), 学术论文 (7); 289 in total.</td></tr>
    </tbody>
  </table>
</div>

Shared Benchmark fields such as `sample_ids` follow [Benchmark Parameters](/en/user_guide/modules/benchmarks/overview). Configure repeated attempts with `--k` and `--attempt-strategy`; see [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation).

<a id="judge-model-spec" />

### Judge model spec

`judge_model` is passed as a dict with the fields `id`, `base_url`, `api_key`, `api_protocol`, and `params`, pointing to the judge model's own endpoint, with inference parameters under `params`.

We recommend **fixing a single judge** across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. The judge need not be especially strong — the A/B/C criterion (semantic match) is relatively objective, so a mid-sized model suffices. AgentCompass recommends `Qwen3.6-35B-A3B`.

## Run examples

The three positional arguments to `agentcompass run` are Benchmark, Harness, and Model. These examples use `browsecomp_zh`, [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent), and `$MODEL_NAME`, with [`host_process`](/en/user_guide/modules/environments/providers/host_process) as the Environment.

Set the following environment variables in your terminal before running the examples:

* Model under test: `MODEL_NAME`, `MODEL_BASE_URL`, and `MODEL_API_KEY`; see [Model connection details](/en/user_guide/modules/models/overview#configure-connection-details).
* Judge Model: `JUDGE_MODEL_NAME`, `JUDGE_MODEL_BASE_URL`, and `JUDGE_MODEL_API_KEY`. Use a fixed, independent judge configuration.
* Search tools: `SERPER_API_KEY` for `search` and `JINA_API_KEY` for `visit`.

See the [run command](/en/user_guide/using_agentcompass/cli/run) for configuration ownership and CLI overrides.

<Tabs>
  <Tab title="Smoke test (single task end-to-end)">
    Use `sample_ids` to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.

    ```bash wrap theme={"system"}
    agentcompass run \
      browsecomp_zh \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        },
        "sample_ids": ["4b286d72-917a-49f9-bd1d-1c8c7335b92b"]
      }' \
      --harness-params '{
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat
    ```
  </Tab>

  <Tab title="Custom parameters">
    Evaluate only a subset of categories to focus analysis on a specific domain; also demonstrates lowering the iteration limit in `--harness-params`.

    ```bash wrap theme={"system"}
    agentcompass run \
      browsecomp_zh \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        },
        "category": ["科技", "地理"]
      }' \
      --harness-params '{
        "max_iterations": 40,
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>

  <Tab title="AgentCompass recommended config">
    Evaluate the full task set. `--benchmark-params` only needs the judge model `judge_model`; use `--task-concurrency` to raise cross-task concurrency.

    ```bash wrap theme={"system"}
    agentcompass run \
      browsecomp_zh \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        }
      }' \
      --harness-params '{
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>
</Tabs>

<a id="outputs" />

## Evaluation Results

For shared result conventions, see [Run Directory](/en/user_guide/other_features/results/overview#directory-layout), [Aggregate Scores](/en/user_guide/other_features/results/summary_analysis), and [Task Files and Shared Fields](/en/user_guide/other_features/results/task_results).

<a id="metric-contract-and-aggregate-series" />

### Scoring Metrics

BrowseComp-ZH's primary metric is binary `correct`: under the [A/B/C verdict](#the-abc-verdict) described above, A maps to `true` and B/C map to `false`. There is no partial credit.

With the default configuration, each task has one attempt and the overall score is accuracy over tasks with valid scores, ranging from 0 to 1; higher is better. For example, if all 100 tasks receive valid judgments and 63 are correct, `0.63` in the report means 63%. When you filter tasks with `category` or `sample_ids`, scores cover only the selected tasks.

See [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation) for repeated attempts, category aggregation, and scoring failure rules.

<a id="per-task-details-details" />

### Task Results and Scoring Evidence

After an attempt is scored successfully, `scoring` under `meta.benchmark` contains:

| Field | Contents |
| - | - |
| `evaluation_type` | Fixed as `llm_judge`, indicating that a judge model evaluated the answer. |
| `correct` | The parsed boolean verdict, matching `metrics.correct`. |
| `model_answer` | The final answer sent to the judge. |
| `ground_truth` | The reference answer sent to the judge. |

This scoring record preserves the verdict and compared answers, but not the judge's raw response or reasoning.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.