> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Documentation Update

Write a Benchmark user guide that lets users prepare their environment, run an evaluation, and interpret its results without reading the source.

When adding or changing a Benchmark, maintain both pages:

```text theme={"system"}
docs/en/user_guide/modules/benchmarks/<benchmark-id>.mdx
docs/zh/user_guide/modules/benchmarks/<benchmark-id>.mdx
```

Follow the [documentation contribution guide](/en/developer_guide/contributing/documentation) for placement, localization, navigation, and public-content boundaries. This page defines the content of a Benchmark guide and the structure of its run examples and evaluation results.

## Required Content

Organize the page in the order users need the information, from understanding tasks to interpreting scores:

* **Evaluation scope and sources**: purpose, official resources, pinned dataset and evaluator versions, task counts and splits, and applicable task browsers, licenses, and access requirements.
* **Preparation and compatibility**: dependencies, task images, credentials, recommended and alternative Harnesses, supported Environments, and image, workspace, resource, and network settings inferred by Recipes.
* **Benchmark parameters**: the meaning, type, default, valid values, and selection guidance for each Benchmark-specific parameter.
* **Run examples**: smoke, custom-parameter, and full-evaluation commands following the [standard structure](#run-examples-section) below.
* **Evaluation results**: metrics, task-level scoring evidence, and artifacts following the [section requirements](#evaluation-results-section) below.
* **Limitations and alignment**: known compatibility constraints, material differences from the official evaluation, and their effects.

Verify fields, defaults, commands, and scoring behavior against the current implementation. Explain why a Harness is recommended, distinguishing an upstream recommendation from an AgentCompass integration recommendation.

## Keep the Page Benchmark-Specific

Link to shared rules and expand only Benchmark-specific differences:

* Link shared parameters to [Shared Benchmark Fields](/en/user_guide/modules/benchmarks/overview#shared-benchmark-fields). Keep `k`, attempt strategy, and the positional Model argument out of the Benchmark-specific parameter table.
* Link CLI parameter ownership and configuration precedence to the [Run Parameter Reference](/en/user_guide/using_agentcompass/cli/run#parameter-reference).
* Link Harness installation, step and cost limits, command timeouts, and Model connection setup to the relevant Harness page. Explain only the additional settings required by this Benchmark here.
* Reference repeated attempts, common result structures, and aggregation through the links in the [Evaluation Results section](#evaluation-results-section), rather than explaining them in several sections.

## Run Examples Section

Use the heading “Run examples” and organize the section as command introduction, three scenario tabs, then other optional Harnesses.

### Command Introduction and Harness Selection

Explain that the three positional arguments to `agentcompass run` are Benchmark, Harness, and Model, in that order. Identify the components used on the page and the required credentials and environment preparation. Put shared environment-variable setup before the tabs.

When distinguishing multiple Harnesses, place the main execution path and the three scenario tabs below under “Recommended Harness.” With only one execution path, show the tabs directly without adding another heading level.

### Three Scenario Tabs

Use `<Tabs>` and `<Tab>` with these titles and this order:

| Tab | Purpose | Command requirements |
| - | - | - |
| **Smoke test (single task end-to-end)** | Verify task preparation, inference, and scoring together. | Select one real, valid task and use the minimum necessary configuration. |
| **Custom parameters** | Demonstrate representative adjustments for this Benchmark. | Explain the reason for each adjustment and its effect on task scope or execution; avoid listing every parameter. |
| **AgentCompass recommended config** | Provide the recommended full evaluation. | State the task scope and remove smoke-only task filters, while retaining settings required for the selected split or version. |

Start each tab with a short explanation of its purpose and task scope, followed by a complete command. After completing the shared prerequisites, users must be able to copy and run any tab's command independently, without combining fragments from other tabs.

If other Harnesses are supported, add “Other optional Harnesses” after the scenario tabs. Explain their use cases and configuration differences, and provide a complete evaluation command for each. These examples may use tabs grouped by Harness; keep them separate from the three main scenarios.

### Command Content and Formatting

Specify only required settings and the overrides demonstrated by the example. Omit parameters whose defaults already meet the need, and include `k` and attempt strategy only when demonstrating repeated attempts. Reference credentials and Model connection settings through environment variables instead of repeating their values in each example.

Apply these formatting rules throughout every Benchmark page in the user guide, including the overview, all tabs, and parameter fragments in other sections:

* **One CLI option per line**: use a trailing backslash (`\`) for continuation, omit it on the final line, and leave no spaces after it.
* **One JSON field per line**: indent nested objects and expand their fields. Keep the entire JSON argument quoted and do not insert shell continuation characters inside it.
* **Wrap rendered lines**: enable Mintlify's `wrap` option for Bash and JSON blocks, opening their fences with ` ```bash wrap ` and ` ```json wrap ` respectively. Let the page wrap long strings without changing parameter values for presentation.

## Evaluation Results Section

Use the heading “Evaluation Results.” Preserve old anchors when changing headings or subsections on published pages.

Start with links to the [run directory](/en/user_guide/other_features/results/overview#directory-layout), [aggregate scores](/en/user_guide/other_features/results/summary_analysis), and [per-task files and common fields](/en/user_guide/other_features/results/task_results), then use these two subsections:

* **Scoring Metrics**: explain the primary and auxiliary metrics, their binary or scalar kinds, ranges or units, score direction, and how to read the default result. Expand only Benchmark-specific aggregation, denominator, or scoring-error rules; link common rules to [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation).
* **Task Results and Scoring Evidence**: explain where each attempt stores Benchmark-specific scoring records, evidence, and artifacts, what their fields mean, and any relevant evidence that is not retained. Leave common directory layouts, filenames, and status fields to the shared documentation.

Link to scoring mechanisms explained earlier on the page. Do not add a separate “Scoring Notes” section that repeats common status, error-handling, `pass@k`, or `avg@k` definitions. Put exceptions that affect score interpretation under “Scoring Metrics” and diagnostic fields under “Task Results and Scoring Evidence.” Distinguish diagnostic fields from aggregate metrics, and logical result objects from files on disk.

Use [BrowseComp](/en/user_guide/modules/benchmarks/browsecomp#evaluation-results) as an example, or [SciCode](/en/user_guide/modules/benchmarks/scicode#evaluation-results) for multiple scoring levels and special denominators.

## Preview and Validate

After editing, check the following in order:

1. **Review the content**: commands, parameters, and scoring descriptions match the implementation; the three scenarios have distinct purposes; full evaluations retain no smoke-only filters; shared explanations are linked.
2. **Check the pages**: English and Chinese paths and content are synchronized, and navigation, references, and old anchors remain valid. Update `docs/docs.json` and required redirects when adding, moving, or removing pages.
3. **Preview the rendering**: run `mint dev` and check tabs, tables, and code blocks on desktop and narrow screens, especially long strings in `--harness-params`. Verify that the page wraps lines while copied text retains the original command; MDX source line breaks alone do not verify the displayed result.
4. **Run validation**: run the link and build checks from the documentation root and fix any reported problems.

```bash wrap theme={"system"}
cd docs
mint broken-links
mint validate
```

Documentation checks do not replace actual evaluation runs. See [Validation and Alignment](/en/developer_guide/extensions/benchmark/validation_and_alignment) for smoke tests, full evaluations, and comparison with official results.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.