Required Content
Organize the page in the order users need the information, from understanding tasks to interpreting scores:- Evaluation scope and sources: purpose, official resources, pinned dataset and evaluator versions, task counts and splits, and applicable task browsers, licenses, and access requirements.
- Preparation and compatibility: dependencies, task images, credentials, recommended and alternative Harnesses, supported Environments, and image, workspace, resource, and network settings inferred by Recipes.
- Benchmark parameters: the meaning, type, default, valid values, and selection guidance for each Benchmark-specific parameter.
- Run examples: smoke, custom-parameter, and full-evaluation commands following the standard structure below.
- Evaluation results: metrics, task-level scoring evidence, and artifacts following the section requirements below.
- Limitations and alignment: known compatibility constraints, material differences from the official evaluation, and their effects.
Keep the Page Benchmark-Specific
Link to shared rules and expand only Benchmark-specific differences:- Link shared parameters to Shared Benchmark Fields. Keep
k, attempt strategy, and the positional Model argument out of the Benchmark-specific parameter table. - Link CLI parameter ownership and configuration precedence to the Run Parameter Reference.
- Link Harness installation, step and cost limits, command timeouts, and Model connection setup to the relevant Harness page. Explain only the additional settings required by this Benchmark here.
- Reference repeated attempts, common result structures, and aggregation through the links in the Evaluation Results section, rather than explaining them in several sections.
Run Examples Section
Use the heading “Run examples” and organize the section as command introduction, three scenario tabs, then other optional Harnesses.Command Introduction and Harness Selection
Explain that the three positional arguments toagentcompass run are Benchmark, Harness, and Model, in that order. Identify the components used on the page and the required credentials and environment preparation. Put shared environment-variable setup before the tabs.
When distinguishing multiple Harnesses, place the main execution path and the three scenario tabs below under “Recommended Harness.” With only one execution path, show the tabs directly without adding another heading level.
Three Scenario Tabs
Use<Tabs> and <Tab> with these titles and this order:
Start each tab with a short explanation of its purpose and task scope, followed by a complete command. After completing the shared prerequisites, users must be able to copy and run any tab’s command independently, without combining fragments from other tabs.
If other Harnesses are supported, add “Other optional Harnesses” after the scenario tabs. Explain their use cases and configuration differences, and provide a complete evaluation command for each. These examples may use tabs grouped by Harness; keep them separate from the three main scenarios.
Command Content and Formatting
Specify only required settings and the overrides demonstrated by the example. Omit parameters whose defaults already meet the need, and includek and attempt strategy only when demonstrating repeated attempts. Reference credentials and Model connection settings through environment variables instead of repeating their values in each example.
Apply these formatting rules throughout every Benchmark page in the user guide, including the overview, all tabs, and parameter fragments in other sections:
- One CLI option per line: use a trailing backslash (
\) for continuation, omit it on the final line, and leave no spaces after it. - One JSON field per line: indent nested objects and expand their fields. Keep the entire JSON argument quoted and do not insert shell continuation characters inside it.
- Wrap rendered lines: enable Mintlify’s
wrapoption for Bash and JSON blocks, opening their fences with```bash wrapand```json wraprespectively. Let the page wrap long strings without changing parameter values for presentation.
Evaluation Results Section
Use the heading “Evaluation Results.” Preserve old anchors when changing headings or subsections on published pages. Start with links to the run directory, aggregate scores, and per-task files and common fields, then use these two subsections:- Scoring Metrics: explain the primary and auxiliary metrics, their binary or scalar kinds, ranges or units, score direction, and how to read the default result. Expand only Benchmark-specific aggregation, denominator, or scoring-error rules; link common rules to Metrics and Aggregation.
- Task Results and Scoring Evidence: explain where each attempt stores Benchmark-specific scoring records, evidence, and artifacts, what their fields mean, and any relevant evidence that is not retained. Leave common directory layouts, filenames, and status fields to the shared documentation.
pass@k, or avg@k definitions. Put exceptions that affect score interpretation under “Scoring Metrics” and diagnostic fields under “Task Results and Scoring Evidence.” Distinguish diagnostic fields from aggregate metrics, and logical result objects from files on disk.
Use BrowseComp as an example, or SciCode for multiple scoring levels and special denominators.
Preview and Validate
After editing, check the following in order:- Review the content: commands, parameters, and scoring descriptions match the implementation; the three scenarios have distinct purposes; full evaluations retain no smoke-only filters; shared explanations are linked.
- Check the pages: English and Chinese paths and content are synchronized, and navigation, references, and old anchors remain valid. Update
docs/docs.jsonand required redirects when adding, moving, or removing pages. - Preview the rendering: run
mint devand check tabs, tables, and code blocks on desktop and narrow screens, especially long strings in--harness-params. Verify that the page wraps lines while copied text retains the original command; MDX source line breaks alone do not verify the displayed result. - Run validation: run the link and build checks from the documentation root and fix any reported problems.
