> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Choose a Benchmark

Select a registered benchmark and configure its complete benchmark-parameter schema.

Benchmarks define what is evaluated. Each benchmark owns its dataset, stable task ids, task preparation, scoring logic,
and aggregate metrics. Select the benchmark as the first positional argument to `agentcompass run`:

```bash wrap theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME"
```

## Find a Benchmark

Use the live registry to see the benchmarks available in your installed AgentCompass revision:

```bash wrap theme={"system"}
agentcompass list benchmark
```

The sidebar links to benchmarks with dedicated task, parameter, compatibility, and run documentation.

## Configure Benchmark Parameters

The [Run Parameter Reference](/en/user_guide/using_agentcompass/cli/run#parameter-reference) introduces
`--benchmark-params <json>`. The `<json>` value is one JSON object containing the complete parameter override for the
selected benchmark:

```bash wrap theme={"system"}
agentcompass run \
  <benchmark> \
  <harness> \
  "$MODEL_NAME" \
  --benchmark-params '{
    "sample_ids": ["<task-id>"],
    "<benchmark-specific-field>": "<value>"
  }'
```

The accepted object combines two schemas:

```text theme={"system"}
benchmark params
  ├─ shared fields from RuntimeBenchmarkConfig
  └─ fields defined by the selected benchmark config
```

### Shared Benchmark Fields

Every benchmark config derived from `RuntimeBenchmarkConfig` supports these user-facing fields. The table shows base
defaults; the selected Benchmark can override them.

<table style={{width:'100%', tableLayout:'fixed'}}>
  <thead>
    <tr><th style={{width:'24%', whiteSpace:'nowrap'}}>Field</th><th style={{width:'25%'}}>Type</th><th style={{width:'18%'}}>Base default</th><th style={{width:'33%'}}>Meaning and when to change it</th></tr>
  </thead>

  <tbody>
    <tr><td style={{width:'24%', whiteSpace:'nowrap'}}><code>sample\_ids</code></td><td><code>list\[str] | null</code></td><td><code>null</code></td><td>Runs only the listed stable task ids. Use it for a smoke test, failed-task rerun, or a controlled subset. Unknown ids fail before execution.</td></tr>
    <tr><td style={{width:'24%', whiteSpace:'nowrap'}}><code>aggregation\_mode</code></td><td><code>"micro\_weighted" | "category\_mean"</code></td><td><code>"micro\_weighted"</code></td><td>Selects how generic metrics combine tasks and categories when <code>category\_hierarchy</code> is not set.</td></tr>
    <tr><td style={{width:'24%', whiteSpace:'nowrap'}}><code>category\_hierarchy</code></td><td><code>object | null</code></td><td><code>null</code></td><td>Uses an explicit category aggregation tree and takes precedence over <code>aggregation\_mode</code>. Leave unset unless the Benchmark documentation defines one.</td></tr>
  </tbody>
</table>

Repeated attempts are configured under <code>execution.attempts</code>, not in this object. See [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation) for the attempt plan, Metric Contracts, and how the aggregation fields above combine task results.

The model id is not part of this JSON object. It remains the third positional argument to `agentcompass run` and is
injected into the benchmark config by the runtime.

Each benchmark also adds its own fields to the shared schema. Whether or not it has a dedicated page, query the complete
field list, types, defaults, and descriptions directly from the installed code:

```bash wrap theme={"system"}
agentcompass config docs benchmark <benchmark-id>
```

When a dedicated benchmark page exists, use it as the source for valid values, recommended settings, required
credentials, and field interactions.

### Build the JSON Object

For example, `swebench_verified` combines shared task-selection and aggregation fields with its own preparation and evaluator fields:

```json wrap theme={"system"}
{
  "sample_ids": ["astropy__astropy-12907"],
  "prepare_mode": "prebaked",
  "workspace_root": "/testbed",
  "eval_timeout": 1800
}
```

This expanded object demonstrates ownership; it is not a recommendation to repeat defaults in every command. Pass only
the fields that must differ from the selected benchmark's effective configuration.

`--benchmark-params` must be valid JSON, so keys and string values use double quotes. CLI values override matching keys
from `benchmark.params` in configuration files. Inspect the merged built-in and configuration-file values before adding
the final CLI override:

```bash wrap theme={"system"}
agentcompass config show \
  --benchmark <benchmark-id> \
  --config <config-file>
```

## Images and Provider Settings

Heavyweight benchmarks usually attach task images, workspace roots, and resource hints to task metadata. Compatible
[recipes](/en/user_guide/other_features/recipes) translate those requirements for Docker, Daytona, or Modal. Keep provider image,
resource, and network overrides in `--env-params`; they are not benchmark parameters.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.