> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Terminal-Bench 2 Verified

Terminal-Bench 2 Verified is the verified Terminal-Bench 2 subset hosted on Hugging Face. It follows the same task execution and verifier flow as [Terminal-Bench 2](/en/user_guide/modules/benchmarks/terminal_bench_2), normally with [`terminus2`](/en/user_guide/modules/harnesses/terminus2).

## How it works

1. **Load tasks.** AgentCompass clones the Hugging Face dataset and retrieves its Git LFS objects before loading the task directories.
2. **Run the agent.** The recipe prepares each task's container image and workspace, then the terminal harness solves the instruction.
3. **Verify the result.** The Harbor verifier executes `tests/test.sh`; a reward of `1` marks the task `correct`.

<Note type="warning">
  `git-lfs` must be installed in the process that loads the dataset. Without it, the verified task assets cannot be retrieved.
</Note>

## Parameters

You can run all tasks in the default dataset without `--benchmark-params`. For task selection and aggregation options, see the shared [Benchmark parameter reference](/en/user_guide/modules/benchmarks/overview#shared-benchmark-fields). Configure phase deadlines and timeout multipliers through `--execution-params`; see [run controls](/en/user_guide/using_agentcompass/run_controls#set-an-appropriate-timeout).

## Run examples

`agentcompass run` takes three positional arguments in order: Benchmark, Harness, and Model. The examples use `terminal_bench_2_verified`; harness choices are described below.

Before running, make sure local [Docker](/en/user_guide/modules/environments/providers/docker) is available and set `MODEL_NAME`, `MODEL_BASE_URL`, and `MODEL_API_KEY` to the model under test, API endpoint, and API key.

### Recommended harness

The recommended terminal agent is [`terminus2`](/en/user_guide/modules/harnesses/terminus2).

<Tabs>
  <Tab title="Smoke test (single task end-to-end)">
    Run the `filter-js-from-html` task to verify container preparation, inference, and scoring, leaving other parameters at their defaults.

    ```bash wrap theme={"system"}
    agentcompass run \
      terminal_bench_2_verified \
      terminus2 \
      "$MODEL_NAME" \
      --env docker \
      --benchmark-params '{
        "sample_ids": ["filter-js-from-html"]
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY"
    ```
  </Tab>

  <Tab title="Custom parameters">
    Evaluate every task in the default dataset with higher inference and evaluation timeout multipliers, for environments with slower model responses or tool execution.

    ```bash wrap theme={"system"}
    agentcompass run \
      terminal_bench_2_verified \
      terminus2 \
      "$MODEL_NAME" \
      --env docker \
      --execution-params '{
        "evaluation_timeout_multiplier": 8,
        "run_timeout_multiplier": 16
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY"
    ```
  </Tab>

  <Tab title="AgentCompass recommended config">
    Run a full evaluation with the AgentCompass recommended configuration. If model service conditions, deployment instance count, task concurrency, or model capability increase task duration, pass `evaluation_timeout_multiplier` and `run_timeout_multiplier` through `--execution-params`.

    ```bash wrap theme={"system"}
    agentcompass run \
      terminal_bench_2_verified \
      terminus2 \
      "$MODEL_NAME" \
      --env docker \
      --task-concurrency 16 \
      --harness-params '{
        "max_turns": 300
      }' \
      --execution-params '{
        "run_timeout_seconds": 14400
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY"
    ```
  </Tab>
</Tabs>

### Other optional harnesses

The commands below evaluate the full dataset. Configure the model variables for an OpenAI Responses API endpoint when using Codex, or an Anthropic Messages API endpoint when using Claude Code.

[`codex`](/en/user_guide/modules/harnesses/codex) and [`claude_code`](/en/user_guide/modules/harnesses/claude_code) are two other harness options. Pass `--recipe terminalbench2_verified_docker_ac` to use the AgentCompass prebuilt image. It includes download dependencies such as Node.js, npm, curl, and wget for Codex, Claude Code, and similar harnesses.

<Tabs>
  <Tab title="Run with the official image">
    Omit `--recipe` to use the official task image. Because it does not include the Node bootstrap dependencies, provide the matching installation command explicitly.

    ```bash wrap theme={"system"}
    # Codex
    agentcompass run \
      terminal_bench_2_verified \
      codex \
      "$MODEL_NAME" \
      --env docker \
      --task-concurrency 16 \
      --harness-params '{
        "install_command": "apt-get update && apt-get install -y curl ca-certificates && curl -fsSL https://deb.nodesource.com/setup_20.x | bash - && apt-get install -y nodejs && npm install -g @openai/codex"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-responses

    # Claude Code
    agentcompass run \
      terminal_bench_2_verified \
      claude_code \
      "$MODEL_NAME" \
      --env docker \
      --task-concurrency 16 \
      --harness-params '{
        "install_command": "apt-get update && apt-get install -y curl ca-certificates && curl -fsSL https://deb.nodesource.com/setup_20.x | bash - && apt-get install -y nodejs && npm install -g @anthropic-ai/claude-code"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol anthropic
    ```
  </Tab>

  <Tab title="Use AgentCompass images">
    Pass `--recipe terminalbench2_verified_docker_ac` to select the AgentCompass prebuilt image without explicitly providing the corresponding installation command.

    ```bash wrap theme={"system"}
    # Codex
    agentcompass run \
      terminal_bench_2_verified \
      codex \
      "$MODEL_NAME" \
      --env docker \
      --recipe terminalbench2_verified_docker_ac \
      --task-concurrency 16 \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-responses

    # Claude Code
    agentcompass run \
      terminal_bench_2_verified \
      claude_code \
      "$MODEL_NAME" \
      --env docker \
      --recipe terminalbench2_verified_docker_ac \
      --task-concurrency 16 \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol anthropic
    ```
  </Tab>
</Tabs>

<a id="output" />

## Evaluation Results

For shared result conventions, see [Run Directory](/en/user_guide/other_features/results/overview#directory-layout), [Aggregate Scores](/en/user_guide/other_features/results/summary_analysis), and [Task Files and Shared Fields](/en/user_guide/other_features/results/task_results).

<a id="aggregate-metrics-summarymd" />

### Scoring Metrics

Terminal-Bench 2 Verified's primary metric is binary `correct`, recording the verdict from the [verification flow](#how-it-works) above. A verifier reward of exactly `1` maps to `true`; other valid reward values map to `false`. There is no partial credit.

With the default configuration, each task has one attempt and the overall score is the pass rate over tasks with valid scores, ranging from 0 to 1; higher is better.

See [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation) for repeated attempts, category aggregation, and scoring failure rules.

<a id="per-task-details-details" />

### Task Results and Scoring Evidence

`eval_raw_data` under `meta.benchmark` preserves the Harbor verifier's scoring evidence:

| Field | Contents |
| - | - |
| `verify_result` | Raw verifier result, when available. Its `rewards` object contains the `reward` used to determine whether the task passed. |
| `testcase_output` | Contents of the verifier's `test-stdout.txt`, including test output. |
| `verifier` | Start and finish times for verification. |
| `error` | Diagnostics when verification encounters a problem. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.