> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# SWE-bench Pro Verified

SWE-bench Pro Verified is a verified version of [SWE-bench Pro](/en/user_guide/modules/benchmarks/swebench_pro) that addresses two sources of evaluation unreliability identified through trajectory analysis: reward hacking caused by leakage of gold solutions or hidden evaluation information, and task quality issues such as misleading problem statements and improperly scoped tests ([paper](https://arxiv.org/abs/2609.08149), [dataset](https://huggingface.co/datasets/opencompass/SWEBench-Pro-Verified), [evaluation scripts](https://github.com/scaleapi/SWE-bench_Pro-os)).

The benchmark contains 731 tasks. Four anti-hacking controls, including repository reconstruction, test artifact concealment, metadata filtering and anonymization, and network blocking, apply to all tasks; task refinement further corrects 102 tasks with confirmed quality issues: 22 misleading descriptions, 75 overly narrow tests, 3 overly broad tests, and 2 other issues.

## How it works

A task has separate inference and evaluation stages:

1. **Load and prepare.** AgentCompass reads `instance_id`, repository, base commit, problem statement, requirements, and any newly introduced interface from the dataset. A provider recipe normally selects the task's prebaked image and exposes its repository at `/app`.
2. **Apply the refined data.** AgentCompass loads the complete task data from `swebench_pro_verified.jsonl` in `opencompass/SWEBench-Pro-Verified`.
3. **Apply anti-hacking controls.** AgentCompass removes evaluation test files, rebuilds the repository as a fresh commit with its Git history erased, blocks code-hosting domains through a blacklist, and filters and anonymizes task metadata such as `instance_id`.
4. **Run the coding agent.** A harness such as [mini-SWE-agent](/en/user_guide/modules/harnesses/mini_swe_agent) or [OpenHands](/en/user_guide/modules/harnesses/openhands) receives the issue and edits the checked-out repository. It must write the final unified diff to `/app/patch.txt` when using the standard recipe layout.
5. **Start a fresh evaluation environment.** Inference changes are not trusted as the evaluation workspace. AgentCompass starts a new environment from the task image, resets `/app` to `base_commit`, and applies the patch.
6. **Run the official instance scripts.** The benchmark loads the task's `run_script.sh` and `parser.py` from the local `run_scripts/<instance_id>/` tree, downloading missing scripts from `SWE-bench_Pro-os`. The parser turns test logs into structured results.
7. **Decide resolution.** A task is `resolved=true` only when every required `FAIL_TO_PASS` and `PASS_TO_PASS` test appears in the passed-test set.

## Environments

This benchmark currently cannot run on Daytona or Modal because neither provider supports blacklists; they support only whitelists or complete outbound-network blocking ([Daytona network limits](https://www.daytona.io/docs/en/network-limits/), [Modal sandbox networking](https://modal.com/docs/guide/sandbox-networking)). Docker is currently the only public provider that can enforce a blacklist; see [Network Configuration](/en/user_guide/modules/environments/configuration/network).

Each task environment requires at least 4 CPU cores and 8 GiB of memory. Any Environment provider applies these two values by default. These values can be overridden manually, but allocating fewer resources than required is not recommended; see [Resource Limits](/en/user_guide/modules/environments/configuration/resource_limits).

## Parameters

Pass benchmark configuration via `--benchmark-params '{...}'`, or through `benchmark.params` in a YAML file given to `--config`; the CLI wins on shared keys.

### Parameter reference

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'1040px', width:'100%'}}>
    <colgroup>
      <col width="18%" />

      <col width="12%" />

      <col width="14%" />

      <col width="24%" />

      <col width="32%" />
    </colgroup>

    <thead>
      <tr><th style={{whiteSpace:'nowrap'}}>Parameter</th><th style={{whiteSpace:'nowrap'}}>Type</th><th style={{whiteSpace:'nowrap'}}>Default</th><th>Choices / values</th><th>Description</th></tr>
    </thead>

    <tbody>
      <tr><td style={{whiteSpace:'nowrap'}}><code>prepare\_mode</code></td><td>string</td><td><code>git\_clone</code></td><td><code>git\_clone</code> / <code>prebaked</code></td><td>How the inference repository is prepared. Built-in provider recipes normally replace this with <code>prebaked</code>.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>workspace\_root</code></td><td>string</td><td><code>/app</code></td><td>absolute environment path</td><td>Root used for task workspaces before recipe overrides.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>repo\_url\_template</code></td><td>string</td><td><code>[https://github.com/\&#123;repo\&#125;.git](https://github.com/\&#123;repo\&#125;.git)</code></td><td>template containing <code>\{repo}</code></td><td>Repository clone URL used in <code>git\_clone</code> mode.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>scripts\_dir</code></td><td>string</td><td><code>""</code></td><td>local directory</td><td>Controller-side directory containing <code>\<instance\_id>/run\_script.sh</code> and <code>parser.py</code>. Empty resolves to the data directory's <code>run\_scripts/</code>.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>dockerfiles\_dir</code></td><td>string</td><td><code>""</code></td><td>local directory</td><td>Controller-side official Dockerfile root used to recover task environment exports. Empty resolves under the data directory.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>evaluation\_repo\_dir</code></td><td>string</td><td><code>/app</code></td><td>absolute environment path</td><td>Repository path in the evaluation image; recipes keep it at <code>/app</code>.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>evaluation\_workspace\_dir</code></td><td>string</td><td><code>/app</code></td><td>absolute environment path</td><td>Directory where the patch, scripts, logs, and parser output are staged during evaluation.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>sample\_ids</code></td><td>list / string / null</td><td><code>null</code></td><td>valid instance ids</td><td>Optional exact task filter. Unknown ids fail fast.</td></tr>
    </tbody>
  </table>
</div>

The model id is the third positional argument to `agentcompass run`, not a `--benchmark-params` field. This benchmark does not expose a `split` parameter: it loads the public `test` split.

### Inference, model, and evaluation controls

| What is limited | mini-SWE-agent | OpenHands | SWE-bench Pro Verified |
| - | - | - | - |
| One model request | `--model-params.timeout` (unset by AgentCompass) | `--model-params.timeout`, otherwise `conversation_timeout=3600` | — |
| One repository command | `command_timeout=2400` | `command_timeout=1800`; no-change soft limit `600` | — |
| Agent loop | `step_limit=250`, `cost_limit=3.0` | `max_iterations=250` | — |
| Whole inference task | `--execution-params.run_timeout_seconds=null` | `--execution-params.run_timeout_seconds=9600` | — |
| Fresh official evaluation | — | — | `--execution-params.evaluation_timeout_seconds=3600` |
| Repeated attempts | — | — | `--k`, `--attempt-strategy` |

`evaluation_timeout_seconds` controls only the fresh `run_script.sh` and parser evaluation after patch collection. It cannot extend inference. Thinking/reasoning belongs in `--model-params`; use the protocol/provider form documented for [mini-SWE-agent](/en/user_guide/modules/harnesses/mini_swe_agent#thinking-and-reasoning) or [OpenHands](/en/user_guide/modules/harnesses/openhands#thinking-and-reasoning).

## Run examples

`agentcompass run` takes three positional arguments in order: Benchmark, Harness, and Model. The examples use `swebench_pro_verified`; harness choices are described below.

Before running, make sure local [Docker](/en/user_guide/modules/environments/providers/docker) is available and set `MODEL_NAME`, `MODEL_BASE_URL`, and `MODEL_API_KEY` to the model under test, API endpoint, and API key.

### Recommended harness

[mini-SWE-agent](/en/user_guide/modules/harnesses/mini_swe_agent) is the recommended harness for SWE-bench Pro. It uses the benchmark-specific mini-SWE-agent configuration and executes repository commands in the task environment.

<Tabs>
  <Tab title="Smoke test (single task end-to-end)">
    Run one task to verify inference, patch collection, and official evaluation end to end.

    ```bash wrap theme={"system"}
    agentcompass run \
      swebench_pro_verified \
      mini_swe_agent \
      "$MODEL_NAME" \
      --env docker \
      --benchmark-params '{
        "sample_ids": ["instance_NodeBB__NodeBB-04998908ba6721d64eba79ae3b65a351dcfbc5b5-vnan"]
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat
    ```
  </Tab>

  <Tab title="Custom parameters">
    Run three attempts for one task and customize the attempt policy, model request, command, task, and evaluation limits.

    ```bash wrap theme={"system"}
    agentcompass run \
      swebench_pro_verified \
      mini_swe_agent \
      "$MODEL_NAME" \
      --env docker \
      --k 3 \
      --attempt-strategy pass \
      --benchmark-params '{
        "sample_ids": ["instance_NodeBB__NodeBB-04998908ba6721d64eba79ae3b65a351dcfbc5b5-vnan"]
      }' \
      --execution-params '{
        "evaluation_timeout_seconds": 4800,
        "run_timeout_seconds": 14400
      }' \
      --harness-params '{
        "step_limit": 300,
        "cost_limit": 5.0,
        "command_timeout": 2400
      }' \
      --model-params '{
        "temperature": 0,
        "max_tokens": 32768,
        "timeout": 3600,
        "reasoning_effort": "high"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat
    ```
  </Tab>

  <Tab title="AgentCompass recommended config">
    Evaluate the complete public split with explicit inference and evaluation limits. Adjust `--task-concurrency` only when required by provider capacity.

    ```bash wrap theme={"system"}
    agentcompass run \
      swebench_pro_verified \
      mini_swe_agent \
      "$MODEL_NAME" \
      --env docker \
      --execution-params '{
        "evaluation_timeout_seconds": 3600,
        "run_timeout_seconds": 12000
      }' \
      --harness-params '{
        "step_limit": 250,
        "cost_limit": 3.0,
        "command_timeout": 2400
      }' \
      --model-params '{
        "temperature": 0,
        "max_tokens": 32768,
        "timeout": 3600,
        "reasoning_effort": "high"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>
</Tabs>

### Other optional harnesses

[OpenHands](/en/user_guide/modules/harnesses/openhands) is also supported. The following command evaluates the full dataset and exposes its independent model-request, terminal-command, agent-loop, whole-task, and evaluation limits:

```bash wrap theme={"system"}
agentcompass run \
  swebench_pro_verified \
  openhands \
  "$MODEL_NAME" \
  --env docker \
  --execution-params '{
    "evaluation_timeout_seconds": 3600,
    "run_timeout_seconds": 12000
  }' \
  --harness-params '{
    "max_iterations": 250,
    "conversation_timeout": 3600,
    "command_timeout": 1800,
    "terminal_no_change_timeout_seconds": 600
  }' \
  --model-params '{
    "temperature": 0,
    "max_output_tokens": 32768,
    "timeout": 3600,
    "reasoning_effort": "high",
    "num_retries": 10,
    "retry_min_wait": 8,
    "retry_max_wait": 64,
    "retry_multiplier": 2
  }' \
  --model-base-url "$MODEL_BASE_URL" \
  --model-api-key "$MODEL_API_KEY" \
  --model-api-protocol openai-chat
```

<a id="outputs" />

## Evaluation Results

For shared result conventions, see [Run Directory](/en/user_guide/other_features/results/overview#directory-layout), [Aggregate Scores](/en/user_guide/other_features/results/summary_analysis), and [Task Files and Shared Fields](/en/user_guide/other_features/results/task_results).

<a id="aggregate-metrics" />

### Scoring Metrics

SWE-bench Pro Verified's primary metric is binary `correct`, matching the evaluator's `resolved` decision. Resolution follows the test rules in [How It Works](#how-it-works); there is no partial credit.

With the default configuration, each task has one attempt and the overall score is the issue resolution rate over tasks with valid scores, ranging from 0 to 1; higher is better.

See [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation) for repeated attempts, category aggregation, and scoring failure rules.

<a id="per-task-details-details" />

### Task Results and Scoring Evidence

`final_answer` contains the unified diff patch submitted to the evaluator. `eval_raw_data` under `meta.benchmark` preserves scoring evidence, with these fields when available:

| Field | Contents |
| - | - |
| `resolved` | The evaluator’s issue resolution decision. |
| `completed` | Whether the evaluator completed its decision. |
| `fail_to_pass`, `pass_to_pass` | Required fix tests and regression tests. |
| `fail_to_pass_missing`, `pass_to_pass_missing` | Required tests absent from the passing set, identifying why an issue remains unresolved. |
| `tests`, `raw_output` | Parsed test list and complete parsed output. |
| `stdout`, `stderr`, `returncode` | Test logs and the evaluation process exit code. |
| `error`, `timed_out` | Diagnostics for patch application or evaluation problems, present only on the relevant paths. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.