> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Run Controls

`agentcompass run` and `agentcompass launch` use the same run controls for scheduling, fault handling, and evaluation artifacts. `execution.task_concurrency` is the single concurrency setting for physical attempt executions; there is no separate task-versus-attempt concurrency knob.

This page explains what each control does and how to use it. See [`agentcompass config`](/en/user_guide/using_agentcompass/cli/config) for configuration-file syntax and precedence, [`agentcompass run`](/en/user_guide/using_agentcompass/cli/run#parameter-reference) for the complete single-request signatures, and [`agentcompass launch`](/en/user_guide/using_agentcompass/cli/launch#validate-before-running) for multi-request orchestration and its CLI overrides.

| Goal | Primary options |
| - | - |
| Control attempt concurrency and provider capacity | `--task-concurrency`, `--provider-limit`, `--env-open-qps` |
| Limit the duration of the evaluation execution phase | `--timeout-seconds` |
| Handle recoverable transient failures | `--max-retries`, `--retry-pattern-list` |
| Organize results and reuse completed tasks | `--results-dir`, `--run-name`, `--run-id`, `--reuse` |
| Preserve state and diagnostics | `--keep-environment`, `--progress`, `--log-level`, `--file-log-level` |

## Scale Concurrency Safely

A [provider](/en/user_guide/modules/environments/overview#choose-a-provider) is the execution backend that creates and manages environments, such as Docker, Daytona, or Modal.

| Control | Scope |
| - | - |
| `--task-concurrency` | Physical attempts executing at the same time, including retries and different `k` attempts of one task when parallel execution is safe. |
| `--provider-limit <provider=count>` | Physical attempt executions handled concurrently by one provider; `0` disables the limit. |
| `--env-open-qps <provider=qps>` | New environments created per second by one provider; `0` disables startup pacing. |

Effective attempt concurrency is bounded by the lower of `task_concurrency` and the applicable provider limit. `env-open-qps` controls only Environment startup pacing. With `strategy: avg`, same-task attempts share this pool and may overlap only when both the Benchmark and Harness declare their state isolated; otherwise those attempts remain serial. Model endpoint capacity, provider quotas, and local CPU and memory can reduce actual concurrency further. CPU and memory limits for one sandbox are Environment parameters; see [Understand the Scope](/en/user_guide/modules/environments/configuration/resource_limits#understand-the-scope).

### CLI Syntax

In the CLI, repeat the latter two options for different providers. When evaluation requests use different environments, apply limits to each provider from one command:

```bash theme={"system"}
agentcompass launch evaluations.yaml \
  --task-concurrency 32 \
  --provider-limit docker=8 \
  --provider-limit modal=24 \
  --env-open-qps modal=4
```

A single `agentcompass run` needs limits only for the providers actually used by that request.

### Configuration File Syntax

In a [`--config` file](/en/user_guide/using_agentcompass/cli/config), use mappings for provider limits instead of repeating YAML keys:

```yaml theme={"system"}
runtime:
  provider_limits:
    docker: 8
    modal: 24
  env_open_qps:
    modal: 4

execution:
  task_concurrency: 32
```

The example above is a regular run configuration. In a `launch` orchestration file, put shared `task_concurrency` at the top level while keeping the provider mappings under `runtime`; see [`agentcompass launch`](/en/user_guide/using_agentcompass/cli/launch#what-the-fields-mean).

When tuning concurrency, first select a few representative Benchmark tasks and validate them with task concurrency set to `1`, then increase it gradually to `2` or `4`. Observe Environment startup latency, model latency, error rates, and memory use. Return to the last stable value when errors increase.

## Set an Appropriate Timeout

Use `--timeout-seconds` to limit the whole run and `--execution-params` to configure individual task phases. Environment startup, single commands, and model requests can also have their own limits. When limits overlap, the first deadline reached ends the corresponding work.

### Overall and component deadlines

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'900px', width:'100%', tableLayout:'fixed', fontVariantLigatures:'none'}}>
    <thead>
      <tr><th style={{width:'17%', whiteSpace:'nowrap'}}>Layer</th><th style={{width:'30%'}}>Parameter location</th><th style={{width:'53%'}}>Scope</th></tr>
    </thead>

    <tbody>
      <tr><td style={{whiteSpace:'nowrap'}}>Evaluation deadline</td><td>CLI: <code>--timeout-seconds</code> <code>\<seconds></code></td><td>All tasks in one <code>run</code> share this limit, as do all requests in one <code>launch</code>. Timing starts after component preflight and covers task loading, preparation, execution, <a href="/en/user_guide/using_agentcompass/cli/analysis#run-with-evaluation">analysis</a>, and summarization. On expiry, unfinished work is cancelled and resource cleanup begins. The default is <code>360000</code> seconds (100 hours). Explicitly setting <code>0</code> disables the evaluation deadline; it does not affect the component-specific timeouts below.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}>Environment creation</td><td>JSON field: <code>setup.build\_timeout\_seconds</code><br />Passed through <code><span>-</span><span>-</span>env-params</code></td><td>A shared startup deadline for every Environment, including <a href="/en/user_guide/modules/environments/providers/daytona#provider-params">Daytona</a> and <a href="/en/user_guide/modules/environments/providers/modal#provider-params">Modal</a>. Each sandbox creation is timed separately. Expiry fails only that creation and does not limit later operations in a successfully created sandbox.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><a href="/en/user_guide/modules/harnesses/overview#configure-harness-parameters">Harness-specific</a></td><td>JSON field: defined by the Harness<br />Passed through <code><span>-</span><span>-</span>harness-params</code></td><td>The exact scope depends on the field. <code>command\_timeout</code> limits one command; <code>request\_timeout</code> limits one service request. Use the common execution fields below for a task phase deadline.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><a href="/en/user_guide/modules/benchmarks/overview#configure-benchmark-parameters">Benchmark-specific</a></td><td>JSON field: defined by the Benchmark<br />Passed through <code><span>-</span><span>-</span>benchmark-params</code></td><td>The exact scope depends on the field. For example, PinchBench <code>judge\_timeout\_seconds</code> limits one judge-model request.</td></tr>
    </tbody>
  </table>
</div>

### Per-task phase deadlines

Configure phase timeouts through `run --execution-params`, YAML `execution`, or SDK `execution_params`. For `launch`, use `defaults.execution` or `requests[].execution`.

The diagram shows phase budgets for one attempt with a Harness and a fresh evaluation environment, rather than the overall task lifecycle deadline. Phases run from left to right; their widths do not represent duration. Some phases may be skipped; downloads and restores depend on the evaluation mode and artifact settings.

<img src="https://mintcdn.com/opencompass-docs-preview-pr-335-0/EIEK6EmKxREMKTr7/images/agentcompass_timeout_coverage.svg?fit=max&auto=format&n=EIEK6EmKxREMKTr7&q=85&s=b6db8fde2b37e311ac3e8c3e5d5ff838" alt="Timeout coverage across a single task attempt: execution, collecting outputs and building RunResult, nested artifact preparation limits, independently timed batch download and restore, and evaluation." style={{ width: "100%", height: "auto" }} width="1840" height="560" data-path="images/agentcompass_timeout_coverage.svg" />

Gray columns mark phases outside the six timeout settings shown: environment initialization, closing the Harness, evaluation environment initialization, and cleanup. Other Environment or lower-level timeouts may still apply.

| Field | Default | Scope |
| - | - | - |
| `run_timeout_seconds` | `null` | Entire Harness `execute_task()` call, or Benchmark `run_task()` without a Harness. |
| `evaluation_timeout_seconds` | `null` | Each Benchmark `evaluate()` call. |
| `timeout_multiplier` | `1.0` | Common multiplier for execution and evaluation. |
| `run_timeout_multiplier` | `null` | Replace the common multiplier for execution. |
| `evaluation_timeout_multiplier` | `null` | Replace the common multiplier for evaluation. |
| `harness_result_timeout_seconds` | `60` | Collect outputs and build `RunResult`. |
| `artifact_collect_timeout_seconds` | `null` | Total budget for all artifact preparation commands; `null` sets no group deadline. |
| `artifact_collect[].timeout_seconds` | `60` | Budget for each preparation command. |
| `artifact_limits.timeout_seconds` | `600` | Budget for each complete batch download or restore. |

Time budgets are in seconds. The execution budget covers preparation inside `execute_task()`, model requests, and tool execution. Environment/session setup before the call and result collection afterward have separate timing. Evaluation starts its clock after the evaluation environment is ready and artifacts have been restored.

For execution and evaluation, an omitted or `null` override inherits the Benchmark plan default, then the task default. Execution also falls back to the Harness default. If none supplies a budget, the phase has no deadline.

The final budget is `base seconds × effective multiplier`. A phase-specific multiplier replaces `timeout_multiplier`; when it is `null`, the common multiplier applies. These multipliers do not multiply together. They affect only execution and evaluation; result collection, artifact preparation, download, and restore keep their own budgets. A multiplier cannot create a deadline when no base budget exists.

```bash theme={"system"}
agentcompass run terminal_bench_2 claude_code "$MODEL_NAME" \
  --env docker \
  --execution-params '{
    "run_timeout_seconds": 3600,
    "evaluation_timeout_seconds": 600,
    "timeout_multiplier": 2,
    "evaluation_timeout_multiplier": 1
  }'
```

This allows 7,200 seconds for execution and 600 seconds for evaluation per attempt. On timeout, the runtime records the failure and attempts applicable result collection and cleanup. Collecting a partial answer does not clear the timeout or guarantee evaluation. These budgets do not extend sandbox lifetime or the outer `--timeout-seconds` deadline.

### Collect outputs and build RunResult

`harness_result_timeout_seconds` limits the Harness `collect_result()` call after successful execution, an execution error, or an execution timeout. This call reads logs, saved state, and requested output files; extracts the answer; and assembles the trajectory, status, and diagnostics into `RunResult` for persistence, analysis, and evaluation. Some Harnesses also write trajectory files during collection.

Collection has its own budget because it may require sandbox I/O after execution has ended. It does not run the agent or request further model responses. It finishes before the Harness session closes; artifact preparation and download follow while the task sandbox remains available. Harness-free Benchmarks skip this hook. A collection failure is recorded alongside any original execution error.

### Prepare and transfer artifacts

Preparation commands run in order under two limits: the whole list shares `artifact_collect_timeout_seconds`, and each command has its own `timeout_seconds`. The first deadline reached stops preparation. For example, three commands with 60 seconds each and a 100-second group budget may use at most 60 seconds per command and 100 seconds in total. With a `null` group budget, only the individual command limits apply.

For transfers, one batch download shares `artifact_limits.timeout_seconds` across all declared files. A subsequent restore starts a new clock with the same budget. At the default, that means up to 600 seconds for download and another 600 seconds for restore.

<span id="post-run-budgets" />

### Configure post-run budgets

This YAML allows 120 seconds for building `RunResult`, 100 seconds for preparation, and 1,800 seconds per batch transfer. It inherits the declared paths, commands, and per-command limits:

```yaml theme={"system"}
execution:
  harness_result_timeout_seconds: 120
  artifact_collect_timeout_seconds: 100
  artifact_limits:
    timeout_seconds: 1800
```

To remove the group preparation deadline while retaining per-command limits, override it with `null`:

```bash theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME" \
  --execution-params '{"artifact_collect_timeout_seconds":null}'
```

CLI and SDK mappings merge over YAML; omitted fields inherit and supplied lists replace existing lists. Dedicated CLI options take precedence over the corresponding `--execution-params` fields. The SDK accepts `execution_params` in `build_run_request()`, `run_evaluation()`, and `async_run_evaluation()`.

### Choose the budget to adjust

| Where work times out | Control to inspect |
| - | - |
| Agent inference or tool loop | `run_timeout_seconds` and its multiplier; check for a shorter model-request or command limit. |
| Collecting logs, answers, or trajectories | `harness_result_timeout_seconds`. |
| Preparing the submission | Both the command's `timeout_seconds` and the group's `artifact_collect_timeout_seconds`. |
| Batch download or restoration | `artifact_limits.timeout_seconds`. |
| Benchmark grading | `evaluation_timeout_seconds` and its multiplier; check for a shorter judge-request limit. |

<span id="collect-submission-artifacts" />

## Configure Artifacts

Use `execution` to control which artifacts are saved, how they are prepared, and whether evaluation restores them. Preparation and transfer deadlines are described under [timeouts](#prepare-and-transfer-artifacts).

### Save and restore artifacts

For Benchmarks with artifact collection enabled and declared paths, transfers depend on the evaluation mode:

| Evaluation mode | Default `save_artifacts` | Batch download | Batch restore |
| - | - | - | - |
| `fresh`: a new evaluation environment | `true`; saving is required | Save declared artifacts. | Restore into the new environment before evaluation. |
| `reuse`: the execution environment | `false` | Only with `save_artifacts=true`. | Skip; use files already in the environment. |
| `none`: no evaluation environment | `false` | Only with `save_artifacts=true`. | Skip. |

An omitted or `null` `save_artifacts` uses the mode's default. Skipped transfers consume no transfer budget. Declared preparation commands and Harness result collection still run when downloads are disabled. Saving artifacts keeps local files; use `keep_environment` to retain the sandbox itself.

Artifact size/count limits default to 16 GiB and 100,000 entries per collection or restore. Under `artifact_limits`, set `max_mb` and `max_files` to adjust them. `max_mb` covers the whole artifact list and also caps each transfer archive, including archive overhead; one MB is 1,048,576 bytes.

Preparation or transfer failures preserve validated files and execution diagnostics, but skip grading and creation of a resumable checkpoint for that attempt. Progress records report the operation, completed bytes, and remaining budget.

### Override artifact paths and preparation commands

Benchmark declarations supply the default paths and commands. Override them under `execution`:

| Field | Omitted or `null` | Supplied list |
| - | - | - |
| `artifacts` | Inherit declared paths. | Replace extra paths; `/logs/artifacts/` remains included. `[]` selects only that directory. |
| `artifact_collect` | Inherit preparation commands. | Replace the ordered command list. `[]` clears it. |

Include any required original paths or commands in a replacement list. These overrides require a Benchmark that supports artifact collection; see [Submission Artifacts](/en/developer_guide/extensions/benchmark/code_implementation/shared_contracts#submission-artifacts-and-replay).

Use absolute sandbox paths for `source` and relative local paths for `destination`. Commands run in the Environment's default workdir; use absolute paths or an explicit `cd` when they depend on the task workspace.

```bash theme={"system"}
agentcompass run deepswe claude_code "$MODEL_NAME" \
  --env docker \
  --model-base-url "$MODEL_BASE_URL" \
  --model-api-key "$MODEL_API_KEY" \
  --execution-params '{
    "save_artifacts": true,
    "artifacts": [
      {"source": "/app/submission", "destination": "submission", "exclude": ["*.tmp"]},
      {"source": "/app/result.json", "destination": "result.json"}
    ],
    "artifact_collect": [
      {"command": "mkdir -p /app/submission && printf example > /app/submission/note.txt", "timeout_seconds": 60}
    ]
  }'
```

This example replaces the preparation commands with one that creates a demonstration file. It saves `/app/submission` under the attempt's local `artifacts/submission/`, excluding `*.tmp`, and `/app/result.json` as `artifacts/result.json`. Missing sources are recorded without failing collection. Fresh evaluation restores files to their original `source` paths.

## Retry Only Transient Failures

`--max-retries` sets the maximum retries inside each logical attempt. For example, `--max-retries 2` permits up to two replacement executions after that attempt's initial failure. A retry never creates a new metric attempt.

`--retry-pattern-list` matches each ERROR issue message and code separately; null and \[] mean FATAL-only retries. FATAL always retries within the shared budget and WARNING never retries.

Retry only transient errors that may recover on another execution, such as dropped network connections, temporary service failures, or sandbox timeouts:

```bash theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME" \
  --env <environment> \
  --max-retries 2 \
  --retry-pattern-list '["(?i)connection.*reset","(?i)temporar","(?i)sandbox.*timeout"]'
```

Judge invalid JSON, missing credentials and Environment failures are FATAL. The default budget is 0; use --max-retries 2 to tolerate transient failures. Unresolved FATAL fails the run and removes official scores.

Retries restart only the current logical attempt, or only its evaluation phase when the retry scope permits. Completed sibling attempts stay checkpointed: if attempt 3 retries, attempts 1 and 2 are not rerun. The final details record `retry_count` and `retry_counts`; discarded executions remain under the attempt’s `retries/` directory for diagnosis.

## Output and Reuse

### Name a New Run

For `agentcompass run`, the three options correspond to different levels of the result path:

```text theme={"system"}
<results-dir>/[<run-name>/]<model>_<benchmark>_<harness>/<run-id>/
```

* `--results-dir` sets the result root and defaults to `results`.
* `--run-name` adds an optional experiment-group directory.
* `--run-id` names this run's directory; the current timestamp is used when it is omitted.

Model, Benchmark, and Harness IDs are normalized to safe names and joined with underscores into one directory component. For `agentcompass launch`, each request uses its normalized `name` in place of that combined component:

```text theme={"system"}
<results-dir>/[<output.run_name>/]<requests[].name>/<run-id>/
```

See [`agentcompass launch`](/en/user_guide/using_agentcompass/cli/launch#mapping-rules) for request naming and collision checks.

The following command uses `ablation` as the experiment group and gives this run the fixed name `baseline`:

```bash theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME" \
  --env <environment> \
  --run-name ablation \
  --run-id baseline
```

With the default result root, the path is:

```text theme={"system"}
results/ablation/<model>_<benchmark>_<harness>/baseline/
```

See [Understanding Evaluation Results](/en/user_guide/other_features/results/overview) for the complete directory and file layout.

### Resume an Interrupted Run

Use `--reuse` to continue an evaluation from an existing run. AgentCompass reuses compatible complete details without errors by task ID and can materialize valid terminal-attempt checkpoints for an unfinished multi-attempt task:

```bash theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME" \
  --env <environment> \
  --reuse
```

Without a value, `--reuse` selects the latest run under this hierarchy:

```text theme={"system"}
<results-dir>/[<run-name>/]<model>_<benchmark>_<harness>/
```

Pass a run ID to select an exact source under that hierarchy:

```bash theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME" \
  --env <environment> \
  --reuse 20260806_120000
```

For `run`, AgentCompass searches only within the same result root, optional run-name prefix, and combined Model/Benchmark/Harness directory. For `launch`, it searches within the request's named output namespace. Existing directories remain untouched; automatic reuse does not search the former Benchmark/Model hierarchy.

The source must use a supported run schema, the same Benchmark ID, and the same attempt plan. Task IDs and attempt numbers identify reusable data; request parameters and task inputs do not need to be identical. You can change evaluation timeouts, resources, or environment variables when continuing a pending fresh evaluation. The new run records the source and preserves reused details or checkpoints; see [reuse validation](/en/user_guide/other_features/results/run_records#reuse-identity).

Reuse is decided per logical attempt using the current task plan. Complete WARNING results and ERROR results whose message/code do not match the current patterns remain reusable. FATAL or matching ERROR results enter the current retry policy with a new shared budget; historical retry counts do not consume it. Zero budget does not turn a failed result into success. Only evaluation failures with complete none/fresh snapshots can resume scoring; reuse mode reruns the attempt.

### Resume Evaluation from Saved Artifacts

After inference and required collection complete without FATAL, none/fresh modes save a v4 checkpoint containing classified inference and prepared-input snapshots, identity, provenance, network policy and artifact integrity. Recoverable ERROR results are allowed. Fresh recovery restores captured artifacts; none recovery verifies declared local input paths and digests. Incomplete transfers or unserializable runtime objects prevent recovery.

| Source state | Default `--reuse` behavior |
| - | - |
| Complete result, no FATAL and no matching ERROR | Reuse the result and its valid score. |
| Evaluation needs retry, complete none/fresh snapshot | Apply the current budget and restore scoring inputs; fresh uses a new verifier. |
| Incomplete input, unsupported checkpoint or reuse mode | Schedule the necessary attempt execution. |
| Interrupted multi-attempt task | Keep reusable terminal attempts and recover only the pending work. |

The new run copies validated checkpoints and captured artifacts. Recovery validates task/attempt identity, source compatibility, network policy and input integrity, then uses the current plan and a fresh copy of the saved inputs. Driver-side file dependencies must still exist and match their digests. Recovery does not restore a live process or sandbox.

To disable automatic checkpoint recovery for pending tasks, add `--no-checkpoint-resume` to `agentcompass run`, pass `checkpoint_resume=False` to the Python SDK, or configure:

```yaml theme={"system"}
runtime:
  checkpoint_resume: false
```

For multi-request launches, set `defaults.runtime.checkpoint_resume` or `requests[].runtime.checkpoint_resume`. The effective default is `true`. This option does not force already completed results to be evaluated again, disable terminal-attempt scheduling checkpoints, or prevent new checkpoints from being saved.

## Keep Environments for Debugging

Add `--keep-environment` when a failure requires direct inspection of task or verifier sandboxes:

```bash theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME" \
  --env <environment> \
  --keep-environment
```

AgentCompass then skips provider cleanup for environments created by the run. Retries and multiple tasks may leave several resources active, so release them later with the provider's tooling. Harness sessions are still closed normally.

## Logs and Progress

| Option | Default | Accepted values | Purpose |
| - | - | - | - |
| `--progress <mode>` | `auto` | `auto`, `plain`, `none` | Controls terminal progress: `auto` shows a live view only in an interactive terminal, `plain` prints text progress suitable for CI or redirected logs, and `none` disables terminal progress. |
| `--log-level <level>` | `INFO` | `DEBUG`, `INFO`, `WARNING`, `ERROR`, `CRITICAL` | Sets the minimum console log level. |
| `--file-log-level <level>` | `DEBUG` | `DEBUG`, `INFO`, `WARNING`, `ERROR`, `CRITICAL` | Sets the minimum level for run logs saved in the result directory. |

`--progress` controls only terminal rendering. AgentCompass still saves progress, logs, and task results in every mode.
See [Results](/en/user_guide/other_features/results/overview#directory-layout) for their locations.

## Environment variable scopes

Configure all variable scopes through `--env-params` (SDK: `environment_params`; YAML: the selected provider under `environments`). `env_variables` supplies common startup/command bindings; `run_env_variables` covers agent setup/run; `evaluation_env_variables` covers evaluator commands. `evaluation_environment_env_variables` independently overrides fresh-verifier startup bindings and requires fresh mode. The former Harness `env` input is rejected.

```bash theme={"system"}
agentcompass run <benchmark> <harness> "$MODEL_NAME" \
  --env docker \
  --env-params '{"evaluation_environment_mode":"fresh","env_variables":{"LANG":"C.UTF-8"},"run_env_variables":{"AGENT_MODE":"test"},"evaluation_environment_env_variables":{"VERIFIER_MODE":"offline"},"evaluation_env_variables":{"GRADER_KEY":"${GRADER_KEY}"}}'
```

The earlier `execution.download_artifacts` option has been replaced by `execution.save_artifacts`; the old configuration key is rejected.

## Error handling and score validity

`RunResult.issues` is the only execution error contract. Each issue has `severity` (`fatal`, `error`, `warning`), `phase` (`setup`, `run`, `collect`, `evaluate`, `cleanup`), a stable `code`, and a redacted `message`. `RunResult.error` has been removed. Full redacted tracebacks and exception chains belong in `artifacts.execution_diagnostics`.

FATAL reports setup, Environment, external tool service, model authentication/quota/server, or Judge failures. Judge timeouts, invalid JSON and missing required scoring fields are FATAL. Model API or Harness timeouts and empty model responses are ERROR. Missing model deliverables, ordinary verifier reward files that are absent or malformed, and cleanup failures are WARNING. A known trusted setup failure remains FATAL even when no reward is produced.

ERROR does not override a valid Benchmark score. Ordinary verifier reward failures remain valid fail/0 observations; the Judge protocol exception above is FATAL. Runtime never fills missing observations with zero based on status alone.

`execution.max_retries` defaults to **0**, and is one shared extra retry budget per logical attempt. FATAL retries whenever budget remains. ERROR retries only when a configured regex matches its `message` or `code`; WARNING never triggers a retry. Both `null` and `[]` mean FATAL-only retries. Consider `execution.max_retries=2` for transient service failures. A zero budget means one unresolved FATAL fails the run. Patterns no longer match traceback text, multiple concatenated issues, or redacted secrets; migrate them to stable codes or retained diagnostic features.

The resolved task plan supplies both budget and patterns. In `none` and `fresh` modes an evaluation retry uses an isolated inference snapshot; `fresh` creates a new evaluation Environment. In `reuse` mode it reruns the whole attempt. Cross-run recovery uses the new run's budget; previous retry counts remain history. Resolved failures remain in retry history and are absent from final issues.

Any logical attempt with a final FATAL invalidates its entire task across every metric, even when another attempt succeeded. `MetricReport.evaluation_failed` then becomes true, all official `value` fields are null, and available values move to explicitly labelled `reference_value` fields. Reference scores exclude invalidated tasks from numerators, denominators and weights. With no valid tasks both values are null. The run is `failed`, CLI exits nonzero, and SDK failure messages retain result paths. Other requests in an orchestration continue.

Counts satisfy `evaluated + unavailable + invalidated = total`. `error` is a separate diagnostic count. Attempt coverage counts each final ERROR/FATAL attempt once, regardless of how many issues it has. WARNING and resolved retry history do not count. Legitimate `pass@k` early stopping is not a missing attempt.

Current task and run-info schemas are v3; evaluation checkpoints are v4 and contain classified inference and prepared-input snapshots, network policy and artifact integrity data. Previous v2 results are converted only at the read boundary using reliable structured evidence. Historical free-text failures remain classification-unknown and require re-execution for current scoring; they remain viewable as historical results. Old checkpoints without classified snapshots are rejected for cross-run materialization with a reason and normal execution resumes. New-format missing `issues` or result-level `error` is a format error.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.