> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Shared Contracts

Convert dataset records into stable `TaskSpec` values first, then decide which fields the Harness may see and which state must remain evaluation-only.

## Distinguish the Four Data Carriers

| Carrier | Primary consumers | Store here | Do not store here |
| - | - | - | - |
| `TaskSpec` | Benchmark, Planner, Recipe | Stable task ID, question, category, upstream metadata, per-task evaluation mode, and network policies | Provider sessions or an already-open Environment |
| `BenchmarkPlan` | Benchmark and Planner for the current task attempt | Resolved configuration, workspace paths, verifier timeout, and typed evaluator state | Provider SDK clients or mutable global state |
| `PreparedTask` | Harness or `HarnessFreeBenchmark.run_task()` | Prompt, messages, files, media, tools, workspace, and expected output | Hidden answers, reference patches, private tests, or grading secrets |
| `RunResult` | Benchmark, runtime, and result consumers | Execution status, answer, trajectory, artifacts, scores, and publishable evaluation evidence | Non-serializable objects or credentials |

A Harness can read the entire `PreparedTask`, including its `ground_truth` and `metadata`. Do not copy `TaskSpec.metadata` into it without filtering.

Keep evaluator-only data in `TaskSpec.ground_truth` or a typed `BenchmarkPlan`, and set `PreparedTask.ground_truth` to `None`. These objects still belong to the runtime and result-audit boundary, so they must not contain credentials or other secret material that must never be persisted.

Write a reference answer to `RunResult.ground_truth` only when it is safe to publish with the result.

## Define Public Configuration

Define Benchmark parameters with `RuntimeBenchmarkConfig` and `config_field()`, and normalize types early in `__post_init__()`:

```python theme={"system"}
from dataclasses import dataclass

from agentcompass.benchmarks.config import RuntimeBenchmarkConfig
from agentcompass.runtime.config import config_field, parse_bool


@dataclass(slots=True)
class ExampleExactMatchConfig(RuntimeBenchmarkConfig):
    case_sensitive: bool = config_field(
        default=False,
        description="Compare answers with case sensitivity.",
    )

    def __post_init__(self) -> None:
        RuntimeBenchmarkConfig.__post_init__(self)
        self.case_sensitive = parse_bool(self.case_sensitive, "case_sensitive")
```

Do not download data, install dependencies, or read credentials during module import. Access data through an explicit loader or dependency-preparation path. If a revision, split, or access requirement is invalid, fail with an actionable message.

## Load Deterministic Tasks

`load_tasks()` should pin the upstream revision and produce stable `task_id` values. The following `metadata` contains only reproducibility information that may appear in logs and execution input; the answer remains separate in `ground_truth`:

```python theme={"system"}
def load_tasks(self, req: RunRequest) -> list[TaskSpec]:
    _ = req
    return [
        TaskSpec(
            task_id="capital-france",
            question="What is the capital of France? Answer with only the city name.",
            category="geography",
            ground_truth="Paris",
            metadata={"dataset_revision": "tutorial-v1"},
        )
    ]
```

The inherited `select_tasks()` method already applies the runtime's shared task-selection logic. Override it only when the Benchmark needs semantics beyond ordinary ID filtering. Whichever rule you use, keep the returned order deterministic.

### Harbor tasks and execution requirements

Use `load_harbor_task()` from `agentcompass.benchmarks.utils.harbor` to convert a Harbor task directory into a `TaskSpec`. In addition to network policies and resources, the adapter maps these execution requirements:

| Harbor declaration | AgentCompass task field | Enforcement |
| - | - | - |
| `agent.timeout_sec` | `run_timeout_seconds` | The `run_task` / `run_harness` phase |
| `verifier.timeout_sec` | `evaluation_timeout_seconds` | Each `evaluate` invocation |
| `environment.docker_image` | `environment_setup.image` | Provider-native registry image selection |
| `environment.workdir` | `environment_setup.workdir` | Default command execution directory inside the environment |
| `environment.build_timeout_sec` | `environment_setup.build_timeout_seconds` | Environment construction/startup, after the open-rate-limit queue |
| `environment.os` | `environment_setup.os` | Provider OS capability validation before allocation |
| `verifier.environment.network_mode` / `allowed_hosts` | `evaluation_baseline_network_policy` | Fresh-verifier startup and baseline reset, separate from execution-phase policy |
| `environment.env` | `env_variables` | Common task bindings |
| `verifier.environment.env` | `evaluation_environment_env_variables` | Fresh-verifier startup bindings |
| `verifier.env` | `evaluation_env_variables` | Evaluation-command bindings, not agent-command bindings |
| `artifacts` | `artifacts` | File/directory download and fresh-verifier restore |
| `verifier.collect` | `artifact_collect` | Ordered submission-generation commands in the agent sandbox |

An explicitly separate verifier environment supplies an independent image, workdir, and build-timeout baseline through `evaluation_environment_setup`: omitted values do not inherit the agent's settings. An explicitly empty baseline also remains independent. Only `None` (no separate baseline declared) inherits the resolved run environment setup, including Recipe defaults. Request-level `EnvironmentSpec.setup` and `evaluation_setup` override task defaults. Provider-native startup selectors are not public inputs. Working directory, startup timeout, environment variables, and outbound rules accept only the shared framework fields; provider-native aliases are rejected.

Registry images have one resolved source: `EnvironmentSpec.setup.image`. Harbor's `environment.docker_image` maps to the task setup; Recipes read the merged plan setup and fill a missing image there instead of reading or writing `params["image"]`. Providers translate this field into their native configuration only when building the environment config. Request parameters `image` and `evaluation_image` are rejected; configure `setup.image` and `evaluation_setup.image` directly. Use `--env-params '{"setup": {"image": "registry.example/runner:v1"}}'`.

Fresh-verifier networking follows the same independence rule: an explicit `verifier.environment` supplies its own baseline, including the default public network when the table is empty. An absent verifier environment inherits the agent baseline. Request-level `evaluation_baseline_network_policy` overrides the common request `baseline_network_policy`, which overrides the task's verifier baseline. The runtime opens the verifier with that baseline, switches to `evaluation_network_policy` for evaluation, and resets the baseline afterward. Neither policy is silently substituted for the other. You can override only the fresh-verifier baseline with `--env-params '{"evaluation_baseline_network_policy": {"network_mode": "no-network"}}'`; this request field requires fresh evaluation.

`workdir` must be an absolute POSIX path inside the environment. It applies when `env.exec()` omits `cwd`; an explicit `cwd` takes precedence. It does not replace an explicit Benchmark/Harness workspace. SWE-Marathon uses the unified workdir for its workspace, falling back to the Dockerfile's `WORKDIR` and then its configured workspace root. You can override the task default with `--env-params '{"setup": {"workdir": "/app"}}'`.

These are separate phase deadlines, not one combined task wall clock, a per-tool timeout, or a sandbox lifetime. `None` leaves a task deadline unspecified; positive finite numbers specify seconds. The adapter maps only explicitly declared timeouts and does not introduce Harbor's implicit defaults. A build timeout does not add Dockerfile-building support to a provider that only launches prebuilt images.

`BenchmarkPlan.run_timeout_seconds` and `evaluation_timeout_seconds` supply task/Benchmark defaults. The Planner resolves the request's `execution.run_timeout_seconds` and `evaluation_timeout_seconds`, then applies the phase-specific multiplier (`run_timeout_multiplier` or `evaluation_timeout_multiplier`) or the common `timeout_multiplier`. A phase multiplier replaces the common multiplier. An omitted or `null` override inherits the default. Harness fallback budgets apply only when no task/Benchmark run default exists.

The final `ExecutionPlan` supplies both the runtime phase watchdog and native Harness/evaluator timeout controls. Harness `execute_task()` (or a Harness-free Benchmark's `run_task()`) and `evaluate()` have independent clocks, with the same behavior in reused and fresh environments. `collect_result()` then recovers saved Harness output before session cleanup, under `harness_result_timeout_seconds` (default 60 seconds). Recovery cannot resume agent execution and does not clear the original execution error. Run/evaluation multipliers do not scale its budget. On timeout the runtime records the phase error and attempts the applicable artifact/cleanup path; cancellation alone does not confirm a remote process has stopped.

Artifact preparation has a separate `artifact_collect_timeout_seconds` deadline (default `null`); downloading/restoring artifacts uses `artifact_limits.timeout_seconds` (default 600 seconds). Configure these under `execution` in YAML, `execution_params` in the SDK, or `--execution-params` in `run`. These budgets do not extend sandbox TTL or override whole-run cancellation.

Unmapped declarations are stored as `TaskSpec.load_warnings`. The runtime emits them only after `select_tasks()`, so excluded samples do not produce warnings. Direct loader callers can inspect this list themselves. Truly unsupported declarations remain visible; do not mark a field mapped merely because it is retained in raw metadata.

Legacy `memory` and `storage` size strings under `environment` or `verifier.environment` count as mapped only when Harbor converts them into the corresponding MiB resource fields. Conflicting old/new values still fail validation; ignored legacy values still produce unmapped warnings.

Task provenance is preserved in `TaskSpec.metadata`: the top-level `source` maps to `metadata["source"]`, while package information, including `task.version`, maps to `metadata["task"]`. Explicit task declarations override same-named entries in the free-form metadata table. Package versions are non-empty strings, not necessarily semantic versions; older Harbor releases that omit this field use the adapter's compatibility validation. The top-level legacy `version` remains an alias for `schema_version`, not the package version. Original declarations remain unchanged in `harbor_raw`.

After Recipe application, the final `ExecutionPlan` is the only execution-time source. Harnesses and verifiers do not re-read task TOML, private timeout fields, or timeout metadata. Recipes must change `plan.run_timeout_seconds` and `plan.evaluation_timeout_seconds`; the Planner maps the run value to the native harness field.

### Task environment variables

The shared task and environment contracts have common, run-only, and evaluation-only bindings: `env_variables`, `run_env_variables`, and `evaluation_env_variables`. `TaskSpec.evaluation_environment_env_variables` supplies a separate fresh-verifier startup baseline: `None` inherits task common bindings, while an explicit map replaces them, including an empty map. Harbor maps the effective separate verifier environment's `env` table to this field, not to command-only bindings. Plans retain references such as `${GRADER_KEY}`; providers resolve them from the launcher's environment before execution. Export only the keys required by your selected tasks. Missing required references fail before allocation, without printing their values. `${NAME:-default}` and references embedded in a string are supported; shell commands are never expanded.

Request-level values override task defaults per key. Within the request, phase-specific bindings override common bindings; the active phase overrides internal `env.exec(env=...)` defaults. Native protocol requirements that cannot be overridden must be checked with `env.require_exec_env()`; conflicting names fail explicitly without printing values. Agent setup/run and evaluator execution use session-local, async-context-local scopes. The scope resets on success, error, or cancellation. Task preparation, result recovery, Harness cleanup, artifact commands, and artifact restore use common bindings only; agent bindings do not carry over to them. Common bindings can also reach sandbox startup where supported; phase-only bindings are applied to commands, not the image entrypoint. This is an injection contract, not a security boundary against an agent that persists processes or files in a reused sandbox. `host_process` continues to inherit the host process environment.

`EnvironmentSpec.evaluation_environment_env_variables` is the request override for fresh startup bindings. It requires `evaluation_environment_mode="fresh"` and merges per key after the task verifier baseline and request common bindings. Evaluation-only request bindings take precedence during evaluator commands. An omitted request map inherits; an empty request map adds no overrides. It does not replace the task baseline, unlike an explicit task map.

The former Harness `env` parameter is rejected; migrate it to `environment.run_env_variables`. Harnesses obtain resolved values from the scoped Environment when adapting nested native tool configurations. They must not mutate `RunRequest` or the host `os.environ`. Local SDK calls need explicit SDK configuration; only commands executed through the Environment receive these bindings.

Harbor's `solution.env` belongs to the reference solution by default. A Benchmark must explicitly opt in with `solution_env_for_agent=True` when those bindings really belong to its agent tasks. Do not inject reference-solution secrets automatically.

### Submission artifacts and replay

`TaskSpec.artifacts` is the resolved file selection supplied by the Benchmark or format adapter. A Benchmark with `collect_artifacts = true` always includes the conventional main `/logs/artifacts/` directory. The Harbor adapter resolves that same directory while loading its task format, including for an empty Harbor declaration, and appends explicit extra paths; native Benchmarks receive the same treatment during planning. `TaskSpec.artifacts = []` or `execution.artifacts = []` means no extra paths, not an opt-out. Only `collect_artifacts = false` disables collection. An explicit convention-directory entry supplies its own exclusions. Destination collisions keep the earlier entry with a warning, as in Harbor; the shared runtime receives a non-overlapping list.

Benchmark declarations provide the collection defaults. `execution.artifacts` replaces task-specific extra paths when supplied, but `/logs/artifacts/` remains for a collecting Benchmark; omitted or `null` inherits, and `[]` selects only that conventional directory. `execution.artifact_collect` independently replaces commands, where `[]` clears the command list. Runtime uses the resolved lists; `collect_artifacts = false` is the only way to skip collection. `collect_artifacts` is a read-only resolved-plan value, not a CLI, YAML, or SDK execution option. `execution.save_artifacts` controls persistent local storage: omitted or `null` resolves to `true` for fresh and `false` for reuse/none. Fresh requires storage; explicitly setting `save_artifacts=false` fails before environment allocation. In reuse/none, disabling storage skips file download but still runs declared preparation commands. Enabling storage alone does not select any additional paths. Storage does not preserve the sandbox (`keep_environment` controls that). Answer/trajectory/log persistence and Benchmark-specific evaluator outputs are unaffected. Evaluation-only recovery supports none/fresh when a classified inference snapshot and complete scoring inputs are available. Reuse requires a new attempt.

When an artifact path depends on the final workspace selected by a Recipe, override `BaseBenchmark.resolve_artifacts(task, req, plan)`. The Planner calls it after all Recipes and before final normalization; an explicit `execution.artifacts` list remains authoritative and bypasses Benchmark defaults.

### Override artifact paths and preparation commands

`execution.artifacts` and `execution.artifact_collect` independently override task/Recipe extras and commands. Omitted or `null` inherits the corresponding declaration; a supplied artifact list replaces the extra paths but retains `/logs/artifacts`, including `[]`, which selects only that directory. CLI lists replace corresponding YAML lists. These settings can also supply paths or commands for a Benchmark with no task-specific defaults, provided its collection capability is enabled. A Benchmark class can declare `collect_artifacts = False`; either explicit override (including `[]`) then fails during planning with an `unsupported` error. Omitted/`null` values are not overrides. This capability defaults to `True` and is not a Benchmark params or execution CLI option. Recipes cannot re-enable it. To retain task-specific entries while adding your own, include the required original entries in the replacement list.

```bash theme={"system"}
agentcompass run deepswe claude_code "$MODEL_NAME" \
  --env docker \
  --model-base-url "$MODEL_BASE_URL" \
  --model-api-key "$MODEL_API_KEY" \
  --execution-params '{
    "save_artifacts": true,
    "artifacts": [
      {"source": "/app/submission", "destination": "submission", "exclude": ["*.tmp"]},
      {"source": "/app/result.json", "destination": "result.json"}
    ],
    "artifact_collect": [
      {"command": "mkdir -p /app/submission && printf example > /app/submission/note.txt", "timeout_seconds": 60}
    ]
  }'
```

`source` is an absolute path in the agent sandbox and may name a file or directory. `destination` is relative to the attempt's local `artifacts/` directory: `/app/submission` above is stored under `artifacts/submission/`, excluding `*.tmp`; the separate file is stored as `artifacts/result.json`. Missing sources are recorded without failing collection. Destination overlaps are rejected before execution. Fresh evaluation restores files to their original `source` paths.

Only the resolved command list runs, in order, before any shared file download; overriding commands does not automatically prepend the Benchmark commands. CLI commands use `timeout_seconds`; Harbor `task.toml` uses `timeout_sec`. Shared preparation and transfer limits apply to the resolved lists. `save_artifacts=false` still runs commands in reuse/none; fresh requires saving. The example command creates a demonstration file and replaces the Benchmark command list. Preserve any commands required by the verifier in your replacement list; otherwise its expected submission may not be produced.

Recipes may adjust the resolved paths. `None` and `[]` mean no resolved paths; they are not reinterpreted as Harbor declarations. Paths are revalidated after Recipe adjustment. To disable the entire collection pipeline, use the execution switch. With `host_process`, absolute sources refer to the host filesystem.

The default directory is optional: a missing source is recorded as `missing`, and a successfully downloaded empty directory is recorded as `collected` with zero content bytes. Neither case blocks evaluation or checkpoint creation. An empty directory is preserved for fresh restore; a missing entry does not upload files. Permission errors, broken ancestor paths, transport failures, timeouts, and size violations remain capture errors, not missing outputs. Their existing failure behavior is unchanged. Empty capture alone does not prove an independent verifier can grade the task; replay still needs every input the verifier requires.

`ArtifactSpec` lives in `agentcompass.runtime.artifacts`. A string declaration names an absolute sandbox source path. An object can also specify a relative artifact-store `destination` and `exclude` patterns. The default destination removes the source's leading slash. Collection supports files, binary data, and directories; exclusions use GNU tar patterns. Overlapping destinations, escaping paths/links, special files, and non-main service declarations fail explicitly.

The runtime first executes declared collection commands, if any. Disabling collection skips the entire pipeline. Otherwise, no commands means preparation is skipped and existing files can still be collected. With storage enabled, selected outputs are copied before releasing the agent environment. Binary contents stay in files below the run's `artifacts/` directory; `RunResult.artifacts.declared_artifacts` contains relative paths, status, sizes, and checksums. Keep that directory with the result JSON when moving a run. Result reuse verifies and copies referenced artifact bytes before publishing the reused detail or terminal attempt checkpoint; missing or corrupt payloads cause that result to rerun. Artifact bytes are not hard-linked between runs. Missing or fully excluded sources are recorded, leaving the verifier to score them. Transport errors and limit violations fail collection.

Runtime executes `ExecutionPlan.artifact_collect` in order inside the agent sandbox. Without commands, preparation is skipped; no Benchmark fallback hook runs. When storage is enabled and paths are selected, shared `download_artifacts()` handles packaging, transfer, validation, and local storage. An empty path list does not suppress declared commands.

Harbor's `verifier.collect` maps to `TaskSpec.artifact_collect`, which Planner copies and validates again after Recipe overrides. Each command maps `command`, `timeout_sec` to `timeout_seconds` (default 60 seconds), and `service` (default `main`). Commands run sequentially with `bash -c`, under the existing run-phase network policy and common environment bindings, before evaluation or environment teardown. They do not receive verifier-only environment variables; variable references expand in the sandbox, not during task loading. Each command has its own positive finite timeout in addition to the optional total collection-phase deadline. Sidecar services and command-level `user` overrides are not supported and fail during loading; unknown command fields also fail explicitly.

Commands still execute when only `save_artifacts` is disabled; paths can be selected without commands. Paths describe existing or expected submissions, not how to generate them. For example:

```toml theme={"system"}
artifacts = ["/logs/artifacts/result.txt"]

[[verifier.collect]]
command = "mkdir -p /logs/artifacts && cp /app/result.txt /logs/artifacts/result.txt"
timeout_sec = 300.0
```

Collection journals are written atomically alongside their submission directory as `submission-*.manifest.json`. Each successfully validated entry remains on disk if a later operation fails. `ArtifactDownloadError.manifest` contains the partial record with `complete=false`, failed/cancelled entries, and uncollected entries. The runtime preserves it in `declared_artifacts`, records preparation/transfer errors separately in telemetry `post_run_errors`, and retains any original run error. Preparation failure still permits best-effort collection of existing files, but either preparation or transfer failure prevents evaluation and creation of a resumable checkpoint for that attempt. Explicit cancellation propagates, leaving the journal for diagnosis. Progress events report the current operation, completed bytes and entries, archive size when known, and remaining budget; provider downloads do not expose continuous byte-level progress.

Fresh evaluation validates the manifest and checksums, then restores collected outputs to their original sandbox source paths before calling `evaluate()`. A restored directory replaces the target's contents, including removing stale and hidden entries, while preserving the target directory itself for mounted workspaces. Parent paths are restored before their descendants regardless of declaration or manifest order, so explicitly captured children survive parent exclusions. Fresh plans reject conflicting exclusion declarations for the same source. Multiple copies of one source must have matching capture status, content and permissions before any upload; identical copies are restored once. Restoration stages the payload first and attempts rollback on installation failure; interrupted recovery staging is retained until environment cleanup. Reuse evaluation never restores over the live workspace. Before rerunning an agent within the same logical attempt, runtime archives the previous artifacts, result and checkpoint under `retries/<id>/` and updates diagnostic artifact references. The new physical run starts with an empty artifact destination; evaluation-only resume preserves its captured inputs. Legacy inline text submissions remain supported when they cover all declared outputs. `ExecutionPlan.artifact_limits` defaults to 16 GiB, 100,000 filesystem entries, and 600 seconds per collection/restore operation; Recipes may override these typed limits. Local file workers finish or acknowledge cancellation before temporary-file cleanup.

The pinned DeepSWE v1.1 tasks declare `verifier.collect` commands that extract the diff from the base commit to `HEAD`, so the agent must commit its changes. Runtime does not auto-commit or execute legacy `pre_artifacts.sh` scripts. The shared downloader transfers the patch; SWE-Marathon directory declarations use the same downloader.

The Harbor adapter maps only official Harbor fields. SWE-Marathon uses `load_harbor_task()` directly; its custom `verifier.type` and `grader.restore_paths` fields are unsupported. Their raw declarations remain in task metadata, and the runtime emits unmapped-field warnings only for selected tasks. AgentCompass does not validate or act on these extensions: they neither select a verifier implementation nor restore files. Normal SWE-Marathon evaluation still executes `tests/test.sh` in the reused agent environment. Checkpoint selection and evaluation resumption are separate runtime responsibilities, not task-format mapping.

`runtime.checkpoint` saves classified inference and prepared-input snapshots, network policy and artifact integrity for none/fresh evaluation recovery. Failed scoring cannot mutate the saved input.

The v4 checkpoint restores `PreparedTask` and the inference result, including issues, trajectory and telemetry. Each scoring attempt receives a separate copy. Driver input files are checked against their saved paths and digests; missing or changed inputs prevent recovery. Client/session objects and unverified remote media URLs cannot form a recoverable snapshot. Legacy artifact-only checkpoints remain readable, but lack classified snapshots for cross-run materialization and may require re-execution. `prepare_evaluation(task, req, plan)` remains the fallback for supported legacy fresh checkpoints. See [Evaluation checkpoints](/en/user_guide/other_features/results/run_records#evaluation-checkpoints).

## Build a Typed Plan for Each Attempt

When evaluator state must be derived from both configuration and the task, define a `BenchmarkPlan` subclass and resolve it once for the current attempt in `build_plan()`:

```python theme={"system"}
from dataclasses import dataclass

from agentcompass.runtime import BenchmarkPlan, EnvironmentSpec, RunRequest, TaskSpec


@dataclass(slots=True)
class ExampleExactMatchPlan(BenchmarkPlan):
    expected: str = ""
    case_sensitive: bool = False


def build_plan(
    self,
    task: TaskSpec,
    req: RunRequest,
    environment: EnvironmentSpec,
) -> ExampleExactMatchPlan:
    _ = environment
    config = self.build_config(req)
    if not isinstance(config, ExampleExactMatchConfig):
        raise TypeError("example_exact_match requires ExampleExactMatchConfig")
    return ExampleExactMatchPlan(
        expected=str(task.ground_truth),
        case_sensitive=config.case_sensitive,
    )
```

`build_plan()` must not open an Environment, call a Model, or mutate `RunRequest`. After the initial `ExecutionPlan` is built, Recipes adjust the plan according to their own contracts. Benchmark documentation therefore must not assume that the runtime enforces one framework-wide precedence rule for Recipe fields. When provider mapping is required, document and test its preservation rules in the corresponding [Recipe Integration](/en/developer_guide/extensions/recipe_integration).

## Prepare Execution Input

`prepare_task()` may create a workspace or upload public material in the task Environment, but its return value may contain only data visible during execution:

```python theme={"system"}
async def prepare_task(
    self,
    task: TaskSpec,
    env: EnvironmentSession,
    req: RunRequest,
    plan: BenchmarkPlan,
) -> PreparedTask:
    _ = env, req
    self._require_plan(plan)
    return PreparedTask(
        task_id=task.task_id,
        category=task.category,
        ground_truth=None,
        input=TaskInput(prompt=task.question),
        output=TaskOutput(answer="Return only the city name."),
        metadata={"dataset_revision": task.metadata["dataset_revision"]},
    )
```

Use the supplied `EnvironmentSession` when you need to create files or directories; do not bypass the Environment and call a provider SDK directly. Retries may invoke this method again, so preparation must be safe to repeat. Otherwise, explicitly clean up the workspace you created before execution.

## Registration and Dependencies

Register the implementation with `@BENCHMARKS.register()` and import its module from `src/agentcompass/benchmarks/__init__.py`:

```python theme={"system"}
from .example_exact_match import ExampleExactMatchBenchmark
```

From the repository root, inspect component discovery and the parameter schema:

```bash theme={"system"}
uv run agentcompass list benchmark
uv run agentcompass config docs benchmark example_exact_match
```

Dependencies required by the framework in every installation belong in the default project dependencies. A Python driver used only by this Benchmark belongs in a dedicated optional dependency group with a declared `DependencySpec`. Runtime dependencies needed by the task or verifier belong in the corresponding Environment and must be pinned there. Successful registration proves only that the module imports; it does not validate data, credentials, the verifier, or a real run.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.