TaskSpec values first, then decide which fields the Harness may see and which state must remain evaluation-only.
Distinguish the Four Data Carriers
A Harness can read the entire
PreparedTask, including its ground_truth and metadata. Do not copy TaskSpec.metadata into it without filtering.
Keep evaluator-only data in TaskSpec.ground_truth or a typed BenchmarkPlan, and set PreparedTask.ground_truth to None. These objects still belong to the runtime and result-audit boundary, so they must not contain credentials or other secret material that must never be persisted.
Write a reference answer to RunResult.ground_truth only when it is safe to publish with the result.
Define Public Configuration
Define Benchmark parameters withRuntimeBenchmarkConfig and config_field(), and normalize types early in __post_init__():
Load Deterministic Tasks
load_tasks() should pin the upstream revision and produce stable task_id values. The following metadata contains only reproducibility information that may appear in logs and execution input; the answer remains separate in ground_truth:
select_tasks() method already applies the runtime’s shared task-selection logic. Override it only when the Benchmark needs semantics beyond ordinary ID filtering. Whichever rule you use, keep the returned order deterministic.
Harbor tasks and execution requirements
Useload_harbor_task() from agentcompass.benchmarks.utils.harbor to convert a Harbor task directory into a TaskSpec. In addition to network policies and resources, the adapter maps these execution requirements:
An explicitly separate verifier environment supplies an independent image, workdir, and build-timeout baseline through
evaluation_environment_setup: omitted values do not inherit the agent’s settings. An explicitly empty baseline also remains independent. Only None (no separate baseline declared) inherits the resolved run environment setup, including Recipe defaults. Request-level EnvironmentSpec.setup and evaluation_setup override task defaults. Provider-native startup selectors are not public inputs. Working directory, startup timeout, environment variables, and outbound rules accept only the shared framework fields; provider-native aliases are rejected.
Registry images have one resolved source: EnvironmentSpec.setup.image. Harbor’s environment.docker_image maps to the task setup; Recipes read the merged plan setup and fill a missing image there instead of reading or writing params["image"]. Providers translate this field into their native configuration only when building the environment config. Request parameters image and evaluation_image are rejected; configure setup.image and evaluation_setup.image directly. Use --env-params '{"setup": {"image": "registry.example/runner:v1"}}'.
Fresh-verifier networking follows the same independence rule: an explicit verifier.environment supplies its own baseline, including the default public network when the table is empty. An absent verifier environment inherits the agent baseline. Request-level evaluation_baseline_network_policy overrides the common request baseline_network_policy, which overrides the task’s verifier baseline. The runtime opens the verifier with that baseline, switches to evaluation_network_policy for evaluation, and resets the baseline afterward. Neither policy is silently substituted for the other. You can override only the fresh-verifier baseline with --env-params '{"evaluation_baseline_network_policy": {"network_mode": "no-network"}}'; this request field requires fresh evaluation.
workdir must be an absolute POSIX path inside the environment. It applies when env.exec() omits cwd; an explicit cwd takes precedence. It does not replace an explicit Benchmark/Harness workspace. SWE-Marathon uses the unified workdir for its workspace, falling back to the Dockerfile’s WORKDIR and then its configured workspace root. You can override the task default with --env-params '{"setup": {"workdir": "/app"}}'.
These are separate phase deadlines, not one combined task wall clock, a per-tool timeout, or a sandbox lifetime. None leaves a task deadline unspecified; positive finite numbers specify seconds. The adapter maps only explicitly declared timeouts and does not introduce Harbor’s implicit defaults. A build timeout does not add Dockerfile-building support to a provider that only launches prebuilt images.
BenchmarkPlan.run_timeout_seconds and evaluation_timeout_seconds supply task/Benchmark defaults. The Planner resolves the request’s execution.run_timeout_seconds and evaluation_timeout_seconds, then applies the phase-specific multiplier (run_timeout_multiplier or evaluation_timeout_multiplier) or the common timeout_multiplier. A phase multiplier replaces the common multiplier. An omitted or null override inherits the default. Harness fallback budgets apply only when no task/Benchmark run default exists.
The final ExecutionPlan supplies both the runtime phase watchdog and native Harness/evaluator timeout controls. Harness execute_task() (or a Harness-free Benchmark’s run_task()) and evaluate() have independent clocks, with the same behavior in reused and fresh environments. collect_result() then recovers saved Harness output before session cleanup, under harness_result_timeout_seconds (default 60 seconds). Recovery cannot resume agent execution and does not clear the original execution error. Run/evaluation multipliers do not scale its budget. On timeout the runtime records the phase error and attempts the applicable artifact/cleanup path; cancellation alone does not confirm a remote process has stopped.
Artifact preparation has a separate artifact_collect_timeout_seconds deadline (default null); downloading/restoring artifacts uses artifact_limits.timeout_seconds (default 600 seconds). Configure these under execution in YAML, execution_params in the SDK, or --execution-params in run. These budgets do not extend sandbox TTL or override whole-run cancellation.
Unmapped declarations are stored as TaskSpec.load_warnings. The runtime emits them only after select_tasks(), so excluded samples do not produce warnings. Direct loader callers can inspect this list themselves. Truly unsupported declarations remain visible; do not mark a field mapped merely because it is retained in raw metadata.
Legacy memory and storage size strings under environment or verifier.environment count as mapped only when Harbor converts them into the corresponding MiB resource fields. Conflicting old/new values still fail validation; ignored legacy values still produce unmapped warnings.
Task provenance is preserved in TaskSpec.metadata: the top-level source maps to metadata["source"], while package information, including task.version, maps to metadata["task"]. Explicit task declarations override same-named entries in the free-form metadata table. Package versions are non-empty strings, not necessarily semantic versions; older Harbor releases that omit this field use the adapter’s compatibility validation. The top-level legacy version remains an alias for schema_version, not the package version. Original declarations remain unchanged in harbor_raw.
After Recipe application, the final ExecutionPlan is the only execution-time source. Harnesses and verifiers do not re-read task TOML, private timeout fields, or timeout metadata. Recipes must change plan.run_timeout_seconds and plan.evaluation_timeout_seconds; the Planner maps the run value to the native harness field.
Task environment variables
The shared task and environment contracts have common, run-only, and evaluation-only bindings:env_variables, run_env_variables, and evaluation_env_variables. TaskSpec.evaluation_environment_env_variables supplies a separate fresh-verifier startup baseline: None inherits task common bindings, while an explicit map replaces them, including an empty map. Harbor maps the effective separate verifier environment’s env table to this field, not to command-only bindings. Plans retain references such as ${GRADER_KEY}; providers resolve them from the launcher’s environment before execution. Export only the keys required by your selected tasks. Missing required references fail before allocation, without printing their values. ${NAME:-default} and references embedded in a string are supported; shell commands are never expanded.
Request-level values override task defaults per key. Within the request, phase-specific bindings override common bindings; the active phase overrides internal env.exec(env=...) defaults. Native protocol requirements that cannot be overridden must be checked with env.require_exec_env(); conflicting names fail explicitly without printing values. Agent setup/run and evaluator execution use session-local, async-context-local scopes. The scope resets on success, error, or cancellation. Task preparation, result recovery, Harness cleanup, artifact commands, and artifact restore use common bindings only; agent bindings do not carry over to them. Common bindings can also reach sandbox startup where supported; phase-only bindings are applied to commands, not the image entrypoint. This is an injection contract, not a security boundary against an agent that persists processes or files in a reused sandbox. host_process continues to inherit the host process environment.
EnvironmentSpec.evaluation_environment_env_variables is the request override for fresh startup bindings. It requires evaluation_environment_mode="fresh" and merges per key after the task verifier baseline and request common bindings. Evaluation-only request bindings take precedence during evaluator commands. An omitted request map inherits; an empty request map adds no overrides. It does not replace the task baseline, unlike an explicit task map.
The former Harness env parameter is rejected; migrate it to environment.run_env_variables. Harnesses obtain resolved values from the scoped Environment when adapting nested native tool configurations. They must not mutate RunRequest or the host os.environ. Local SDK calls need explicit SDK configuration; only commands executed through the Environment receive these bindings.
Harbor’s solution.env belongs to the reference solution by default. A Benchmark must explicitly opt in with solution_env_for_agent=True when those bindings really belong to its agent tasks. Do not inject reference-solution secrets automatically.
Submission artifacts and replay
TaskSpec.artifacts is the resolved file selection supplied by the Benchmark or format adapter. A Benchmark with collect_artifacts = true always includes the conventional main /logs/artifacts/ directory. The Harbor adapter resolves that same directory while loading its task format, including for an empty Harbor declaration, and appends explicit extra paths; native Benchmarks receive the same treatment during planning. TaskSpec.artifacts = [] or execution.artifacts = [] means no extra paths, not an opt-out. Only collect_artifacts = false disables collection. An explicit convention-directory entry supplies its own exclusions. Destination collisions keep the earlier entry with a warning, as in Harbor; the shared runtime receives a non-overlapping list.
Benchmark declarations provide the collection defaults. execution.artifacts replaces task-specific extra paths when supplied, but /logs/artifacts/ remains for a collecting Benchmark; omitted or null inherits, and [] selects only that conventional directory. execution.artifact_collect independently replaces commands, where [] clears the command list. Runtime uses the resolved lists; collect_artifacts = false is the only way to skip collection. collect_artifacts is a read-only resolved-plan value, not a CLI, YAML, or SDK execution option. execution.save_artifacts controls persistent local storage: omitted or null resolves to true for fresh and false for reuse/none. Fresh requires storage; explicitly setting save_artifacts=false fails before environment allocation. In reuse/none, disabling storage skips file download but still runs declared preparation commands. Enabling storage alone does not select any additional paths. Storage does not preserve the sandbox (keep_environment controls that). Answer/trajectory/log persistence and Benchmark-specific evaluator outputs are unaffected. Evaluation-only recovery supports none/fresh when a classified inference snapshot and complete scoring inputs are available. Reuse requires a new attempt.
When an artifact path depends on the final workspace selected by a Recipe, override BaseBenchmark.resolve_artifacts(task, req, plan). The Planner calls it after all Recipes and before final normalization; an explicit execution.artifacts list remains authoritative and bypasses Benchmark defaults.
Override artifact paths and preparation commands
execution.artifacts and execution.artifact_collect independently override task/Recipe extras and commands. Omitted or null inherits the corresponding declaration; a supplied artifact list replaces the extra paths but retains /logs/artifacts, including [], which selects only that directory. CLI lists replace corresponding YAML lists. These settings can also supply paths or commands for a Benchmark with no task-specific defaults, provided its collection capability is enabled. A Benchmark class can declare collect_artifacts = False; either explicit override (including []) then fails during planning with an unsupported error. Omitted/null values are not overrides. This capability defaults to True and is not a Benchmark params or execution CLI option. Recipes cannot re-enable it. To retain task-specific entries while adding your own, include the required original entries in the replacement list.
source is an absolute path in the agent sandbox and may name a file or directory. destination is relative to the attempt’s local artifacts/ directory: /app/submission above is stored under artifacts/submission/, excluding *.tmp; the separate file is stored as artifacts/result.json. Missing sources are recorded without failing collection. Destination overlaps are rejected before execution. Fresh evaluation restores files to their original source paths.
Only the resolved command list runs, in order, before any shared file download; overriding commands does not automatically prepend the Benchmark commands. CLI commands use timeout_seconds; Harbor task.toml uses timeout_sec. Shared preparation and transfer limits apply to the resolved lists. save_artifacts=false still runs commands in reuse/none; fresh requires saving. The example command creates a demonstration file and replaces the Benchmark command list. Preserve any commands required by the verifier in your replacement list; otherwise its expected submission may not be produced.
Recipes may adjust the resolved paths. None and [] mean no resolved paths; they are not reinterpreted as Harbor declarations. Paths are revalidated after Recipe adjustment. To disable the entire collection pipeline, use the execution switch. With host_process, absolute sources refer to the host filesystem.
The default directory is optional: a missing source is recorded as missing, and a successfully downloaded empty directory is recorded as collected with zero content bytes. Neither case blocks evaluation or checkpoint creation. An empty directory is preserved for fresh restore; a missing entry does not upload files. Permission errors, broken ancestor paths, transport failures, timeouts, and size violations remain capture errors, not missing outputs. Their existing failure behavior is unchanged. Empty capture alone does not prove an independent verifier can grade the task; replay still needs every input the verifier requires.
ArtifactSpec lives in agentcompass.runtime.artifacts. A string declaration names an absolute sandbox source path. An object can also specify a relative artifact-store destination and exclude patterns. The default destination removes the source’s leading slash. Collection supports files, binary data, and directories; exclusions use GNU tar patterns. Overlapping destinations, escaping paths/links, special files, and non-main service declarations fail explicitly.
The runtime first executes declared collection commands, if any. Disabling collection skips the entire pipeline. Otherwise, no commands means preparation is skipped and existing files can still be collected. With storage enabled, selected outputs are copied before releasing the agent environment. Binary contents stay in files below the run’s artifacts/ directory; RunResult.artifacts.declared_artifacts contains relative paths, status, sizes, and checksums. Keep that directory with the result JSON when moving a run. Result reuse verifies and copies referenced artifact bytes before publishing the reused detail or terminal attempt checkpoint; missing or corrupt payloads cause that result to rerun. Artifact bytes are not hard-linked between runs. Missing or fully excluded sources are recorded, leaving the verifier to score them. Transport errors and limit violations fail collection.
Runtime executes ExecutionPlan.artifact_collect in order inside the agent sandbox. Without commands, preparation is skipped; no Benchmark fallback hook runs. When storage is enabled and paths are selected, shared download_artifacts() handles packaging, transfer, validation, and local storage. An empty path list does not suppress declared commands.
Harbor’s verifier.collect maps to TaskSpec.artifact_collect, which Planner copies and validates again after Recipe overrides. Each command maps command, timeout_sec to timeout_seconds (default 60 seconds), and service (default main). Commands run sequentially with bash -c, under the existing run-phase network policy and common environment bindings, before evaluation or environment teardown. They do not receive verifier-only environment variables; variable references expand in the sandbox, not during task loading. Each command has its own positive finite timeout in addition to the optional total collection-phase deadline. Sidecar services and command-level user overrides are not supported and fail during loading; unknown command fields also fail explicitly.
Commands still execute when only save_artifacts is disabled; paths can be selected without commands. Paths describe existing or expected submissions, not how to generate them. For example:
submission-*.manifest.json. Each successfully validated entry remains on disk if a later operation fails. ArtifactDownloadError.manifest contains the partial record with complete=false, failed/cancelled entries, and uncollected entries. The runtime preserves it in declared_artifacts, records preparation/transfer errors separately in telemetry post_run_errors, and retains any original run error. Preparation failure still permits best-effort collection of existing files, but either preparation or transfer failure prevents evaluation and creation of a resumable checkpoint for that attempt. Explicit cancellation propagates, leaving the journal for diagnosis. Progress events report the current operation, completed bytes and entries, archive size when known, and remaining budget; provider downloads do not expose continuous byte-level progress.
Fresh evaluation validates the manifest and checksums, then restores collected outputs to their original sandbox source paths before calling evaluate(). A restored directory replaces the target’s contents, including removing stale and hidden entries, while preserving the target directory itself for mounted workspaces. Parent paths are restored before their descendants regardless of declaration or manifest order, so explicitly captured children survive parent exclusions. Fresh plans reject conflicting exclusion declarations for the same source. Multiple copies of one source must have matching capture status, content and permissions before any upload; identical copies are restored once. Restoration stages the payload first and attempts rollback on installation failure; interrupted recovery staging is retained until environment cleanup. Reuse evaluation never restores over the live workspace. Before rerunning an agent within the same logical attempt, runtime archives the previous artifacts, result and checkpoint under retries/<id>/ and updates diagnostic artifact references. The new physical run starts with an empty artifact destination; evaluation-only resume preserves its captured inputs. Legacy inline text submissions remain supported when they cover all declared outputs. ExecutionPlan.artifact_limits defaults to 16 GiB, 100,000 filesystem entries, and 600 seconds per collection/restore operation; Recipes may override these typed limits. Local file workers finish or acknowledge cancellation before temporary-file cleanup.
The pinned DeepSWE v1.1 tasks declare verifier.collect commands that extract the diff from the base commit to HEAD, so the agent must commit its changes. Runtime does not auto-commit or execute legacy pre_artifacts.sh scripts. The shared downloader transfers the patch; SWE-Marathon directory declarations use the same downloader.
The Harbor adapter maps only official Harbor fields. SWE-Marathon uses load_harbor_task() directly; its custom verifier.type and grader.restore_paths fields are unsupported. Their raw declarations remain in task metadata, and the runtime emits unmapped-field warnings only for selected tasks. AgentCompass does not validate or act on these extensions: they neither select a verifier implementation nor restore files. Normal SWE-Marathon evaluation still executes tests/test.sh in the reused agent environment. Checkpoint selection and evaluation resumption are separate runtime responsibilities, not task-format mapping.
runtime.checkpoint saves classified inference and prepared-input snapshots, network policy and artifact integrity for none/fresh evaluation recovery. Failed scoring cannot mutate the saved input.
The v4 checkpoint restores PreparedTask and the inference result, including issues, trajectory and telemetry. Each scoring attempt receives a separate copy. Driver input files are checked against their saved paths and digests; missing or changed inputs prevent recovery. Client/session objects and unverified remote media URLs cannot form a recoverable snapshot. Legacy artifact-only checkpoints remain readable, but lack classified snapshots for cross-run materialization and may require re-execution. prepare_evaluation(task, req, plan) remains the fallback for supported legacy fresh checkpoints. See Evaluation checkpoints.
Build a Typed Plan for Each Attempt
When evaluator state must be derived from both configuration and the task, define aBenchmarkPlan subclass and resolve it once for the current attempt in build_plan():
build_plan() must not open an Environment, call a Model, or mutate RunRequest. After the initial ExecutionPlan is built, Recipes adjust the plan according to their own contracts. Benchmark documentation therefore must not assume that the runtime enforces one framework-wide precedence rule for Recipe fields. When provider mapping is required, document and test its preservation rules in the corresponding Recipe Integration.
Prepare Execution Input
prepare_task() may create a workspace or upload public material in the task Environment, but its return value may contain only data visible during execution:
EnvironmentSession when you need to create files or directories; do not bypass the Environment and call a provider SDK directly. Retries may invoke this method again, so preparation must be safe to repeat. Otherwise, explicitly clean up the workspace you created before execution.
Registration and Dependencies
Register the implementation with@BENCHMARKS.register() and import its module from src/agentcompass/benchmarks/__init__.py:
DependencySpec. Runtime dependencies needed by the task or verifier belong in the corresponding Environment and must be pinned there. Successful registration proves only that the module imports; it does not validate data, credentials, the verifier, or a real run.