Skip to main content
agentcompass run and agentcompass launch use the same run controls for scheduling, fault handling, and evaluation artifacts. execution.task_concurrency is the single concurrency setting for physical attempt executions; there is no separate task-versus-attempt concurrency knob. This page explains what each control does and how to use it. See agentcompass config for configuration-file syntax and precedence, agentcompass run for the complete single-request signatures, and agentcompass launch for multi-request orchestration and its CLI overrides.

Scale Concurrency Safely

A provider is the execution backend that creates and manages environments, such as Docker, Daytona, or Modal. Effective attempt concurrency is bounded by the lower of task_concurrency and the applicable provider limit. env-open-qps controls only Environment startup pacing. With strategy: avg, same-task attempts share this pool and may overlap only when both the Benchmark and Harness declare their state isolated; otherwise those attempts remain serial. Model endpoint capacity, provider quotas, and local CPU and memory can reduce actual concurrency further. CPU and memory limits for one sandbox are Environment parameters; see Understand the Scope.

CLI Syntax

In the CLI, repeat the latter two options for different providers. When evaluation requests use different environments, apply limits to each provider from one command:
A single agentcompass run needs limits only for the providers actually used by that request.

Configuration File Syntax

In a --config file, use mappings for provider limits instead of repeating YAML keys:
The example above is a regular run configuration. In a launch orchestration file, put shared task_concurrency at the top level while keeping the provider mappings under runtime; see agentcompass launch. When tuning concurrency, first select a few representative Benchmark tasks and validate them with task concurrency set to 1, then increase it gradually to 2 or 4. Observe Environment startup latency, model latency, error rates, and memory use. Return to the last stable value when errors increase.

Set an Appropriate Timeout

Use --timeout-seconds to limit the whole run and --execution-params to configure individual task phases. Environment startup, single commands, and model requests can also have their own limits. When limits overlap, the first deadline reached ends the corresponding work.

Overall and component deadlines

LayerParameter locationScope
Evaluation deadlineCLI: —timeout-seconds <seconds>All tasks in one run share this limit, as do all requests in one launch. Timing starts after component preflight and covers task loading, preparation, execution, analysis, and summarization. On expiry, unfinished work is cancelled and resource cleanup begins. The default is 360000 seconds (100 hours). Explicitly setting 0 disables the evaluation deadline; it does not affect the component-specific timeouts below.
Environment creationJSON field: setup.build_timeout_seconds
Passed through --env-params
A shared startup deadline for every Environment, including Daytona and Modal. Each sandbox creation is timed separately. Expiry fails only that creation and does not limit later operations in a successfully created sandbox.
Harness-specificJSON field: defined by the Harness
Passed through --harness-params
The exact scope depends on the field. command_timeout limits one command; request_timeout limits one service request. Use the common execution fields below for a task phase deadline.
Benchmark-specificJSON field: defined by the Benchmark
Passed through --benchmark-params
The exact scope depends on the field. For example, PinchBench judge_timeout_seconds limits one judge-model request.

Per-task phase deadlines

Configure phase timeouts through run --execution-params, YAML execution, or SDK execution_params. For launch, use defaults.execution or requests[].execution. The diagram shows phase budgets for one attempt with a Harness and a fresh evaluation environment, rather than the overall task lifecycle deadline. Phases run from left to right; their widths do not represent duration. Some phases may be skipped; downloads and restores depend on the evaluation mode and artifact settings. Timeout coverage across a single task attempt: execution, collecting outputs and building RunResult, nested artifact preparation limits, independently timed batch download and restore, and evaluation. Gray columns mark phases outside the six timeout settings shown: environment initialization, closing the Harness, evaluation environment initialization, and cleanup. Other Environment or lower-level timeouts may still apply. Time budgets are in seconds. The execution budget covers preparation inside execute_task(), model requests, and tool execution. Environment/session setup before the call and result collection afterward have separate timing. Evaluation starts its clock after the evaluation environment is ready and artifacts have been restored. For execution and evaluation, an omitted or null override inherits the Benchmark plan default, then the task default. Execution also falls back to the Harness default. If none supplies a budget, the phase has no deadline. The final budget is base seconds × effective multiplier. A phase-specific multiplier replaces timeout_multiplier; when it is null, the common multiplier applies. These multipliers do not multiply together. They affect only execution and evaluation; result collection, artifact preparation, download, and restore keep their own budgets. A multiplier cannot create a deadline when no base budget exists.
This allows 7,200 seconds for execution and 600 seconds for evaluation per attempt. On timeout, the runtime records the failure and attempts applicable result collection and cleanup. Collecting a partial answer does not clear the timeout or guarantee evaluation. These budgets do not extend sandbox lifetime or the outer --timeout-seconds deadline.

Collect outputs and build RunResult

harness_result_timeout_seconds limits the Harness collect_result() call after successful execution, an execution error, or an execution timeout. This call reads logs, saved state, and requested output files; extracts the answer; and assembles the trajectory, status, and diagnostics into RunResult for persistence, analysis, and evaluation. Some Harnesses also write trajectory files during collection. Collection has its own budget because it may require sandbox I/O after execution has ended. It does not run the agent or request further model responses. It finishes before the Harness session closes; artifact preparation and download follow while the task sandbox remains available. Harness-free Benchmarks skip this hook. A collection failure is recorded alongside any original execution error.

Prepare and transfer artifacts

Preparation commands run in order under two limits: the whole list shares artifact_collect_timeout_seconds, and each command has its own timeout_seconds. The first deadline reached stops preparation. For example, three commands with 60 seconds each and a 100-second group budget may use at most 60 seconds per command and 100 seconds in total. With a null group budget, only the individual command limits apply. For transfers, one batch download shares artifact_limits.timeout_seconds across all declared files. A subsequent restore starts a new clock with the same budget. At the default, that means up to 600 seconds for download and another 600 seconds for restore.

Configure post-run budgets

This YAML allows 120 seconds for building RunResult, 100 seconds for preparation, and 1,800 seconds per batch transfer. It inherits the declared paths, commands, and per-command limits:
To remove the group preparation deadline while retaining per-command limits, override it with null:
CLI and SDK mappings merge over YAML; omitted fields inherit and supplied lists replace existing lists. Dedicated CLI options take precedence over the corresponding --execution-params fields. The SDK accepts execution_params in build_run_request(), run_evaluation(), and async_run_evaluation().

Choose the budget to adjust

Configure Artifacts

Use execution to control which artifacts are saved, how they are prepared, and whether evaluation restores them. Preparation and transfer deadlines are described under timeouts.

Save and restore artifacts

For Benchmarks with artifact collection enabled and declared paths, transfers depend on the evaluation mode: An omitted or null save_artifacts uses the mode’s default. Skipped transfers consume no transfer budget. Declared preparation commands and Harness result collection still run when downloads are disabled. Saving artifacts keeps local files; use keep_environment to retain the sandbox itself. Artifact size/count limits default to 16 GiB and 100,000 entries per collection or restore. Under artifact_limits, set max_mb and max_files to adjust them. max_mb covers the whole artifact list and also caps each transfer archive, including archive overhead; one MB is 1,048,576 bytes. Preparation or transfer failures preserve validated files and execution diagnostics, but skip grading and creation of a resumable checkpoint for that attempt. Progress records report the operation, completed bytes, and remaining budget.

Override artifact paths and preparation commands

Benchmark declarations supply the default paths and commands. Override them under execution: Include any required original paths or commands in a replacement list. These overrides require a Benchmark that supports artifact collection; see Submission Artifacts. Use absolute sandbox paths for source and relative local paths for destination. Commands run in the Environment’s default workdir; use absolute paths or an explicit cd when they depend on the task workspace.
This example replaces the preparation commands with one that creates a demonstration file. It saves /app/submission under the attempt’s local artifacts/submission/, excluding *.tmp, and /app/result.json as artifacts/result.json. Missing sources are recorded without failing collection. Fresh evaluation restores files to their original source paths.

Retry Only Transient Failures

--max-retries sets the maximum retries inside each logical attempt. For example, --max-retries 2 permits up to two replacement executions after that attempt’s initial failure. A retry never creates a new metric attempt. --retry-pattern-list matches each ERROR issue message and code separately; null and [] mean FATAL-only retries. FATAL always retries within the shared budget and WARNING never retries. Retry only transient errors that may recover on another execution, such as dropped network connections, temporary service failures, or sandbox timeouts:
Judge invalid JSON, missing credentials and Environment failures are FATAL. The default budget is 0; use —max-retries 2 to tolerate transient failures. Unresolved FATAL fails the run and removes official scores. Retries restart only the current logical attempt, or only its evaluation phase when the retry scope permits. Completed sibling attempts stay checkpointed: if attempt 3 retries, attempts 1 and 2 are not rerun. The final details record retry_count and retry_counts; discarded executions remain under the attempt’s retries/ directory for diagnosis.

Output and Reuse

Name a New Run

For agentcompass run, the three options correspond to different levels of the result path:
  • --results-dir sets the result root and defaults to results.
  • --run-name adds an optional experiment-group directory.
  • --run-id names this run’s directory; the current timestamp is used when it is omitted.
Model, Benchmark, and Harness IDs are normalized to safe names and joined with underscores into one directory component. For agentcompass launch, each request uses its normalized name in place of that combined component:
See agentcompass launch for request naming and collision checks. The following command uses ablation as the experiment group and gives this run the fixed name baseline:
With the default result root, the path is:
See Understanding Evaluation Results for the complete directory and file layout.

Resume an Interrupted Run

Use --reuse to continue an evaluation from an existing run. AgentCompass reuses compatible complete details without errors by task ID and can materialize valid terminal-attempt checkpoints for an unfinished multi-attempt task:
Without a value, --reuse selects the latest run under this hierarchy:
Pass a run ID to select an exact source under that hierarchy:
For run, AgentCompass searches only within the same result root, optional run-name prefix, and combined Model/Benchmark/Harness directory. For launch, it searches within the request’s named output namespace. Existing directories remain untouched; automatic reuse does not search the former Benchmark/Model hierarchy. The source must use a supported run schema, the same Benchmark ID, and the same attempt plan. Task IDs and attempt numbers identify reusable data; request parameters and task inputs do not need to be identical. You can change evaluation timeouts, resources, or environment variables when continuing a pending fresh evaluation. The new run records the source and preserves reused details or checkpoints; see reuse validation. Reuse is decided per logical attempt using the current task plan. Complete WARNING results and ERROR results whose message/code do not match the current patterns remain reusable. FATAL or matching ERROR results enter the current retry policy with a new shared budget; historical retry counts do not consume it. Zero budget does not turn a failed result into success. Only evaluation failures with complete none/fresh snapshots can resume scoring; reuse mode reruns the attempt.

Resume Evaluation from Saved Artifacts

After inference and required collection complete without FATAL, none/fresh modes save a v4 checkpoint containing classified inference and prepared-input snapshots, identity, provenance, network policy and artifact integrity. Recoverable ERROR results are allowed. Fresh recovery restores captured artifacts; none recovery verifies declared local input paths and digests. Incomplete transfers or unserializable runtime objects prevent recovery. The new run copies validated checkpoints and captured artifacts. Recovery validates task/attempt identity, source compatibility, network policy and input integrity, then uses the current plan and a fresh copy of the saved inputs. Driver-side file dependencies must still exist and match their digests. Recovery does not restore a live process or sandbox. To disable automatic checkpoint recovery for pending tasks, add --no-checkpoint-resume to agentcompass run, pass checkpoint_resume=False to the Python SDK, or configure:
For multi-request launches, set defaults.runtime.checkpoint_resume or requests[].runtime.checkpoint_resume. The effective default is true. This option does not force already completed results to be evaluated again, disable terminal-attempt scheduling checkpoints, or prevent new checkpoints from being saved.

Keep Environments for Debugging

Add --keep-environment when a failure requires direct inspection of task or verifier sandboxes:
AgentCompass then skips provider cleanup for environments created by the run. Retries and multiple tasks may leave several resources active, so release them later with the provider’s tooling. Harness sessions are still closed normally.

Logs and Progress

--progress controls only terminal rendering. AgentCompass still saves progress, logs, and task results in every mode. See Results for their locations.

Environment variable scopes

Configure all variable scopes through --env-params (SDK: environment_params; YAML: the selected provider under environments). env_variables supplies common startup/command bindings; run_env_variables covers agent setup/run; evaluation_env_variables covers evaluator commands. evaluation_environment_env_variables independently overrides fresh-verifier startup bindings and requires fresh mode. The former Harness env input is rejected.
The earlier execution.download_artifacts option has been replaced by execution.save_artifacts; the old configuration key is rejected.

Error handling and score validity

RunResult.issues is the only execution error contract. Each issue has severity (fatal, error, warning), phase (setup, run, collect, evaluate, cleanup), a stable code, and a redacted message. RunResult.error has been removed. Full redacted tracebacks and exception chains belong in artifacts.execution_diagnostics. FATAL reports setup, Environment, external tool service, model authentication/quota/server, or Judge failures. Judge timeouts, invalid JSON and missing required scoring fields are FATAL. Model API or Harness timeouts and empty model responses are ERROR. Missing model deliverables, ordinary verifier reward files that are absent or malformed, and cleanup failures are WARNING. A known trusted setup failure remains FATAL even when no reward is produced. ERROR does not override a valid Benchmark score. Ordinary verifier reward failures remain valid fail/0 observations; the Judge protocol exception above is FATAL. Runtime never fills missing observations with zero based on status alone. execution.max_retries defaults to 0, and is one shared extra retry budget per logical attempt. FATAL retries whenever budget remains. ERROR retries only when a configured regex matches its message or code; WARNING never triggers a retry. Both null and [] mean FATAL-only retries. Consider execution.max_retries=2 for transient service failures. A zero budget means one unresolved FATAL fails the run. Patterns no longer match traceback text, multiple concatenated issues, or redacted secrets; migrate them to stable codes or retained diagnostic features. The resolved task plan supplies both budget and patterns. In none and fresh modes an evaluation retry uses an isolated inference snapshot; fresh creates a new evaluation Environment. In reuse mode it reruns the whole attempt. Cross-run recovery uses the new run’s budget; previous retry counts remain history. Resolved failures remain in retry history and are absent from final issues. Any logical attempt with a final FATAL invalidates its entire task across every metric, even when another attempt succeeded. MetricReport.evaluation_failed then becomes true, all official value fields are null, and available values move to explicitly labelled reference_value fields. Reference scores exclude invalidated tasks from numerators, denominators and weights. With no valid tasks both values are null. The run is failed, CLI exits nonzero, and SDK failure messages retain result paths. Other requests in an orchestration continue. Counts satisfy evaluated + unavailable + invalidated = total. error is a separate diagnostic count. Attempt coverage counts each final ERROR/FATAL attempt once, regardless of how many issues it has. WARNING and resolved retry history do not count. Legitimate pass@k early stopping is not a missing attempt. Current task and run-info schemas are v3; evaluation checkpoints are v4 and contain classified inference and prepared-input snapshots, network policy and artifact integrity data. Previous v2 results are converted only at the read boundary using reliable structured evidence. Historical free-text failures remain classification-unknown and require re-execution for current scoring; they remain viewable as historical results. Old checkpoints without classified snapshots are rejected for cross-run materialization with a reason and normal execution resumes. New-format missing issues or result-level error is a format error.