Skip to main content
Environment resource parameters limit what a local instance can use or specify the CPU, memory, storage, and GPU requested for a remote instance. Use them to keep one task from consuming excessive resources or to request a remote instance that matches Benchmark requirements. They do not limit the AgentCompass host process, model service, or other external services.
host_process runs commands directly on the host and cannot enforce Environment-level CPU, memory, storage, or GPU limits. Choose another provider when you need resource isolation.

Separate Resources from Scheduling

These four settings solve different problems: For example, Docker resources.cpu: 2 limits each container to two cores, while --task-concurrency 8 permits up to eight tasks to be processed concurrently. Neither setting replaces the other. Task and verifier Environments are also separate resource allocations. When a Benchmark requires a fresh verifier Environment, AgentCompass normally closes the task Environment before creating it. They overlap only when --keep-environment retains the task Environment. See Run Controls for the full behavior of concurrency, open rate, and --keep-environment.

Provider Capabilities and Units

AgentCompass exposes one provider-neutral resource model. Provider adapters convert these values into their platform-native fields and units. Providers support subsets of that model: AgentCompass reports an unsupported requested field before opening an Environment. storage_mb is the exception: a provider that cannot enforce it ignores it with a warning, so task-declared disk requirements do not prevent otherwise compatible runs. GPU model constraints remain fail-closed by default. If a Benchmark requires a specific model but you intentionally want the selected provider to allocate any available GPU, set ignore_gpu_type: true explicitly:
This removes only gpu_type; it does not remove or change gpu. AgentCompass logs a warning when it applies the directive. The directive follows the same precedence and phase scoping as other resource fields, so it can also be placed in run_resources or evaluation_resources. A higher-precedence gpu_type restores strict model selection for that phase. Do not use this override when the task depends on a particular GPU architecture, memory capacity, or performance profile, and record it when comparing results. Use the following command to inspect the exact fields and defaults accepted by the installed revision:
Each provider page explains its field formats, account quotas, and operating requirements in more detail.

Configure Resources

Scope Resources for Fresh Evaluation

Benchmarks with evaluation_environment_mode="fresh" create separate run and evaluation Environments. Use these fields inside --env-params to control their resources: Phase-specific fields take precedence over the same fields in resources. For example, this configuration gives both Environments 8192 MiB of memory, while assigning 4 CPUs to the run Environment and 2 CPUs to the evaluation Environment:
Benchmark task metadata remains the lowest-precedence resource source. A common explicit resources value overrides the corresponding task fields for both fresh Environments; run_resources and evaluation_resources then override the corresponding common fields. When a fresh Benchmark does not define separate evaluation resources, the evaluation Environment falls back to its common task resources. The following examples use the same unified resource shape with Docker, Daytona, and Modal. Each one selects a single task through sample_ids and assigns 2 CPU cores and 6144 MiB of memory to each Environment. These values demonstrate the syntax; they are not a recommended Benchmark configuration. The examples use agentcompass run. See Configure an Environment for configuration-file, Python SDK, and launch orchestration-file forms.

Docker

The Docker adapter translates these values into Docker CPU and memory arguments. See Docker resource parameters for provider requirements.

Daytona

The recipe for this combination selects the task image, so the request applies to an image-based sandbox. See Daytona resource parameters for provider requirements.
The Modal adapter passes the MiB memory value to Modal. See Modal resource parameters for provider requirements.
Docker applies storage_mb through --storage-opt size=.... If the Docker daemon reports that its storage driver cannot enforce this option, AgentCompass logs a warning and starts the container without a storage limit. A remote provider may also reject a request because of account quota, regional capacity, or an unavailable instance shape.

Recipe Resources and Explicit Overrides

Some recipes read task resource requirements from a Benchmark and place them in the unified resource model. The selected provider adapter then performs platform-specific validation and conversion. Built-in recipes usually preserve compatible explicit resource values, but the exact adaptation still depends on the recipe and Benchmark. As a result:
  • to reproduce the Benchmark resource conditions, start with the defaults supplied by its recipe; and
  • to compare another resource profile, override it explicitly and record the change with the results.
Resource changes can affect task completion and scores. Do not combine runs made under different resource limits as though they used the same evaluation conditions.

Estimate Aggregate Capacity

Estimate capacity in this order:
  1. Start with the resource requirements supplied by the Benchmark or recipe.
  2. Run one representative task and observe peak memory, CPU use, disk growth, and verifier needs.
  3. Leave headroom for dependency installation, compilation, and caches.
  4. Estimate aggregate use from per-instance resources and actual concurrency, then adjust task and provider limits.
  5. Increase concurrency gradually, reducing it when OOM failures, creation errors, or sustained queueing appear.
Capacity planning should focus on how many Environment instances can exist at the same time. --env-open-qps changes only how quickly new instances begin creation; it does not limit the number of running instances.

Troubleshoot Resource Problems