Separate Resources from Scheduling
These four settings solve different problems:
For example, Docker
resources.cpu: 2 limits each container to two cores, while --task-concurrency 8 permits up to eight tasks to be processed concurrently. Neither setting replaces the other.
Task and verifier Environments are also separate resource allocations. When a Benchmark requires a fresh verifier Environment, AgentCompass normally closes the task Environment before creating it. They overlap only when --keep-environment retains the task Environment.
See Run Controls for the full behavior of concurrency, open rate, and --keep-environment.
Provider Capabilities and Units
AgentCompass exposes one provider-neutral resource model. Provider adapters convert these values into their platform-native fields and units.
Providers support subsets of that model:
AgentCompass reports an unsupported requested field before opening an Environment.
storage_mb is the exception: a provider that cannot enforce it ignores it with a warning, so task-declared disk requirements do not prevent otherwise compatible runs.
GPU model constraints remain fail-closed by default. If a Benchmark requires a specific model but you intentionally want the selected provider to allocate any available GPU, set ignore_gpu_type: true explicitly:
gpu_type; it does not remove or change gpu. AgentCompass logs a warning when it applies the directive. The directive follows the same precedence and phase scoping as other resource fields, so it can also be placed in run_resources or evaluation_resources. A higher-precedence gpu_type restores strict model selection for that phase. Do not use this override when the task depends on a particular GPU architecture, memory capacity, or performance profile, and record it when comparing results.
Use the following command to inspect the exact fields and defaults accepted by the installed revision:
Configure Resources
Scope Resources for Fresh Evaluation
Benchmarks withevaluation_environment_mode="fresh" create separate run and evaluation Environments. Use these fields inside --env-params to control their resources:
Phase-specific fields take precedence over the same fields in
resources. For example, this configuration gives both Environments 8192 MiB of memory, while assigning 4 CPUs to the run Environment and 2 CPUs to the evaluation Environment:
resources value overrides the corresponding task fields for both fresh Environments; run_resources and evaluation_resources then override the corresponding common fields. When a fresh Benchmark does not define separate evaluation resources, the evaluation Environment falls back to its common task resources.
The following examples use the same unified resource shape with Docker, Daytona, and Modal. Each one selects a single task through sample_ids and assigns 2 CPU cores and 6144 MiB of memory to each Environment. These values demonstrate the syntax; they are not a recommended Benchmark configuration.
The examples use agentcompass run. See Configure an Environment for configuration-file, Python SDK, and launch orchestration-file forms.
Docker
Daytona
Modal
Docker applies
storage_mb through --storage-opt size=.... If the Docker daemon reports that its storage driver cannot enforce this option, AgentCompass logs a warning and starts the container without a storage limit. A remote provider may also reject a request because of account quota, regional capacity, or an unavailable instance shape.Recipe Resources and Explicit Overrides
Some recipes read task resource requirements from a Benchmark and place them in the unified resource model. The selected provider adapter then performs platform-specific validation and conversion. Built-in recipes usually preserve compatible explicit resource values, but the exact adaptation still depends on the recipe and Benchmark. As a result:- to reproduce the Benchmark resource conditions, start with the defaults supplied by its recipe; and
- to compare another resource profile, override it explicitly and record the change with the results.
Estimate Aggregate Capacity
Estimate capacity in this order:- Start with the resource requirements supplied by the Benchmark or recipe.
- Run one representative task and observe peak memory, CPU use, disk growth, and verifier needs.
- Leave headroom for dependency installation, compilation, and caches.
- Estimate aggregate use from per-instance resources and actual concurrency, then adjust task and provider limits.
- Increase concurrency gradually, reducing it when OOM failures, creation errors, or sustained queueing appear.
--env-open-qps changes only how quickly new instances begin creation; it does not limit the number of running instances.
