task.toml, an instruction, a task-specific image with its workspace at /app, and an
official verifier under tests/.
AgentCompass pins the upstream task set to commit
422b9bb9
and supports the docker, daytona, and modal environment providers. The Harbor adapter loads task resources, and
the provider recipe selects the image automatically. Verification runs in the same sandbox after the agent, which
matches the Harbor task contract and preserves all workspace changes made during the rollout.
Execution contract
- AgentCompass reads the 17 task metadata files from a managed sparse checkout under
data/frontier_swe/. - It downloads
tests/only for the selectedsample_ids, starts the task’s published GHCR image, and exposes/appto the harness. - The harness edits the existing workspace for the task’s
agent.timeout_secbudget. - AgentCompass uploads the official verifier to
/tests, runs/tests/test.shwith the task’sverifier.timeout_sec, and reads/logs/verifier/reward.jsonor/logs/verifier/reward.txtfor scoring. - The raw reward is converted to the official Frontier-SWE gated score. This conversion matters because performance
tasks combine correctness and speedup,
frogsgame-rlreports a board count, andnotebook-compressionreports a lower-is-better compression ratio.
scripts/score_from_reward.py result. The Frontier-SWE scoring guide describes
a separate post-hoc anti-cheat audit that can zero a leaderboard trial; that unpublished audit is not part of the
Harbor task verifier and is therefore not run by AgentCompass.
Resources and network
Task defaults range from 4 to 16 CPUs, 8 to 128 GiB of memory, and 10 to 150 GiB of storage. Five tasks require one H100 or B200 GPU. AgentCompass maps these Harbor fields into its unified resource model, so explicit CLIresources and run_resources values override task defaults field by field. Frontier-SWE verifies in the run
environment, so it does not use separate evaluation_resources.
Docker applies CPU, memory, GPU-count, and best-effort storage limits but cannot select a GPU model. To run a GPU task
on an appropriate Docker host without enforcing its declared H100 or B200 type, set
resources.ignore_gpu_type=true. Daytona maps all five unified resource fields but rejects GPU models unsupported by
the installed Daytona SDK or target. Modal maps CPU, memory, GPU count, and GPU type; it ignores storage_mb with a
warning, so ensure the selected backend has enough free storage.
The legacy Harbor environment.allow_internet value is applied to environment startup, rollout, and verification.
Most tasks use no-network; frogsgame-rl and pcqm4mv2-autoresearch use public network access. A local
mini_swe_agent keeps model calls on the AgentCompass host. When a Harness calls the model inside the sandbox,
AgentCompass preserves the restricted policy but automatically permits the explicitly configured model endpoint. Set
--model-base-url; planning fails before sandbox creation if the endpoint cannot be resolved.
frogsgame-rl also requires TINKER_API_KEY during the agent rollout and verifier. Export it before selecting that
task, expose it through the selected provider’s env_variables setting, and keep it available to the AgentCompass process. AgentCompass resolves the verifier’s Harbor
${TINKER_API_KEY} declaration and fails with a missing-variable message if the controller cannot supply it.
Parameters
Pass benchmark-owned values through--benchmark-params '{...}'.
Use the unified
run_timeout_multiplier and evaluation_timeout_multiplier fields in --execution-params to adjust
the task’s agent and verifier timeouts. Explicit environment parameters override recipe defaults. Frontier-SWE tasks
are intentionally large and long-lived; check provider quotas before running the complete set.
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. These examples use frontier_swe, mini_swe_agent, and the tested Model named by MODEL_NAME. This Harness calls the model from the host and executes commands in the task sandbox. Set the model connection details first:
frogsgame-rl; export TINKER_API_KEY on the host before running the recommended configuration, which also passes it to the sandbox.
- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Run the CPU task
pyright-type-checking-optimization to check image startup, agent execution, and the official verifier. The Docker Recipe applies the task defaults of 8 CPUs, 32 GiB of memory, and the /app workspace.Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
Frontier-SWE uses scalar primary metricscore and scalar auxiliary metric correctness. Score conversion follows the execution contract above; the verifier’s raw reward is not interchangeable with the final score.
With the default configuration, each scalar is averaged over valid observations in the selected tasks, with breakdowns for
implementation, performance, and ml_research. The scalar primary metric does not support the pass execution strategy.
See Metrics and Aggregation for repeated attempts, category aggregation, and shared scoring failure rules.
Task Results and Scoring Evidence
Within each attempt’smeta.benchmark, eval_raw_data contains:
The run-level report’s
extra also records dataset_revision and scoring to identify the task version and scoring method.