openevolve harness and
the Docker recipe. It pins the upstream Frontier-Engineering source repository to a reproducible revision, prepares
each task’s initial program and verifier materials, and returns the best candidate together with the verifier’s metrics.
Frontier Engineering does not use an LLM judge or pairwise output comparison.
How It Works
A Frontier Engineering run has four stages:- Select tasks. AgentCompass loads the packaged task matrix, applies
task_set, and then applies any exactsample_idsfilter. - Prepare the baseline. The pinned upstream repository is cached under the AgentCompass data directory. For every selected task, AgentCompass resolves the task metadata, uploads the benchmark materials, and places the shipped initial program in the task workspace.
- Evolve the program. The
openevolveharness asks the model under test to propose program edits. Each candidate is evaluated by the task’s official command, and the evaluator’s feedback is available to later generations. The harness submits the best program it finds. - Verify and aggregate. AgentCompass evaluates the submitted program, records the task score and artifacts, and aggregates scores across tasks. Read-only benchmark files are checked so that a candidate cannot change the verifier or its reference data.
Task families
The released matrix spans engineering and scientific optimization families including computer systems, cryptography, GPU kernels, quantum computing, job-shop and inventory optimization, robotics, optics, energy storage, structural optimization, astrodynamics, sustainable data-center control, and EngDesign. The exact task ids come from the selected matrix and can be inspected in this directory:Verifier scoring
Each task’s official evaluator writes numeric metrics such ascombined_score, score, or raw_score. AgentCompass
uses the evaluator’s combined score when available and exposes it as the scalar Metric Contract observation
metrics.score; it never substitutes a judge-model opinion. When available, the evaluator’s valid value remains
diagnostic metadata under meta.benchmark.frontier_engineering.evaluation.valid, not a final correct flag or Metric Contract observation.
Only authoritative score fields produce metrics.score; telemetry such as runtime_s cannot imply a score. Missing or invalid ordinary verifier output records WARNING and no observation, without generic zero fill. Known setup/Environment failures are FATAL; execution failures without a confirmed cause are ERROR.
See Scoring Metrics for medal scores, ranks, and coverage.
Parameters
Pass benchmark-owned configuration through--benchmark-params '{...}'. The table below intentionally contains only
Frontier Engineering task-selection fields; harness evolution settings and provider settings are documented with the
selected harness and environment.
Select a task matrix with
task_set. The counts below are the entries in the matrices shipped with this AgentCompass revision:
sample_ids is matched against the matrix labels, for example
InventoryOptimization/disruption_eoqd or Optics/holographic_multiplane_focusing. Use agentcompass list benchmark
and the benchmark data files to inspect the registry and available ids.
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. These examples use frontier_engineering, openevolve, and $MODEL_NAME, with docker as the Environment.
Set the following environment variables in your terminal before running the examples:
- Model under test:
MODEL_NAME,MODEL_BASE_URL, andMODEL_API_KEY; see Model connection details.
host_process, install the frontier-engineering extra first.
- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Run one representative task for one evolution iteration. This checks task preparation, model access, candidate
collection, and official verification end to end.
Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
Frontier Engineering’s scalar primary metricscore is displayed as “Raw Score”; see Verifier scoring above for its task-specific meaning. Raw scores can use different units. Their default overall mean is not a normalized percentage.
Each task contributes
1, 0.67, or 0.33 when its score reaches the gold, silver, or bronze threshold, respectively, and 0 below the bronze threshold. The medal score sums these credits and divides by the corresponding baseline’s fixed total task count.
Unselected baseline tasks keep a zero contribution. A selected applicable task without a valid observation makes the corresponding medal score unavailable. Rank comparisons use only tasks that overlap the reference table and pass score filters. Report-level extra entries frontier_engineering_rank and frontier_engineering_medal preserve coverage, reasons, and official, reference, or unavailable labels.
The primary metric is scalar, so the pass execution strategy is unsupported.
See Metrics and Aggregation for repeated attempts, category aggregation, and shared scoring failure rules.
Task Results and Scoring Evidence
The main entries in each attempt’sartifacts are:
Under
meta.benchmark, frontier_engineering records the evaluation command, workspace, source task information, and evaluation diagnostics. evaluation.valid is an evaluator validity check, not a separate scoring metric. Inspect it together with the raw score, candidate program, and verifier evidence when investigating a low score.