MetricContract, writes those observations to RunResult.metrics, and returns a validated MetricReport from aggregate_metrics(). Execution status and metric observations are separate contracts.
Declare the Metric Contract
Usemake_metric_contract() for the standard binary or scalar shapes:
Additional observations may be binary or scalar. Undeclared observations, non-boolean binary values, and non-finite scalar values fail validation. A scalar primary cannot use the
pass strategy.
Write Evaluated Observations
Preserve the Harness result withdataclasses.replace() and write only Benchmark observations to metrics:
exception_issue() requires an explicit severity, phase and code, formats a redacted summary, and preserves an existing StageFailure.issue. failed_result() accepts the classified issue and retains partial results and the traceback. Do not ask an outer handler to infer the cause from a source string. Within the same run phase, model timeouts are ERROR, while authentication, quota and server failures are FATAL; model API adapters can reuse runtime.llm.errors.model_exception_issue(). Explicitly classify judge failures as FATAL at the scoring boundary and cleanup failures as WARNING at the release boundary.
The detail serializer does not read or write legacy top-level correct and score fields. It always keeps status, metrics, final_answer, trajectory, issues, artifacts, analysis_result, and the benchmark and harness metadata namespaces so task detail files have a stable shape.
Organize Scoring Logic
The Benchmark declares metrics, coordinates evaluation inevaluate(), and writes observations to a RunResult. Organize the grading logic according to the evaluation approach:
- Answer-based evaluation: keep simple grading directly in
evaluate(), or optionally encapsulate prompts, answer parsing, and grading rules in a separate scorer, as in BrowseComp. A scorer is an optional way to organize internal code, not a component that requires separate registration. - Test- or verifier-based evaluation: benchmarks such as SWE-bench and TerminalBench can run tests or call a verifier directly and map the results to their declared metrics, without adding a scorer wrapper.
Understand the Two Aggregation Stages
Aggregation has two independent stages:- The runtime applies the selected attempt strategy and metric reducers to the
kobservations of each task. Atk=1it emitsnative@1; atk>1,avgrequires all attempts, whilepasscan stop after the first successful binary observation. Benchmark.aggregate_metrics()combines the reduced task values according to the Benchmark’s official cross-task formula.
BaseBenchmark.aggregate_metrics() implements task mean, category mean, and category hierarchy aggregation for every observation in the contract. A Benchmark whose official metric is a mean can inherit it unchanged:
MetricReport. Each MetricSeries records its metric_id, reducer, k, value, aggregation method, independent total/evaluated/error/unavailable counts, and optional category and hierarchy breakdowns.
Preserve an Official Corpus Formula
Overrideaggregate_metrics() only when the official result is not a mean of reduced task values. Start from the shared report and use BenchmarkAggregationContext so attempt reduction, coverage counts, categories, and report validation remain shared:
replace_with_sum(), append_ratio_series(), append_benchmark_series(), and update_extra() for totals and Benchmark-derived series. Do not reimplement attempt selection inside this hook or mutate the persisted details supplied by the runtime.
Extend Strategies and Reducers
Repeated-attempt execution and metric calculation are registry-backed extension points:- An
AttemptStrategyselects reducers, supplies the scheduling policy, and decides whether an attempt can stop further work. - A
MetricReducerreduces exact observations for one task and declares which metric kinds it supports.
register_attempt_strategy() and register_metric_reducer(). Keep scheduling policy in the strategy and numeric semantics in the reducer; Benchmark-specific cross-task formulas remain in aggregate_metrics().
Set parallel_attempts_safe = True on a Benchmark only when its per-attempt state is isolated and the selected Harness also supports parallel attempts. The runtime otherwise serializes attempts for that task while still applying request-level task concurrency.
Check Before Aggregating
- The contract declares every emitted metric and exactly one
correctorscoreprimary. - Missing observations are never filled from status. Final FATAL invalidates the task; ordinary verifier reward failures must explicitly emit WARNING and fail/0.
avg@khas allkvalid observations, while a positivepass@khas at least one valid success.- Custom corpus formulas retain the official denominator and return
context.build(). - Every task has a stable, non-empty
task_id, which is required for attempts, resume, reuse, and metric coverage.
