Skip to main content
A Benchmark declares its attempt observations with a MetricContract, writes those observations to RunResult.metrics, and returns a validated MetricReport from aggregate_metrics(). Execution status and metric observations are separate contracts.

Declare the Metric Contract

Use make_metric_contract() for the standard binary or scalar shapes:
Every contract has exactly one primary observation: Additional observations may be binary or scalar. Undeclared observations, non-boolean binary values, and non-finite scalar values fail validation. A scalar primary cannot use the pass strategy.

Write Evaluated Observations

Preserve the Harness result with dataclasses.replace() and write only Benchmark observations to metrics:
A valid wrong answer carries a valid zero observation. Preserve input issues, trajectory, artifacts and telemetry when grading; ERROR and WARNING may coexist with valid scores. Store full diagnostics in artifacts.execution_diagnostics. Classify exceptions at the actual call boundary. exception_issue() requires an explicit severity, phase and code, formats a redacted summary, and preserves an existing StageFailure.issue. failed_result() accepts the classified issue and retains partial results and the traceback. Do not ask an outer handler to infer the cause from a source string. Within the same run phase, model timeouts are ERROR, while authentication, quota and server failures are FATAL; model API adapters can reuse runtime.llm.errors.model_exception_issue(). Explicitly classify judge failures as FATAL at the scoring boundary and cleanup failures as WARNING at the release boundary. The detail serializer does not read or write legacy top-level correct and score fields. It always keeps status, metrics, final_answer, trajectory, issues, artifacts, analysis_result, and the benchmark and harness metadata namespaces so task detail files have a stable shape.

Organize Scoring Logic

The Benchmark declares metrics, coordinates evaluation in evaluate(), and writes observations to a RunResult. Organize the grading logic according to the evaluation approach:
  • Answer-based evaluation: keep simple grading directly in evaluate(), or optionally encapsulate prompts, answer parsing, and grading rules in a separate scorer, as in BrowseComp. A scorer is an optional way to organize internal code, not a component that requires separate registration.
  • Test- or verifier-based evaluation: benchmarks such as SWE-bench and TerminalBench can run tests or call a verifier directly and map the results to their declared metrics, without adding a scorer wrapper.

Understand the Two Aggregation Stages

Aggregation has two independent stages:
  1. The runtime applies the selected attempt strategy and metric reducers to the k observations of each task. At k=1 it emits native@1; at k>1, avg requires all attempts, while pass can stop after the first successful binary observation.
  2. Benchmark.aggregate_metrics() combines the reduced task values according to the Benchmark’s official cross-task formula.
BaseBenchmark.aggregate_metrics() implements task mean, category mean, and category hierarchy aggregation for every observation in the contract. A Benchmark whose official metric is a mean can inherit it unchanged:
The default and custom hooks both return MetricReport. Each MetricSeries records its metric_id, reducer, k, value, aggregation method, independent total/evaluated/error/unavailable counts, and optional category and hierarchy breakdowns.

Preserve an Official Corpus Formula

Override aggregate_metrics() only when the official result is not a mean of reduced task values. Start from the shared report and use BenchmarkAggregationContext so attempt reduction, coverage counts, categories, and report validation remain shared:
The context also provides replace_with_sum(), append_ratio_series(), append_benchmark_series(), and update_extra() for totals and Benchmark-derived series. Do not reimplement attempt selection inside this hook or mutate the persisted details supplied by the runtime.

Extend Strategies and Reducers

Repeated-attempt execution and metric calculation are registry-backed extension points:
  • An AttemptStrategy selects reducers, supplies the scheduling policy, and decides whether an attempt can stop further work.
  • A MetricReducer reduces exact observations for one task and declares which metric kinds it supports.
Register new implementations through register_attempt_strategy() and register_metric_reducer(). Keep scheduling policy in the strategy and numeric semantics in the reducer; Benchmark-specific cross-task formulas remain in aggregate_metrics(). Set parallel_attempts_safe = True on a Benchmark only when its per-attempt state is isolated and the selected Harness also supports parallel attempts. The runtime otherwise serializes attempts for that task while still applying request-level task concurrency.

Check Before Aggregating

  • The contract declares every emitted metric and exactly one correct or score primary.
  • Missing observations are never filled from status. Final FATAL invalidates the task; ordinary verifier reward failures must explicitly emit WARNING and fail/0.
  • avg@k has all k valid observations, while a positive pass@k has at least one valid success.
  • Custom corpus formulas retain the official denominator and return context.build().
  • Every task has a stable, non-empty task_id, which is required for attempts, resume, reuse, and metric coverage.
For the persisted output shape, see Task Results and Metrics and Aggregation.