> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Results and Aggregation

A Benchmark declares its attempt observations with a `MetricContract`, writes those observations to `RunResult.metrics`, and returns a validated `MetricReport` from `aggregate_metrics()`. Execution status and metric observations are separate contracts.

## Declare the Metric Contract

Use `make_metric_contract()` for the standard binary or scalar shapes:

```python theme={"system"}
from agentcompass.runtime.metrics import make_metric_contract


class ExampleBenchmark(BaseBenchmark):
    metric_contract = make_metric_contract(
        primary="correct",
        binary=("correct", "format_valid"),
        scalar=("reward",),
        labels={"correct": "Accuracy", "reward": "Reward"},
    )
```

Every contract has exactly one primary observation:

| Primary | Kind | Repeated-attempt support |
| - | - | - |
| `correct` | `binary_success` | `avg` and `pass` |
| `score` | `scalar` | `avg` only |

Additional observations may be binary or scalar. Undeclared observations, non-boolean binary values, and non-finite scalar values fail validation. A scalar primary cannot use the `pass` strategy.

## Write Evaluated Observations

Preserve the Harness result with `dataclasses.replace()` and write only Benchmark observations to `metrics`:

```python theme={"system"}
from dataclasses import replace

from agentcompass.runtime import ExecutionIssue, RunResult, TaskStatus


def apply_verifier_result(result: RunResult, verifier) -> RunResult:
    if verifier.timed_out or verifier.returncode not in {0, 1}:
        detail = (verifier.stderr or verifier.stdout).strip()
        status = (
            TaskStatus.ERROR
            if result.status == TaskStatus.RUN_ERROR
            else TaskStatus.EVAL_ERROR
        )
        return replace(
            result,
            status=status,
            issues=[*result.issues, ExecutionIssue("error", "evaluate", "verifier_failed", detail or "verifier failed")],
            metrics={},
        )

    passed = verifier.returncode == 0
    return replace(
        result,
        metrics={
            "correct": passed,
            "format_valid": verifier.format_valid,
            "reward": verifier.reward,
        },
    )
```

A valid wrong answer carries a valid zero observation. Preserve input issues, trajectory, artifacts and telemetry when grading; ERROR and WARNING may coexist with valid scores. Store full diagnostics in artifacts.execution\_diagnostics.

Classify exceptions at the actual call boundary. `exception_issue()` requires an explicit severity, phase and code, formats a redacted summary, and preserves an existing `StageFailure.issue`. `failed_result()` accepts the classified issue and retains partial results and the traceback. Do not ask an outer handler to infer the cause from a `source` string. Within the same `run` phase, model timeouts are ERROR, while authentication, quota and server failures are FATAL; model API adapters can reuse `runtime.llm.errors.model_exception_issue()`. Explicitly classify judge failures as FATAL at the scoring boundary and cleanup failures as WARNING at the release boundary.

The detail serializer does not read or write legacy top-level `correct` and `score` fields. It always keeps `status`, `metrics`, `final_answer`, `trajectory`, `issues`, `artifacts`, `analysis_result`, and the `benchmark` and `harness` metadata namespaces so task detail files have a stable shape.

## Organize Scoring Logic

The Benchmark declares metrics, coordinates evaluation in `evaluate()`, and writes observations to a `RunResult`. Organize the grading logic according to the evaluation approach:

* Answer-based evaluation: keep simple grading directly in `evaluate()`, or optionally encapsulate prompts, answer parsing, and grading rules in a separate scorer, as in [BrowseComp](https://github.com/open-compass/AgentCompass/blob/main/src/agentcompass/benchmarks/browsecomp.py). A scorer is an optional way to organize internal code, not a component that requires separate registration.
* Test- or verifier-based evaluation: benchmarks such as SWE-bench and TerminalBench can run tests or call a verifier directly and map the results to their declared metrics, without adding a scorer wrapper.

## Understand the Two Aggregation Stages

Aggregation has two independent stages:

1. The runtime applies the selected attempt strategy and metric reducers to the `k` observations of each task. At `k=1` it emits `native@1`; at `k>1`, `avg` requires all attempts, while `pass` can stop after the first successful binary observation.
2. `Benchmark.aggregate_metrics()` combines the reduced task values according to the Benchmark's official cross-task formula.

`BaseBenchmark.aggregate_metrics()` implements task mean, category mean, and category hierarchy aggregation for every observation in the contract. A Benchmark whose official metric is a mean can inherit it unchanged:

```python theme={"system"}
class ExampleBenchmark(BaseBenchmark):
    metric_contract = make_metric_contract(
        primary="score",
        scalar=("score",),
    )

    # load_tasks(), prepare_task(), evaluate() ...
```

The default and custom hooks both return `MetricReport`. Each `MetricSeries` records its `metric_id`, reducer, `k`, value, aggregation method, independent `total`/`evaluated`/`error`/`unavailable` counts, and optional category and hierarchy breakdowns.

## Preserve an Official Corpus Formula

Override `aggregate_metrics()` only when the official result is not a mean of reduced task values. Start from the shared report and use `BenchmarkAggregationContext` so attempt reduction, coverage counts, categories, and report validation remain shared:

```python theme={"system"}
from agentcompass.runtime.metrics import BenchmarkAggregationContext


def aggregate_metrics(self, results, req, config):
    report = super().aggregate_metrics(results, req, config)
    context = BenchmarkAggregationContext(
        results,
        contract=self.metric_contract,
        report=report,
        category_hierarchy=getattr(config, "category_hierarchy", None),
    )
    context.replace_with_ratio_of_sums(
        output_metric="score",
        numerator_metric="points_earned",
        denominator_metric="points_available",
        label="Score",
    )
    return context.build()
```

The context also provides `replace_with_sum()`, `append_ratio_series()`, `append_benchmark_series()`, and `update_extra()` for totals and Benchmark-derived series. Do not reimplement attempt selection inside this hook or mutate the persisted details supplied by the runtime.

## Extend Strategies and Reducers

Repeated-attempt execution and metric calculation are registry-backed extension points:

* An `AttemptStrategy` selects reducers, supplies the scheduling policy, and decides whether an attempt can stop further work.
* A `MetricReducer` reduces exact observations for one task and declares which metric kinds it supports.

Register new implementations through `register_attempt_strategy()` and `register_metric_reducer()`. Keep scheduling policy in the strategy and numeric semantics in the reducer; Benchmark-specific cross-task formulas remain in `aggregate_metrics()`.

Set `parallel_attempts_safe = True` on a Benchmark only when its per-attempt state is isolated and the selected Harness also supports parallel attempts. The runtime otherwise serializes attempts for that task while still applying request-level task concurrency.

## Check Before Aggregating

* The contract declares every emitted metric and exactly one `correct` or `score` primary.
* Missing observations are never filled from status. Final FATAL invalidates the task; ordinary verifier reward failures must explicitly emit WARNING and fail/0.
* `avg@k` has all `k` valid observations, while a positive `pass@k` has at least one valid success.
* Custom corpus formulas retain the official denominator and return `context.build()`.
* Every task has a stable, non-empty `task_id`, which is required for attempts, resume, reuse, and metric coverage.

For the persisted output shape, see [Task Results](/en/user_guide/other_features/results/task_results) and [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.