Files at a Glance
| File | When it is generated | What to read it for |
|---|---|---|
summary.md | Benchmark metric aggregation succeeds. | Readable metrics adapted to single or repeated attempts. |
metrics.json | From the same metric report as summary.md. | Canonical values, per-series counts, categories, and hierarchy. |
analysis_summary.json | At least one saved attempt contains aggregatable analyzer output. | Analyzer statistics, analyzer error counts, bad-case file indexes, and distributions. |
analysis_summary.md | From the same analyzer aggregation as the JSON file. | Readable overall, category, and distribution analysis. |
summary or analysis command.
Benchmark Metric Outputs
Both Benchmark metric files are projections of one strictly validated metric report. They do not run independent aggregators.summary.md
Start here for a readable result. For every k, the Markdown starts with the Model, Total, Evaluated, Error, and Metrics. The top-level counts come from the primary series selected by the current strategy; other series can use different valid samples, so use the counts on their own rows when reading them.
For k=1, Metrics keeps the traditional two-column name/value table, followed by available category or hierarchy details. For k>1, only the metric area expands: the summary adds a one-line attempt plan, and the Metrics table distinguishes headline and auxiliary series with Role while listing each reducer, actual run-level formula, value, and independent coverage counts. Complete structured breakdowns remain in metrics.json.
For k>1, the summary does not include a first result. Under the avg strategy, a binary primary can show both avg@k and pass@k, while a scalar primary itself produces only avg@k. If that scalar-primary contract also declares binary auxiliary metrics, each binary auxiliary can still produce its own avg@k and pass@k series. The Benchmark still cannot use the pass execution strategy because that strategy is controlled by the scalar primary. See Metrics and Aggregation.
metrics.json
metrics.json is the machine-readable source of truth:
| Field | Meaning |
|---|---|
k, strategy | Resolved repeated-attempt plan. |
aggregation | Actual run-level policy: micro_weighted, category_mean, or category_hierarchy. |
series[].series_id | Stable <metric>.<reducer>@<k> identity. |
series[].kind | binary_success or scalar. |
series[].role | headline for the Benchmark primary metric, otherwise auxiliary. |
series[].aggregation | The formula that produced this series, such as micro_weighted, ratio_of_sums, sum, or benchmark. |
series[].value | Aggregated value, or null when it cannot be computed exactly. |
series[].counts | Independent total, evaluated, error, unavailable, and invalidated task counts for this series. |
series[].categories | Category keys mapped to their own value, counts, and optional formula aggregation_weight. |
series[].hierarchy | Hierarchy paths mapped to the same breakdown fields; populated only for explicit hierarchy aggregation. |
series[].extra | Auditable inputs for this formula, such as numerator and denominator metric IDs and totals. |
extra | Benchmark-level structured diagnostics, such as rank or medal comparison details. |
evaluated + unavailable + invalidated equals total, while error is an independent diagnostic and may overlap either count.
Analysis Summaries
Analyzer output first appears underattempts.<N>.analysis_result.<analyzer-family>. Analysis aggregation is independent of the Benchmark Metric Contract.
Combine Multiple Attempts
AgentCompass does not choose a generic “best attempt.” For each task and analyzer family it combines saved attempts as follows:is_badcaseis true when any attempt reports true.- Numeric analyzer scores are averaged across attempts that provide one.
- The latest non-empty analyzer payload supplies diagnostic fields used by distributions.
- The task is counted at most once for that analyzer.
avg_score uses the maximum available family score for each task before averaging tasks.
analysis_summary.json
| Field | Contents |
|---|---|
per_category_per_analyzer | One statistics row for each retained category and analyzer combination. |
per_category_overall | One row per category, combining analyzer families. |
overall_per_analyzer | One row per analyzer across categories, with an items list of matching bad-case detail files. |
overall | Statistics combined across categories and analyzers. |
distributions | Analyzer-declared value counts or numeric distributions. |
| Field | Meaning |
|---|---|
category | Task category; overall rows use overall, and uncategorized tasks use (no category). |
analyzer | Analyzer family ID; combined rows use overall. |
total | Tasks in this scope that contain the analysis result. |
badcase_count, badcase_ratio | Number and fraction of tasks marked as bad cases. |
error_count | Tasks whose selected analysis result contains a non-empty error. |
avg_score | Average available analyzer score, or null. |
items | Only in overall_per_analyzer; filenames marked by that analyzer. |
value_counts or numeric_stats. Value counts keep the 50 most frequent values. Numeric output contains count, min, mean, p50, p90, p95, and max when data exists.
In rows that combine all analyzers, badcase_count is the number of tasks marked by at least one analyzer and error_count is the number with at least one analyzer error; neither is the sum of the per-analyzer rows. Bad-case analyzers with neither bad cases nor errors can be omitted, while statistics-only analyzers and analyzers with errors remain. The Markdown version presents Total, Badcase, Error, Badcase Ratio, and Avg Score tables plus distributions, but omits the full items indexes.
Generate or Regenerate Outputs
agentcompass run and agentcompass launch generate Benchmark outputs after aggregation. agentcompass summary strictly reloads task details and the persisted attempt plan, then replaces summary.md and metrics.json without rerunning attempts or analyzers. Both paths update run_info.json.metric_artifacts with the files’ source and report plan.
agentcompass analysis can update per-attempt analysis_result and regenerate the two analysis summary files without rerunning the agent or Benchmark verifier. It does not update Benchmark metric outputs.
Share Results Safely
Generated files can contain task IDs, categories, analyzer values, or other integration data. They do not receive a complete content-safety or Markdown sanitization pass. Inspect them before sharing.Related Pages
A final FATAL setsevaluation_failed=true. All official value fields become null; reference_value uses only non-invalidated tasks and is null when none can be scored. The CLI exits nonzero and the failed run retains its result paths.