Skip to main content
AgentCompass checkpoints completed attempts, writes each completed task detail atomically, and then aggregates those details into request-level metrics. Reuse always materializes compatible data into a newly reserved run directory; it never modifies the source run. The main implementation is RunStore in src/agentcompass/runtime/results/store.py. Detail shaping lives in detail.py, summary construction in summary.py, rendering in render.py, attempt checkpoints under runtime/attempts/, and validated metric types under runtime/metrics/.

Result Directory

RunStore reserves a directory according to the entry point. A single run uses:
A named request in launch uses:
Component IDs and request names are normalized to safe directory components. Launch validates that request output namespaces are distinct after normalization, even if their explicit run IDs differ. Distinct namespaces can use the same run ID. An explicit run ID must not already exist. Without one, the store generates a timestamp-like ID and advances it when necessary to reserve a unique directory. Temporary files are created next to their target and atomically replaced; no permanent staging directory is required. Sensitive configuration is recursively redacted. Metric observations, ground truth, and final answers are evaluation facts and remain unchanged by configuration-key redaction. Task directories use the readable task ID plus its full SHA-256 suffix. <state> is running until the task result is written, then fatal, error or normal according to the most severe final issue of any attempt; warning-only results are normal. Moving a task rewrites the run-relative artifact directories stored in its attempt records. Reuse writes to a new run directory and places tasks whose attempts will be retried in running/. New output has no task-level result.json or error filename prefix. Shared fields live in task.json; each attempt owns its result. The attempt mapping must use attempt-<n>. See Legacy for the previous layout and reuse support.

Result Layers

Each result layer has one producer and one responsibility: Every persisted attempt always contains status, metrics, final_answer, trajectory, error, artifacts, analysis_result, and meta.benchmark/meta.harness. Optional namespace contents may be empty, but the fields remain present. Each attempt result stores its own retry_count; the logical task total and retry_counts map are derived on read. MetricReport records the resolved attempt plan and metric series. Counts satisfy evaluated + unavailable + invalidated = total; error is an overlapping diagnostic. ERROR may retain valid observations. Final FATAL invalidates the whole task across all series, sets evaluation_failed=true, and suppresses every official value. reference_value reports only valid tasks, with explicit coverage; missing observations are never filled with zero.

Resume and Reuse

Resume and reuse share the same persisted records but select their source differently:
  • Resume reads details and attempt checkpoints already present in the current run directory.
  • Reuse resolves another run through reuse_run_id or the newest run within the same output namespace, validates compatibility, then copies usable details and checkpoints into a new run directory.
Existing result directories are not moved or renamed. Automatic reuse searches only the current output hierarchy; it does not fall back to the former Benchmark/Model hierarchy. Before reuse, the store requires the same Benchmark ID and execution.attempts plan. Tasks are matched by task ID and checkpoints by attempt number. Request parameters and task-input fingerprints are not compared, so pending fresh evaluations can use updated evaluation settings. Complete results retain ordinary reuse behavior. Current run-info v3 uses structured issues. Previous v2 results are adapted only at read boundaries; classification-unknown failures cannot be reused for current scoring. Historical directories are not modified. A complete compatible detail without errors is copied as a unit. A detail containing execution or evaluation errors is not reused as a whole, and its k-attempt plan is scheduled again. A fresh-mode attempt with a clean post-agent completion record can resume evaluation from its artifacts; otherwise the agent runs again. If an interrupted source task has no complete detail, compatible attempt checkpoints are materialized independently. For example, if attempts 1 and 2 of k=3 were completed before interruption, resume or reuse schedules only attempt 3, then builds the final detail from all three attempts. Runtime retries within an attempt also leave its completed sibling attempts intact. Resolved execution plans for materialized attempts are copied into the target run_info.json, keeping the new run self-contained. Incompatible Benchmark IDs, attempt plans, or task/attempt identities, and malformed details or checkpoints that must be read, fail explicitly rather than being translated.

Change Checklist

  • Keep every selected task_id stable, non-empty, and unique.
  • Keep attempt observations in metrics, Benchmark evidence in meta.benchmark, and Harness diagnostics in meta.harness.telemetry.
  • Preserve atomic writes, task/attempt identity checks, artifact integrity checks, retry isolation, and per-attempt checkpoints.
  • Validate changes against both a new run and an interrupted k>1 run.
  • Keep every MetricSeries.counts aligned with the series’ actual denominator.
Continue with Runtime Contracts and Planning for in-memory types, or Execution, Scheduling, and Cleanup for attempt and retry behavior.