RunStore in src/agentcompass/runtime/results/store.py. Detail shaping lives in detail.py, summary construction in summary.py, rendering in render.py, attempt checkpoints under runtime/attempts/, and validated metric types under runtime/metrics/.
Result Directory
RunStore reserves a directory according to the entry point. A single run uses:
launch uses:
Temporary files are created next to their target and atomically replaced; no permanent staging directory is required. Sensitive configuration is recursively redacted. Metric observations, ground truth, and final answers are evaluation facts and remain unchanged by configuration-key redaction.
Task directories use the readable task ID plus its full SHA-256 suffix.
<state> is running until the task result is written, then fatal, error or normal according to the most severe final issue of any attempt; warning-only results are normal. Moving a task rewrites the run-relative artifact directories stored in its attempt records. Reuse writes to a new run directory and places tasks whose attempts will be retried in running/. New output has no task-level result.json or error filename prefix. Shared fields live in task.json; each attempt owns its result. The attempt mapping must use attempt-<n>. See Legacy for the previous layout and reuse support.
Result Layers
Each result layer has one producer and one responsibility:
Every persisted attempt always contains
status, metrics, final_answer, trajectory, error, artifacts, analysis_result, and meta.benchmark/meta.harness. Optional namespace contents may be empty, but the fields remain present. Each attempt result stores its own retry_count; the logical task total and retry_counts map are derived on read.
MetricReport records the resolved attempt plan and metric series. Counts satisfy evaluated + unavailable + invalidated = total; error is an overlapping diagnostic. ERROR may retain valid observations. Final FATAL invalidates the whole task across all series, sets evaluation_failed=true, and suppresses every official value. reference_value reports only valid tasks, with explicit coverage; missing observations are never filled with zero.
Resume and Reuse
Resume and reuse share the same persisted records but select their source differently:- Resume reads details and attempt checkpoints already present in the current run directory.
- Reuse resolves another run through
reuse_run_idor the newest run within the same output namespace, validates compatibility, then copies usable details and checkpoints into a new run directory.
execution.attempts plan. Tasks are matched by task ID and checkpoints by attempt number. Request parameters and task-input fingerprints are not compared, so pending fresh evaluations can use updated evaluation settings. Complete results retain ordinary reuse behavior.
Current run-info v3 uses structured issues. Previous v2 results are adapted only at read boundaries; classification-unknown failures cannot be reused for current scoring. Historical directories are not modified.
A complete compatible detail without errors is copied as a unit. A detail containing execution or evaluation errors is not reused as a whole, and its k-attempt plan is scheduled again. A fresh-mode attempt with a clean post-agent completion record can resume evaluation from its artifacts; otherwise the agent runs again.
If an interrupted source task has no complete detail, compatible attempt checkpoints are materialized independently. For example, if attempts 1 and 2 of k=3 were completed before interruption, resume or reuse schedules only attempt 3, then builds the final detail from all three attempts. Runtime retries within an attempt also leave its completed sibling attempts intact.
Resolved execution plans for materialized attempts are copied into the target run_info.json, keeping the new run self-contained. Incompatible Benchmark IDs, attempt plans, or task/attempt identities, and malformed details or checkpoints that must be read, fail explicitly rather than being translated.
Change Checklist
- Keep every selected
task_idstable, non-empty, and unique. - Keep attempt observations in
metrics, Benchmark evidence inmeta.benchmark, and Harness diagnostics inmeta.harness.telemetry. - Preserve atomic writes, task/attempt identity checks, artifact integrity checks, retry isolation, and per-attempt checkpoints.
- Validate changes against both a new run and an interrupted
k>1run. - Keep every
MetricSeries.countsaligned with the series’ actual denominator.
