Skip to main content
Each task owns details/<state>/<readable-task-id>--<sha256>/, where <state> is normal, error or fatal according to the most severe final issue of any attempt; warning-only results are normal, and unfinished tasks stay in running. The hash is computed from the exact task ID; sanitization and truncation affect only its readable prefix.

Files

Issues are recorded in the result, not in an _error_ filename prefix; the state directory only groups tasks. Reports, reuse, the result browser and offline analysis reconstruct the logical task record by following the index. The example below shows that reconstructed record; on disk the complete attempts payloads live in separate files.

Persisted Task Metadata

The following example shows the task-level task.json after three attempts have been allocated. New directories use attempt-<n>, where n is the logical attempt number starting at 1. During execution, the identity mapping is written first; shared result metadata is published after the attempt payloads are durable.
There is no task-level result.json. The task retry total and sparse retry_counts map are derived from each attempt’s stored retry_count. An allocated attempt without a result is not a completed observation; a terminal failure without a result payload can retain a retry-only record. Attempt directory names must match their logical numbers: "1": "attempt-1". Cross-run reuse writes attempt-<n> directories in the new run and updates artifact references without changing the source. Retrying an attempt keeps its logical number and archives previous outputs under its own retries/ directory. See Legacy for the previous layout and reuse support.

Task Detail Example

Every task detail uses this fixed standard shell. Empty values remain explicit as null, "", or {}. The ... entries only indicate component-specific content; a namespace with no content is written as {}.

Task-Level Fields

Attempt-Level Fields

Every attempt writes all standard fields. Empty values remain present so task-detail consumers see the same field set for every attempt.

meta Namespaces

Both namespaces are always present and use {} when empty. Their internal component-specific keys are not part of the common task-detail schema. Resolved Environment, Recipe, and network plans are stored in run_info.json.resolved_execution_plans; see Run Records and Diagnostics.

Trajectory Shape

When present, trajectory uses the AgentCompass ACTF_v1.0 shape: Common step fields include step_id, prompts, assistant content, tool calls, observations, timestamps, and metric token/timing values. Harnesses can omit data they do not produce.

Analysis Results

analysis_result.<analyzer-family> can contain is_badcase, score, details, error, and extra. These are analyzer diagnostics and do not alter attempt.metrics or its status. Run-level analysis files combine analyzer output across attempts separately from Benchmark metric aggregation; see Summary and Analysis Results.

Retry Details

When a failure matches the retry policy and budget remains, AgentCompass writes agentcompass.retry.v1 diagnostics before rerunning the current logical attempt:
attempt identifies the unchanged logical attempt; retry is its one-based retry number. stage locates the failure, while scope says whether the runtime repeats the complete attempt or only evaluation. discarded_result is diagnostic and is not required to match the strict task-detail shape. Completed sibling attempts remain checkpointed. For example, a retry in attempt 3 does not rerun attempts 1 and 2. The final task detail records the consumed retry counts, while only the terminal attempt payload contributes observations.

Sensitive Content

AgentCompass redacts recognized credential fields before persistence, but answers, trajectories, errors, artifacts, and integration-specific metadata can still contain sensitive task content. Review them before sharing a run directory.