How it works
- Prepare the task. AgentCompass downloads and validates the dataset when necessary. The Docker recipe selects the WildClawBench OpenClaw image, prepares the public task data, skills, warm-up commands, and task workspace, while keeping private ground truth outside the inference environment.
- Run OpenClaw. The prompt and task-specific timeout are passed to the OpenClaw harness, which operates in the prepared workspace. WildClawBench requires a Brave Search credential.
- Run Automated Checks. After inference, AgentCompass decrypts and uploads only the current task’s ground truth, executes the task’s Automated Checks inside the same environment, converts
overall_scoreinto the task score, and appliespass_thresholdto producepassed.
Parameters
Configure WildClawBench-specific options with--benchmark-params '{...}'.
| Parameter | Type | Default | Choices / values | Description |
|---|---|---|---|---|
category | string / list | ”all" | "all”, one category, or a list | Filter tasks by category; a list takes the union. |
pass_threshold | float | 1.0 | numeric score | Minimum Automated Checks score required for passed=true. |
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. These examples use wildclawbench, openclaw, and $MODEL_NAME, with docker as the Environment.
Set the following environment variables in your terminal before running the examples:
- Model under test:
MODEL_NAME,MODEL_BASE_URL, andMODEL_API_KEY; see Model connection details. - Web search:
BRAVE_API_KEY.
- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Verify the end-to-end flow works —
sample_ids selects which case to run, with all other parameters using their defaults.Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
WildClawBench’s primary metric is scalarscore, taken from overall_score returned by the automated checks described above, falling back to score when that field is absent. Higher is better. Each task’s checker defines the scale; AgentCompass does not further normalize or clamp it to 0–1.
Auxiliary binary passed indicates whether the score reaches pass_threshold, which defaults to 1.0. With the default configuration, aggregate results show the mean score and pass rate; filters restrict scores to the selected tasks. Repeated attempts support only the avg execution strategy; auxiliary passed does not enable pass. See Metrics and Aggregation.
Task Results and Scoring Evidence
scoring under meta.benchmark preserves automated-check evidence:
The raw checker may also provide
correct; the aggregate pass rate uses metrics.passed, calculated from the configured threshold.