Skip to main content
WildClawBench (arxiv) evaluates an agent on real-world, long-horizon productivity tasks in executable workspaces. AgentCompass uses OpenClaw to perform each task and runs the task’s Automated Checks afterward. By default, a missing optional Python dependency reports the required extra and installation command; when auto-install is enabled, AgentCompass first attempts to install it. See Dependencies.

How it works

  1. Prepare the task. AgentCompass downloads and validates the dataset when necessary. The Docker recipe selects the WildClawBench OpenClaw image, prepares the public task data, skills, warm-up commands, and task workspace, while keeping private ground truth outside the inference environment.
  2. Run OpenClaw. The prompt and task-specific timeout are passed to the OpenClaw harness, which operates in the prepared workspace. WildClawBench requires a Brave Search credential.
  3. Run Automated Checks. After inference, AgentCompass decrypts and uploads only the current task’s ground truth, executes the task’s Automated Checks inside the same environment, converts overall_score into the task score, and applies pass_threshold to produce passed.

Parameters

Configure WildClawBench-specific options with --benchmark-params '{...}'.
ParameterTypeDefaultChoices / valuesDescription
categorystring / list”all""all”, one category, or a listFilter tasks by category; a list takes the union.
pass_thresholdfloat1.0numeric scoreMinimum Automated Checks score required for passed=true.

Run examples

The three positional arguments to agentcompass run are Benchmark, Harness, and Model. These examples use wildclawbench, openclaw, and $MODEL_NAME, with docker as the Environment. Set the following environment variables in your terminal before running the examples:
  • Model under test: MODEL_NAME, MODEL_BASE_URL, and MODEL_API_KEY; see Model connection details.
  • Web search: BRAVE_API_KEY.
See the run command for configuration ownership and CLI overrides. The Docker Recipe provides the task image, and OpenClaw uses Brave for web search.
Verify the end-to-end flow works — sample_ids selects which case to run, with all other parameters using their defaults.

Evaluation Results

For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.

Scoring Metrics

WildClawBench’s primary metric is scalar score, taken from overall_score returned by the automated checks described above, falling back to score when that field is absent. Higher is better. Each task’s checker defines the scale; AgentCompass does not further normalize or clamp it to 0–1. Auxiliary binary passed indicates whether the score reaches pass_threshold, which defaults to 1.0. With the default configuration, aggregate results show the mean score and pass rate; filters restrict scores to the selected tasks. Repeated attempts support only the avg execution strategy; auxiliary passed does not enable pass. See Metrics and Aggregation.

Task Results and Scoring Evidence

scoring under meta.benchmark preserves automated-check evidence: The raw checker may also provide correct; the aggregate pass rate uses metrics.passed, calculated from the configured threshold.