openclaw, or another compatible productivity / coding harness) to have the model under test complete tasks inside a container in a remote environment; the judge (judge harness) then runs, by default, in a separate new evaluation environment (see Evaluation Environment, Timeout, and Re-judging).
How It Works
End to end, GDPval-AC mainly does two things:-
Inference: the model under test, acting as an agent, completes the GDPVal tasks one by one inside the harness-driven container, writing the required deliverables (usually xlsx / docx / pdf files) into its own workspace. This set of deliverables is the candidate output (output A); after the run it is collected under this layout:
-
Pairwise judging: a judge agent scores the candidate output (A) against the fixed baseline output (B) criterion by criterion, deciding A’s win or loss relative to B. The judge is specified by
judge_model— the command-line--model-*is the model under test, not the judge.
output_a (candidate output), output_b (baseline output), reference (task reference files) and task.json (prompt + rubric). The two sides are shown only under neutral labels A / B with their identities hidden, so the model-under-test’s identity does not bias judging (A is always the candidate, B is always the baseline). The judge evaluates the rubric in batches by window, rather than the whole rubric at once:
judge_rubric_windowsets how many rubric criteria one judge call covers (default16;1= one at a time,0= the whole rubric in one call).- Multiple windows within one task run concurrently, bounded by
judge_concurrency(default8). - A window is the failure blast-radius: if a window call fails or returns an invalid result, only the criteria it covers are affected; the other windows are untouched.
- After the first pass, all failed criteria are collected across windows and re-judged by window, for up to
judge_max_retriesrounds (default3); each round opens a fresh judge session and merges back only the results judged successfully that round.
Fixed Baseline (output B)
Pairwise judging needs a fixed opponent, which is the fixed baseline (output B): the set of deliverables produced by another reference model running inference over all GDPVal tasks, saved as a fixed directory. Every model under test is then compared against the same B, so scores can be compared across models. It is a model-generated set of deliverables — it is neither an official human annotation nor a ground-truth answer. By default the fixed baseline is auto-downloaded viabaseline_zip_url on the first run and extracted into <data_dir>/gdpval_baseline, then the local copy is reused. AgentCompass’s default fixed baseline is generated by claude-opus-4-8, covering all 220 tasks.
Parameters
Parameters fall into two groups: data and inference (which tasks to select, how they land in the container) and pairwise judging (judge model and judging scheduling).Parameter Overview
| Parameter | Type | Default | Allowed values | Description |
|---|---|---|---|---|
sectors | list | [] | One of 9 sectors (full list below) | Filter tasks by sector; empty list = no filter. Intersected with occupations when both are given. |
occupations | list | [] | One of GDPVal’s 44 occupations (full list below) | Filter tasks by occupation; empty list = no filter. Case-insensitive, matched by full name. |
judge_harness | string | openclaw | harness id | Harness used for judging. |
judge_model | dict | null | id, base_url, api_key, api_protocol, params | Judge model spec, required (see Model spec conventions and recommendations). |
judge_harness_params | dict | null | judge_harness params | Params for the judge harness. When the judge uses the same harness as the run, it inherits —harness-params and these take precedence. |
judge_max_turns | int | 100 | integer ≥ 1 | Max turns per judge call; no effect on harnesses without a turn limit, such as openclaw. |
judge_concurrency | int | 8 | integer ≥ 1 | Number of judging windows run concurrently within one task; 1 = serial. |
judge_rubric_window | int | 16 | integer ≥ 0 | How many rubric criteria per judge call: 1 = per-item, N > 1 = N per window, 0 = whole rubric in one call. |
judge_max_retries | int | 3 | integer ≥ 0 | Re-judge rounds after a rubric criterion fails; 0 = disabled. |
All 9 possible values for sectors (click to expand)
All 9 possible values for sectors (click to expand)
Each list item is one complete value accepted by
sectors:- Finance and Insurance
- Government
- Health Care and Social Assistance
- Information
- Manufacturing
- Professional, Scientific, and Technical Services
- Real Estate and Rental and Leasing
- Retail Trade
- Wholesale Trade
All 44 possible values for occupations (click to expand)
All 44 possible values for occupations (click to expand)
Each list item is one complete value accepted by
occupations:- Accountants and Auditors
- Administrative Services Managers
- Audio and Video Technicians
- Buyers and Purchasing Agents
- Child, Family, and School Social Workers
- Compliance Officers
- Computer and Information Systems Managers
- Concierges
- Counter and Rental Clerks
- Customer Service Representatives
- Editors
- Film and Video Editors
- Financial Managers
- Financial and Investment Analysts
- First-Line Supervisors of Non-Retail Sales Workers
- First-Line Supervisors of Office and Administrative Support Workers
- First-Line Supervisors of Police and Detectives
- First-Line Supervisors of Production and Operating Workers
- First-Line Supervisors of Retail Sales Workers
- General and Operations Managers
- Industrial Engineers
- Lawyers
- Mechanical Engineers
- Medical Secretaries and Administrative Assistants
- Medical and Health Services Managers
- News Analysts, Reporters, and Journalists
- Nurse Practitioners
- Order Clerks
- Personal Financial Advisors
- Pharmacists
- Private Detectives and Investigators
- Producers and Directors
- Project Management Specialists
- Property, Real Estate, and Community Association Managers
- Real Estate Brokers
- Real Estate Sales Agents
- Recreation Workers
- Registered Nurses
- Sales Managers
- Sales Representatives, Wholesale and Manufacturing, Except Technical and Scientific Products
- Sales Representatives, Wholesale and Manufacturing, Technical and Scientific Products
- Securities, Commodities, and Financial Services Sales Agents
- Shipping, Receiving, and Inventory Clerks
- Software Developers
Model Spec Conventions and Recommendations
judge_model is passed as a dict with the fields id, base_url, api_key, api_protocol, and params, pointing at the judge model’s own endpoint, with model inference parameters under params. Specify a fixed and sufficiently strong judge, since it decides the evaluation’s win/loss; using the model under test as its own judge is neither fair nor comparable across models.
Judging Scheduling
Concurrency and fault tolerance within a single task are controlled by three parameters; they generally need no change and should be adjusted only when judge throughput or stability becomes a bottleneck:judge_rubric_window— balances “how many rubric criteria per call” against “failure blast-radius”: larger reduces the number of calls and grows the per-call context, smaller is more fine-grained with a smaller failure footprint.judge_concurrency— the number of windows judged simultaneously within one task; larger improves per-task judge-stage throughput (across tasks is already parallelized by--task-concurrency).judge_max_retries— the number of re-judge rounds for judge-stage failures (timeouts, invalid schema, etc.), each round opening a fresh judge session.
Evaluation Environment, Timeout, and Re-judging
Evaluation environment. After inference, the runtime takes a snapshot of the task workspace<workspace_root>/<task_id> and saves it in the run directory, closes the inference Environment, opens a new evaluation Environment, restores the snapshot, uploads the reference files again, and runs the judge there. The snapshot holds only what the model under test produced: reference files are left out, and symbolic links and special files such as pipes and sockets are dropped (for example a virtual environment’s links to the image’s interpreter). Files and directories that the model made unreadable get their owner’s read permission back and are included; entries that still cannot be read are skipped and listed in the attempt’s gdpval_ac_snapshot_skipped. Inference and judging are therefore isolated:
- A judging failure reopens only the evaluation Environment and judges again; inference is not rerun (bounded by
max_retriesin--execution-params). - The saved snapshot can be judged again later; see “Re-judging only failed tasks” below.
- Files written before an inference timeout are still part of the snapshot and are judged normally.
- This mode requires an absolute
workspace_rootand cannot be combined withartifacts/artifact_collectin--execution-params; each task opens one more Environment. The workspace takes a second copy of disk space inside the inference Environment, and each snapshot is bounded byartifact_limits(by default 16 GiB, 100,000 files, and 600 seconds of transfer).
run_timeout_seconds and evaluation_timeout_seconds in --execution-params, or scale them with timeout_multiplier, run_timeout_multiplier, and evaluation_timeout_multiplier. The judging budget covers the whole judging stage, including every window and re-judge round; each judge run uses the remaining budget as its wall clock, and when the budget runs out the task is recorded as a judging failure, its rubric maximum still counts in the denominator, and it is judged again under the rules above.
Re-judging only failed tasks. Rerun the same command with --reuse <run-id> and max_retries ≥ 1: tasks whose judging failed are restored from the saved snapshot and only judged again, while completed tasks are reused as they are. Judge parameters such as judge_model may change between the two runs. Results produced in reuse mode have no snapshot, so those tasks rerun inference as well.
Pre-run checks. After tasks are loaded and before inference starts, two checks run. Every selected task’s reference files are resolved on the host (missing ones are downloaded); files that cannot be obtained are all listed together and the run stops. One minimal request is sent to the judge model; the run stops when the endpoint rejects the key or the model (HTTP 400 / 401 / 403 / 404 / 422), while an unreachable endpoint or a transient error is only logged as a warning.
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. These examples use gdpval_ac, openclaw, and $MODEL_NAME, with docker as the Environment.
Set the following environment variables in your terminal before running the examples:
- Model under test:
MODEL_NAME,MODEL_BASE_URL, andMODEL_API_KEY; see Model connection details. - Judge Model:
JUDGE_MODEL_NAME,JUDGE_MODEL_BASE_URL, andJUDGE_MODEL_API_KEY. Use a fixed, independent judge configuration.
- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Verify the pipeline runs end to end — use
sample_ids to run just one task all the way through inference and judging, leaving everything else at defaults.Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
GDPVal’s scalar primary metricscore is candidate A’s normalized rubric score, ranging from 0 to 1; higher is better. It measures a different outcome from A’s win rate against fixed baseline B.
Overall
score is weighted by rubric maximum, rather than a simple mean of per-task normalized scores. Use the per-criterion evidence below to inspect scores and wins. All original observations are scalar, so the pass execution strategy is unsupported.
See Metrics and Aggregation for repeated attempts, category aggregation, and shared scoring failure rules.
Task Results and Scoring Evidence
Within each attempt’smeta.benchmark, gdpval_ac_pairwise stores task_a (candidate) and task_b (baseline), each containing:
Collected deliverables are indexed by
gdpval_ac_deliverable_files in artifacts. Required and missing files are recorded in gdpval_ac_expected_deliverables and, when present, gdpval_ac_missing_deliverables. Successfully collected workspace files and raw judgments are under the run directory:
