host_process) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.
How it works
An HLE-Verified run has two stages — inference and judging.Inference and judging
- Inference. The model under test acts as a search agent and, driven by the harness (default
naive_search_agent), completes multi-turn tool loops such as search / visit per task, producing a short natural-language final answer. - Judging. The judge model (
judge_model) receives “question + ground truth + answer under test” and grades it with the built-in A/B/C protocol. The judge compares only the final answer, ignoring reasoning and formatting differences; equivalent expressions are accepted. The judge and the model under test are two separate endpoints;judge_modelmust be specified explicitly.
The A/B/C verdict
The judge returns exactly one verdict, and only A counts as correct:- A — CORRECT: the answer semantically matches the ground truth (equivalent expressions and formatting allowed).
- B — INCORRECT: any deviation from the ground truth.
- C — INCOMPLETE / REPETITIVE / REFUSAL: an invalid answer (cut off mid-sentence, looping repetition, or an explicit refusal).
Parameters
Pass a JSON object via--benchmark-params '{...}', or a benchmark.params block in the YAML given to --config; the CLI wins on shared keys. See the Benchmark overview for merge precedence.
Parameter reference
| Parameter | Type | Default | Choices / values | Description |
|---|---|---|---|---|
judge_model | dict | null | id, base_url, api_key, api_protocol, params | Judge model spec, required (see Judge model spec). It decides grading, and is not the CLI —model-*. |
category | string / list | ”all" | "all”, Math, Physics, Chemistry, Biology/Medicine, Computer Science/AI, Engineering, Humanities/Social Science, Other | Filter tasks by category (same subject taxonomy as HLE); “all” = no filter, a list takes the union. Task counts by category — Math (976), Computer Science/AI (224), Biology/Medicine (222), Physics (202), Humanities/Social Science (193), Other (176), Chemistry (101), Engineering (64); 2158 in total. |
subset | string / list | ”all" | "all”, or one/more of gold / revision / uncertain | Filter by verified subset: gold = Gold subset, revision = Revision subset, uncertain = Uncertain subset. “all” = no filter; a list takes the union. |
sample_ids follow Benchmark Parameters. Configure repeated attempts with --k and --attempt-strategy; see Metrics and Aggregation.
Judge model spec
judge_model is passed as a dict with the fields id, base_url, api_key, api_protocol, and params, pointing to the judge model’s own endpoint, with inference parameters under params.
We recommend fixing a single judge across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. The judge need not be especially strong — the A/B/C criterion (semantic match) is relatively objective, so a mid-sized model suffices. AgentCompass recommends Qwen3.6-35B-A3B.
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. These examples use hle_verified, naive_search_agent, and $MODEL_NAME, with host_process as the Environment.
Set the following environment variables in your terminal before running the examples:
- Model under test:
MODEL_NAME,MODEL_BASE_URL, andMODEL_API_KEY; see Model connection details. - Judge Model:
JUDGE_MODEL_NAME,JUDGE_MODEL_BASE_URL, andJUDGE_MODEL_API_KEY. Use a fixed, independent judge configuration. - Search tools:
SERPER_API_KEYforsearchandJINA_API_KEYforvisit.
- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Use
sample_ids to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
HLE-Verified’s primary metric is binarycorrect: under the A/B/C verdict described above, A maps to true and B/C map to false. There is no partial credit.
With the default configuration, each task has one attempt and the overall score is accuracy over tasks with valid scores, ranging from 0 to 1; higher is better. For example, if all 100 tasks receive valid judgments and 63 are correct, 0.63 in the report means 63%. When you filter tasks with category, subset, modality, or sample_ids, scores cover only the selected tasks.
See Metrics and Aggregation for repeated attempts, category aggregation, and scoring failure rules.
Task Results and Scoring Evidence
After an attempt is scored successfully,scoring under meta.benchmark contains:
This scoring record preserves the verdict and compared answers, but not the judge’s raw response or reasoning.
