correct observation.
FrontierScience uses single-sided judging. The judge only assesses the agent-under-test’s answer against the reference, without comparing to any baseline. Both inference and judging run in the local process (host_process) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.
How it works
A FrontierScience run has two stages — inference and judging — where the judging stage applies the grading rule that matches each task’s type.Inference and judging
- Inference. The model under test acts as a search agent and, driven by the harness (default
naive_search_agent), completes multi-turn tool loops such as search / visit per task, producing a natural-language answer. - Judging. The judge model (
judge_model) receives the question, the reference (a short reference answer or a scoring rubric, depending on the task type), and the answer under test, then grades it. The judge and the model under test are two separate endpoints;judge_modelmust be specified explicitly.
How the two task types are graded
The grading rule is not selected by a run parameter — it is determined by the task itself. A task whosecategory is research is graded as **FrontierScience-Research ; otherwise (olympiad) it is graded as ** FrontierScience-Olympiad. Because this is decided per task, a single run can contain both.
-
FrontierScience-Olympiad — short-answer grading. The reference is one short answer: a number, a symbolic expression, or a short phrase. The judge compares the candidate’s final answer to it and:
- accepts mathematically equivalent expressions and harmless formatting differences;
- accepts minor wording differences that preserve the same scientific meaning;
- grades incorrect ** if the candidate states ** multiple conflicting final answers;
- judges only what the candidate actually wrote, without supplying missing steps on its behalf.
correct. -
FrontierScience-Research — rubric grading. The reference is a multi-item scoring rubric worth 10 points in total. The judge scores the answer ** item by item , awarding partial credit per item (each capped at that item’s max points), then sums the awarded points into a total on a 0–10 scale. Both final conclusions and intermediate reasoning steps can earn points, but only what the answer actually supports is credited — unstated work earns nothing. The task is ** correct ** when the total score is ** at least the pass threshold
research_pass_threshold(default7.0).
Parameters
Pass a JSON object via--benchmark-params '{...}', or a benchmark.params block in the YAML given to --config; the CLI wins on shared keys. See the Benchmark overview for merge precedence.
Parameter reference
| Parameter | Type | Default | Choices / values | Description |
|---|---|---|---|---|
judge_model | dict | null | id, base_url, api_key, api_protocol, params | Judge model spec, required (see Judge model spec). It decides grading, and is not the CLI —model-*. |
category | string / list | ”all" | "all”, olympiad, research | Filter tasks by category; “all” = no filter, a list takes the union. The two categories are the benchmark’s task types — olympiad (100 tasks) and research (60), 160 in total — so filtering by category also determines which grading rule the run uses. |
subject | string | ”all” | all / physics / chemistry / biology | Filter by scientific subject — a single value, not a list. all = no filter. Must be non-empty. |
research_pass_threshold | float | 7.0 | 0.0–10.0 | Pass mark for FrontierScience-Research tasks, on the rubric’s 0–10 scale: such a task is correct when its total rubric score is ≥ this value. Raise it to be stricter, lower it to be more lenient. Has no effect on olympiad short-answer tasks. |
sample_ids follow Benchmark Parameters. Configure repeated attempts with --k and --attempt-strategy; see Metrics and Aggregation.
Judge model spec
judge_model is passed as a dict with the fields id, base_url, api_key, api_protocol, and params, pointing to the judge model’s own endpoint, with inference parameters under params.
We recommend fixing a single judge across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. AgentCompass recommends Qwen3.6-35B-A3B. Note that research-rubric grading is more nuanced than short-answer grading — it involves item-by-item scoring with partial credit — so a stronger, more capable judge improves rubric reliability.
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. These examples use frontierscience, naive_search_agent, and $MODEL_NAME, with host_process as the Environment.
Set the following environment variables in your terminal before running the examples:
- Model under test:
MODEL_NAME,MODEL_BASE_URL, andMODEL_API_KEY; see Model connection details. - Judge Model:
JUDGE_MODEL_NAME,JUDGE_MODEL_BASE_URL, andJUDGE_MODEL_API_KEY. Use a fixed, independent judge configuration. - Search tools:
SERPER_API_KEYforsearchandJINA_API_KEYforvisit.
- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Use
sample_ids to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
FrontierScience’s primary metric is binarycorrect, determined by the task-type grading rules: Olympiad tasks use the judge’s boolean verdict, while Research tasks pass when their rubric total reaches research_pass_threshold. Research item scores and total_score explain the threshold decision; they are not separate aggregate metrics.
With the default configuration, each task has one attempt and the overall score is the pass rate over tasks with valid scores, ranging from 0 to 1; higher is better. When both task types are included, each task has equal weight; the overall score is not the mean of the two task-type pass rates. Scores cover only the tasks selected by category, subject, or sample_ids.
See Metrics and Aggregation for repeated attempts, category aggregation, and scoring failure rules.
Task Results and Scoring Evidence
Each attempt’s scoring record is stored inscoring under meta.benchmark. Successful grading provides the following fields for each task type:
FrontierScience-Olympiad:
FrontierScience-Research:
On either grading path, inspect
error for judge-call or response-parsing failures; parsing failures may also preserve a truncated raw_response.