Skip to main content
FrontierScience (arxiv) evaluates an agent’s ability to perform expert-level scientific tasks: given a scientific question that requires research and reasoning, the agent researches and produces a final answer, which an LLM judge ** then grades against the reference. The benchmark spans two task types — ** FrontierScience-Olympiad ** (short-answer problems) and ** FrontierScience-Research (open-ended research questions) — and each is graded by its own rule. A run may mix both types; both rules produce the same binary correct observation. FrontierScience uses single-sided judging. The judge only assesses the agent-under-test’s answer against the reference, without comparing to any baseline. Both inference and judging run in the local process (host_process) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.

How it works

A FrontierScience run has two stages — inference and judging — where the judging stage applies the grading rule that matches each task’s type.

Inference and judging

  • Inference. The model under test acts as a search agent and, driven by the harness (default naive_search_agent), completes multi-turn tool loops such as search / visit per task, producing a natural-language answer.
  • Judging. The judge model (judge_model) receives the question, the reference (a short reference answer or a scoring rubric, depending on the task type), and the answer under test, then grades it. The judge and the model under test are two separate endpoints; judge_model must be specified explicitly.

How the two task types are graded

The grading rule is not selected by a run parameter — it is determined by the task itself. A task whose category is research is graded as **FrontierScience-Research ; otherwise (olympiad) it is graded as ** FrontierScience-Olympiad. Because this is decided per task, a single run can contain both.
  • FrontierScience-Olympiad — short-answer grading. The reference is one short answer: a number, a symbolic expression, or a short phrase. The judge compares the candidate’s final answer to it and:
    • accepts mathematically equivalent expressions and harmless formatting differences;
    • accepts minor wording differences that preserve the same scientific meaning;
    • grades incorrect ** if the candidate states ** multiple conflicting final answers;
    • judges only what the candidate actually wrote, without supplying missing steps on its behalf.
    The verdict is a boolean correct.
  • FrontierScience-Research — rubric grading. The reference is a multi-item scoring rubric worth 10 points in total. The judge scores the answer ** item by item , awarding partial credit per item (each capped at that item’s max points), then sums the awarded points into a total on a 0–10 scale. Both final conclusions and intermediate reasoning steps can earn points, but only what the answer actually supports is credited — unstated work earns nothing. The task is ** correct ** when the total score is ** at least the pass threshold research_pass_threshold (default 7.0).
An empty model answer is graded incorrect outright (a research task also gets total score 0). If the judge returns malformed output, the scorer retries once with a stricter formatting instruction; if it still cannot be parsed, evaluation fails; see Evaluation Results for the recorded failure evidence.

Parameters

Pass a JSON object via --benchmark-params '{...}', or a benchmark.params block in the YAML given to --config; the CLI wins on shared keys. See the Benchmark overview for merge precedence.

Parameter reference

ParameterTypeDefaultChoices / valuesDescription
judge_modeldictnullid, base_url, api_key, api_protocol, paramsJudge model spec, required (see Judge model spec). It decides grading, and is not the CLI —model-*.
categorystring / list”all""all”, olympiad, researchFilter tasks by category; “all” = no filter, a list takes the union. The two categories are the benchmark’s task types — olympiad (100 tasks) and research (60), 160 in total — so filtering by category also determines which grading rule the run uses.
subjectstring”all”all / physics / chemistry / biologyFilter by scientific subject — a single value, not a list. all = no filter. Must be non-empty.
research_pass_thresholdfloat7.00.0–10.0Pass mark for FrontierScience-Research tasks, on the rubric’s 0–10 scale: such a task is correct when its total rubric score is ≥ this value. Raise it to be stricter, lower it to be more lenient. Has no effect on olympiad short-answer tasks.
Shared Benchmark fields such as sample_ids follow Benchmark Parameters. Configure repeated attempts with --k and --attempt-strategy; see Metrics and Aggregation.

Judge model spec

judge_model is passed as a dict with the fields id, base_url, api_key, api_protocol, and params, pointing to the judge model’s own endpoint, with inference parameters under params. We recommend fixing a single judge across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. AgentCompass recommends Qwen3.6-35B-A3B. Note that research-rubric grading is more nuanced than short-answer grading — it involves item-by-item scoring with partial credit — so a stronger, more capable judge improves rubric reliability.

Run examples

The three positional arguments to agentcompass run are Benchmark, Harness, and Model. These examples use frontierscience, naive_search_agent, and $MODEL_NAME, with host_process as the Environment. Set the following environment variables in your terminal before running the examples:
  • Model under test: MODEL_NAME, MODEL_BASE_URL, and MODEL_API_KEY; see Model connection details.
  • Judge Model: JUDGE_MODEL_NAME, JUDGE_MODEL_BASE_URL, and JUDGE_MODEL_API_KEY. Use a fixed, independent judge configuration.
  • Search tools: SERPER_API_KEY for search and JINA_API_KEY for visit.
See the run command for configuration ownership and CLI overrides.
Use sample_ids to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.

Evaluation Results

For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.

Scoring Metrics

FrontierScience’s primary metric is binary correct, determined by the task-type grading rules: Olympiad tasks use the judge’s boolean verdict, while Research tasks pass when their rubric total reaches research_pass_threshold. Research item scores and total_score explain the threshold decision; they are not separate aggregate metrics. With the default configuration, each task has one attempt and the overall score is the pass rate over tasks with valid scores, ranging from 0 to 1; higher is better. When both task types are included, each task has equal weight; the overall score is not the mean of the two task-type pass rates. Scores cover only the tasks selected by category, subject, or sample_ids. See Metrics and Aggregation for repeated attempts, category aggregation, and scoring failure rules.

Task Results and Scoring Evidence

Each attempt’s scoring record is stored in scoring under meta.benchmark. Successful grading provides the following fields for each task type: FrontierScience-Olympiad: FrontierScience-Research: On either grading path, inspect error for judge-call or response-parsing failures; parsing failures may also preserve a truncated raw_response.