> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# FrontierScience

FrontierScience ([arxiv](https://arxiv.org/abs/2601.21165)) evaluates an agent's ability to perform expert-level scientific tasks: given a scientific question that requires research and reasoning, the agent researches and produces a final answer, which an **LLM judge \*\* then grades against the reference. The benchmark spans two task types — \*\* FrontierScience-Olympiad \*\* (short-answer problems) and \*\* FrontierScience-Research** (open-ended research questions) — and each is graded by its own rule. A run may mix both types; both rules produce the same binary `correct` observation.

FrontierScience uses single-sided judging. The judge only assesses the agent-under-test's answer against the reference, without comparing to any baseline. Both inference and judging run in the local process (`host_process`) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.

## How it works

A FrontierScience run has two stages — inference and judging — where the judging stage applies the grading rule that matches each task's type.

### Inference and judging

* **Inference.** The model under test acts as a search agent and, driven by the harness (default [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent)), completes multi-turn tool loops such as search / visit per task, producing a natural-language answer.
* **Judging.** The judge model (`judge_model`) receives the question, the reference (a short reference answer or a scoring rubric, depending on the task type), and the answer under test, then grades it. The judge and the model under test are two separate endpoints; `judge_model` must be specified explicitly.

### How the two task types are graded

The grading rule is not selected by a run parameter — it is determined by the task itself. A task whose `category` is `research` is graded as \*\*FrontierScience-Research **; otherwise (`olympiad`) it is graded as \*\* FrontierScience-Olympiad**. Because this is decided per task, a single run can contain both.

* **FrontierScience-Olympiad — short-answer grading.** The reference is one short answer: a number, a symbolic expression, or a short phrase. The judge compares the candidate's final answer to it and:

  * accepts mathematically equivalent expressions and harmless formatting differences;
  * accepts minor wording differences that preserve the same scientific meaning;
  * grades **incorrect \*\* if the candidate states \*\* multiple conflicting final answers**;
  * judges only what the candidate actually wrote, without supplying missing steps on its behalf.

  The verdict is a boolean `correct`.

* **FrontierScience-Research — rubric grading.** The reference is a multi-item scoring rubric worth 10 points in total. The judge scores the answer \*\* item by item **, awarding partial credit per item (each capped at that item's max points), then sums the awarded points into a total on a 0–10 scale. Both final conclusions and intermediate reasoning steps can earn points, but only what the answer actually supports is credited — unstated work earns nothing. The task is \*\* correct \*\* when the total score is \*\* at least** the pass threshold `research_pass_threshold` (default `7.0`).

An empty model answer is graded incorrect outright (a research task also gets total score 0). If the judge returns malformed output, the scorer retries once with a stricter formatting instruction; if it still cannot be parsed, evaluation fails; see [Evaluation Results](#outputs) for the recorded failure evidence.

## Parameters

Pass a JSON object via `--benchmark-params '{...}'`, or a `benchmark.params` block in the YAML given to `--config`; the CLI wins on shared keys. See the [Benchmark overview](/en/user_guide/modules/benchmarks/overview) for merge precedence.

### Parameter reference

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'1040px', width:'100%'}}>
    <colgroup>
      <col width="18%" />

      <col width="16%" />

      <col width="15%" />

      <col width="20%" />

      <col width="31%" />
    </colgroup>

    <thead>
      <tr><th style={{whiteSpace:'nowrap'}}>Parameter</th><th style={{whiteSpace:'nowrap'}}>Type</th><th style={{whiteSpace:'nowrap'}}>Default</th><th>Choices / values</th><th>Description</th></tr>
    </thead>

    <tbody>
      <tr><td style={{whiteSpace:'nowrap'}}><code>judge\_model</code></td><td style={{whiteSpace:'nowrap'}}>dict</td><td style={{whiteSpace:'nowrap'}}><code>null</code></td><td><code>id</code>, <code>base\_url</code>, <code>api\_key</code>, <code>api\_protocol</code>, <code>params</code></td><td>Judge model spec, <strong>required</strong> (see <a href="#judge-model-spec">Judge model spec</a>). It decides grading, and is not the CLI <code>--model-\*</code>.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>category</code></td><td style={{whiteSpace:'nowrap'}}>string / list</td><td style={{whiteSpace:'nowrap'}}><code>"all"</code></td><td><code>"all"</code>, <code>olympiad</code>, <code>research</code></td><td>Filter tasks by category; <code>"all"</code> = no filter, a list takes the union. The two categories are the benchmark's task types — <code>olympiad</code> (100 tasks) and <code>research</code> (60), 160 in total — so filtering by <code>category</code> also determines which grading rule the run uses.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>subject</code></td><td style={{whiteSpace:'nowrap'}}>string</td><td style={{whiteSpace:'nowrap'}}><code>"all"</code></td><td><code>all</code> / <code>physics</code> / <code>chemistry</code> / <code>biology</code></td><td>Filter by scientific subject — a <strong>single value</strong>, not a list. <code>all</code> = no filter. Must be non-empty.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>research\_pass\_threshold</code></td><td style={{whiteSpace:'nowrap'}}>float</td><td style={{whiteSpace:'nowrap'}}><code>7.0</code></td><td><code>0.0</code>–<code>10.0</code></td><td>Pass mark for <strong>FrontierScience-Research</strong> tasks, on the rubric's 0–10 scale: such a task is correct when its total rubric score is ≥ this value. Raise it to be stricter, lower it to be more lenient. Has no effect on olympiad short-answer tasks.</td></tr>
    </tbody>
  </table>
</div>

Shared Benchmark fields such as `sample_ids` follow [Benchmark Parameters](/en/user_guide/modules/benchmarks/overview). Configure repeated attempts with `--k` and `--attempt-strategy`; see [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation).

<a id="judge-model-spec" />

### Judge model spec

`judge_model` is passed as a dict with the fields `id`, `base_url`, `api_key`, `api_protocol`, and `params`, pointing to the judge model's own endpoint, with inference parameters under `params`.

We recommend **fixing a single judge** across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. AgentCompass recommends `Qwen3.6-35B-A3B`. Note that research-rubric grading is more nuanced than short-answer grading — it involves item-by-item scoring with partial credit — so a stronger, more capable judge improves rubric reliability.

## Run examples

The three positional arguments to `agentcompass run` are Benchmark, Harness, and Model. These examples use `frontierscience`, [`naive_search_agent`](/en/user_guide/modules/harnesses/naive_search_agent), and `$MODEL_NAME`, with [`host_process`](/en/user_guide/modules/environments/providers/host_process) as the Environment.

Set the following environment variables in your terminal before running the examples:

* Model under test: `MODEL_NAME`, `MODEL_BASE_URL`, and `MODEL_API_KEY`; see [Model connection details](/en/user_guide/modules/models/overview#configure-connection-details).
* Judge Model: `JUDGE_MODEL_NAME`, `JUDGE_MODEL_BASE_URL`, and `JUDGE_MODEL_API_KEY`. Use a fixed, independent judge configuration.
* Search tools: `SERPER_API_KEY` for `search` and `JINA_API_KEY` for `visit`.

See the [run command](/en/user_guide/using_agentcompass/cli/run) for configuration ownership and CLI overrides.

<Tabs>
  <Tab title="Smoke test (single task end-to-end)">
    Use `sample_ids` to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.

    ```bash wrap theme={"system"}
    agentcompass run \
      frontierscience \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        },
        "sample_ids": ["frontierscience_research_0000"]
      }' \
      --harness-params '{
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat
    ```
  </Tab>

  <Tab title="Custom parameters">
    Evaluate only a specific subject and raise the research pass threshold to a stricter value; also demonstrates lowering the iteration limit in `--harness-params`.

    ```bash wrap theme={"system"}
    agentcompass run \
      frontierscience \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        },
        "subject": "physics",
        "research_pass_threshold": 8
      }' \
      --harness-params '{
        "max_iterations": 40,
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>

  <Tab title="AgentCompass recommended config">
    Evaluate the full task set. `--benchmark-params` only needs the judge model `judge_model`; use `--task-concurrency` to raise cross-task concurrency.

    ```bash wrap theme={"system"}
    agentcompass run \
      frontierscience \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {
          "id": "'"$JUDGE_MODEL_NAME"'",
          "base_url": "'"$JUDGE_MODEL_BASE_URL"'",
          "api_key": "'"$JUDGE_MODEL_API_KEY"'"
        }
      }' \
      --harness-params '{
        "serper_api_key": "${SERPER_API_KEY}",
        "jina_api_key": "${JINA_API_KEY}"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>
</Tabs>

<a id="outputs" />

## Evaluation Results

For shared result conventions, see [Run Directory](/en/user_guide/other_features/results/overview#directory-layout), [Aggregate Scores](/en/user_guide/other_features/results/summary_analysis), and [Task Files and Shared Fields](/en/user_guide/other_features/results/task_results).

<a id="metric-contract-and-aggregate-series" />

### Scoring Metrics

FrontierScience's primary metric is binary `correct`, determined by the [task-type grading rules](#how-the-two-task-types-are-graded): Olympiad tasks use the judge's boolean verdict, while Research tasks pass when their rubric total reaches `research_pass_threshold`. Research item scores and `total_score` explain the threshold decision; they are not separate aggregate metrics.

With the default configuration, each task has one attempt and the overall score is the pass rate over tasks with valid scores, ranging from 0 to 1; higher is better. When both task types are included, each task has equal weight; the overall score is not the mean of the two task-type pass rates. Scores cover only the tasks selected by `category`, `subject`, or `sample_ids`.

See [Metrics and Aggregation](/en/user_guide/other_features/results/metrics_aggregation) for repeated attempts, category aggregation, and scoring failure rules.

<a id="per-task-details-details" />

### Task Results and Scoring Evidence

Each attempt's scoring record is stored in `scoring` under `meta.benchmark`. Successful grading provides the following fields for each task type:

**FrontierScience-Olympiad**:

| Field | Contents |
| - | - |
| `evaluation_type` | Fixed as `frontierscience_olympiad_judge`. |
| `correct` | The judge's boolean verdict, matching `metrics.correct`. |
| `reason` | The judge's grading rationale, or `empty_model_response` for an empty answer. |

**FrontierScience-Research**:

| Field | Contents |
| - | - |
| `evaluation_type` | Fixed as `frontierscience_research_rubric`. |
| `correct` | After successful grading, whether `total_score` reaches `passing_threshold`. An empty answer is always `false`. |
| `total_score` | The sum of awarded points across rubric items; 0 for an empty answer. |
| `passing_threshold` | The `research_pass_threshold` used for this attempt. |
| `rubric_items` | Per-item scores with `item`, `max_points`, `awarded_points`, and `reason`; an empty list for an empty answer. |
| `summary` | The judge's grading summary, or `empty_model_response` for an empty answer. |

On either grading path, inspect `error` for judge-call or response-parsing failures; parsing failures may also preserve a truncated `raw_response`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.