Skip to main content
SGI Deep Research (arxiv) — “Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows”; AgentCompass uses the SGI Deep Research subset. It evaluates a deep-research agent on scientific-research questions that require multi-step web retrieval and evidence gathering to produce a detailed research answer, which an LLM judge then grades as correct or not against the ground truth. Like BrowseComp, SGI Deep Research uses single-sided judging. The judge only compares the agent-under-test’s answer against the ground truth, without comparing to any baseline. Both inference and judging run in the local process (host_process) — the harness first drives the model under test through the search loop to produce a final answer, then the judge model grades it.

How it works

An SGI Deep Research run has two stages — inference and judging.

Inference and judging

  • Inference. The model under test acts as a research agent and, driven by the harness (default naive_search_agent), completes multi-turn tool loops such as search / visit per task, gathering evidence across sources to produce a detailed natural-language research answer.
  • Judging. The judge model (judge_model) receives “question + ground truth + answer under test” and grades it with the built-in A/B/C protocol. The judge compares only the final answer, ignoring reasoning and formatting differences; equivalent expressions are accepted. The judge and the model under test are two separate endpoints; judge_model must be specified explicitly.

The A/B/C verdict

The judge returns exactly one verdict, and only A counts as correct:
  • A — CORRECT: the answer semantically matches the ground truth (equivalent expressions and formatting allowed).
  • B — INCORRECT: any deviation from the ground truth.
  • C — INCOMPLETE / REPETITIVE / REFUSAL: an invalid answer (cut off mid-sentence, looping repetition, or an explicit refusal).

Parameters

Pass a JSON object via --benchmark-params '{...}', or a benchmark.params block in the YAML given to --config; the CLI wins on shared keys. See the Benchmark overview for merge precedence.

Parameter reference

ParameterTypeDefaultChoices / valuesDescription
judge_modeldictnullid, base_url, api_key, api_protocol, paramsJudge model spec, required (see Judge model spec). It decides grading, and is not the CLI —model-*.
categorystring / list”all""all”, life, earth, material, physics, mathematics, neuroscience, information, astronomy, chemistry, energyFilter tasks by category; “all” = no filter, a list takes the union. Task counts by category — life (87), earth (54), material (38), physics (32), mathematics (25), neuroscience (24), information (20), astronomy (17), chemistry (11), energy (10); 318 in total.
Shared Benchmark fields such as sample_ids follow Benchmark Parameters. Configure repeated attempts with --k and --attempt-strategy; see Metrics and Aggregation.

Judge model spec

judge_model is passed as a dict with the fields id, base_url, api_key, api_protocol, and params, pointing to the judge model’s own endpoint, with inference parameters under params. We recommend fixing a single judge across all models under test. Grading directly decides the scores, so switching judges makes scores no longer comparable across models; likewise, the model under test should not serve as its own judge, as that is neither fair nor comparable. The judge need not be especially strong — the A/B/C criterion (semantic match) is relatively objective, so a mid-sized model suffices. AgentCompass recommends Qwen3.6-35B-A3B.

Run examples

The three positional arguments to agentcompass run are Benchmark, Harness, and Model. These examples use sgi_deep_research, naive_search_agent, and $MODEL_NAME, with host_process as the Environment. Set the following environment variables in your terminal before running the examples:
  • Model under test: MODEL_NAME, MODEL_BASE_URL, and MODEL_API_KEY; see Model connection details.
  • Judge Model: JUDGE_MODEL_NAME, JUDGE_MODEL_BASE_URL, and JUDGE_MODEL_API_KEY. Use a fixed, independent judge configuration.
  • Search tools: SERPER_API_KEY for search and JINA_API_KEY for visit.
See the run command for configuration ownership and CLI overrides.
Use sample_ids to evaluate a single task, verifying that the end-to-end inference and judging flow works; defaults for the rest.

Evaluation Results

For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.

Scoring Metrics

SGI Deep Research’s primary metric is binary correct: under the A/B/C verdict described above, A maps to true and B/C map to false. There is no partial credit. With the default configuration, each task has one attempt and the overall score is accuracy over tasks with valid scores, ranging from 0 to 1; higher is better. For example, if all 100 tasks receive valid judgments and 63 are correct, 0.63 in the report means 63%. When you filter tasks with category or sample_ids, scores cover only the selected tasks. See Metrics and Aggregation for repeated attempts, category aggregation, and scoring failure rules.

Task Results and Scoring Evidence

After an attempt is scored successfully, scoring under meta.benchmark contains: This scoring record preserves the verdict and compared answers, but not the judge’s raw response or reasoning.