vtllms/sealqa dataset revision.
Use seal_0 and seal_hard to evaluate a search-enabled Harness. Use longseal to evaluate long-context evidence synthesis from documents supplied directly in the prompt.
How it works
Inference and judging
Forseal_0 and seal_hard, the Benchmark sends each question to the configured Harness. The recommended naive_search_agent Harness can search the web before producing its answer.
For longseal, the Benchmark builds a prompt containing the question and a deterministic selection of evidence documents. Pair it with openai_chat to measure long-context reasoning without adding another search step.
After inference, the judge model (judge_model) receives the question, ground truth, and answer under test, then uses the official SealQA judge prompt to assign one of three verdicts. The judge and the model under test are two separate endpoints; judge_model must be specified explicitly:
A: correctB: incorrectC: not attempted
A receives a score of 1; B and C receive 0. The SealQA paper uses gpt-4o-mini as the judge model and reports 98% agreement with human evaluation. AgentCompass uses the open-weight Qwen3.5-35B-A3B as the judge model for its evaluations.
Categories and task IDs
The pinned default dataset revision contains:seal_hard includes all seal_0 questions and adds harder questions. AgentCompass treats the categories as separate runs; selecting seal_hard does not also run seal_0.
LongSeal document construction
For eachlongseal task, AgentCompass reads the hard-negative documents from the dataset column selected by longseal_document_count and, when available, inserts one pseudo-randomly selected gold document. The selection and insertion position are deterministic for a given task and longseal_seed.
The resulting prompt normally contains the configured number of hard negatives plus one gold document. A dataset row can produce fewer documents when its source lists are shorter or no gold document is available. AgentCompass records the actual total as longseal_document_count and the one-based gold position as longseal_gold_position in the task result so you can audit the constructed context.
For reproducible LongSeal comparisons, keep dataset_revision, longseal_document_count, and longseal_seed fixed, and use a Harness that does not add external search.
Parameters
Pass a JSON object via--benchmark-params '{...}', or a benchmark.params block in the YAML given to --config; the CLI wins on shared keys. See the Benchmark overview for merge precedence.
Parameter reference
| Parameter | Type | Default | Choices / values | Description |
|---|---|---|---|---|
category | string | ”seal_0” | seal_0, seal_hard, longseal | Selects the dataset configuration. Hyphenated aliases such as seal-hard are also accepted. |
judge_model | dict | null | id, base_url, api_key, api_protocol, params | Judge model spec, required (see Judge model spec). It decides grading, and is not the CLI —model-*. |
dataset_revision | string | ”267b8197ae75680ee0db180c4c2e96bd4e1001b4” | Non-empty Hugging Face revision | Pins the remote dataset snapshot for reproducibility (see Dataset source and cache). |
longseal_document_count | integer | 12 | 12, 20, or 30 | Selects the number of LongSeal hard-negative documents. Following the official setting, AgentCompass adds one gold document. |
longseal_seed | integer | 0 | Any integer | Controls deterministic gold-document selection and placement for LongSeal tasks. |
sample_ids follow Benchmark Parameters. Configure repeated attempts with --k and --attempt-strategy; see Metrics and Aggregation.
Judge model spec
judge_model is passed as a dict with the fields id, base_url, api_key, api_protocol, and params, pointing to the judge model’s own endpoint, with inference parameters under params:
Qwen3.5-35B-A3B.
To specify inference parameters for the judge model, add a params object to its configuration.
Dataset source and cache
AgentCompass downloads the selected configuration from Hugging Face and caches it under<data_dir>/sealqa. The default dataset_revision pins commit 267b8197ae75680ee0db180c4c2e96bd4e1001b4; change it explicitly if you want newer upstream data. You can compare available revisions in the dataset’s commit history.
The upstream dataset is licensed under Apache-2.0. Review its dataset card before redistributing cached data.
Run examples
The three positional arguments toagentcompass run are Benchmark, Harness, and Model. This page uses sealqa: pair seal_0 and seal_hard with naive_search_agent for search and answering, and longseal with openai_chat to process the document context directly. All examples use host_process and the tested Model named by MODEL_NAME.
Set connection details for the tested Model and an independent judge. Search examples also need Serper and Jina credentials; LongSeal does not call search services.
--benchmark-params, and retrieval-tool settings in --harness-params. See the Run Parameter Reference for shared rules.
Recommended Harness
The three scenarios below start from the default search evaluation; the custom scenario also shows how to switch to LongSeal document evaluation.- Smoke test (single task end-to-end)
- Custom parameters
- AgentCompass recommended config
Run
seal_0-001 from the default category with the search Harness to verify inference and scoring.Other optional Harnesses
openai_chat evaluates LongSeal’s long-context capability. This command covers all 254 longseal tasks using the default 12 hard-negative documents and seed 0, without search-service credentials.
Evaluation Results
For shared result conventions, see Run Directory, Aggregate Scores, and Task Files and Shared Fields.Scoring Metrics
SealQA’s primary metric is binarycorrect: under the official judge protocol above, A maps to true and B/C map to false. There is no partial credit.
With the default configuration, the overall score is accuracy over selected tasks with valid scores, ranging from 0 to 1; higher is better, and 0.63 means 63%. Category breakdowns use the dataset’s topic, which differs from the category parameter used to select a dataset subset.
See Metrics and Aggregation for repeated attempts, category aggregation, and shared scoring failure rules.
Task Results and Scoring Evidence
After an attempt is scored successfully,scoring under meta.benchmark contains the following evidence, including the letter grade and raw judge response:
The same
meta.benchmark also stores dataset_category and dataset_revision. LongSeal adds longseal_document_count and longseal_gold_position to trace data provenance and document construction.