> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# NaiveSearchAgent

The `naive_search_agent` harness runs the AgentCompass built-in **deep-search agent**, having the model under test complete research benchmarks such as [GAIA](/en/user_guide/modules/benchmarks/gaia), [SealQA](/en/user_guide/modules/benchmarks/sealqa), [DeepSearchQA](/en/user_guide/modules/benchmarks/deepsearchqa), [FrontierScience](/en/user_guide/modules/benchmarks/frontierscience), and [WideSearch](/en/user_guide/modules/benchmarks/widesearch) with one agent or a coordinator and parallel child agents.

The Harness supports only `host_process` and runs the agent and its tools directly in the AgentCompass Python process. It drives multi-turn retrieval and collects the final answer and search trajectory. Provide Model credentials through `--model-base-url` / `--model-api-key`, plus the API keys for the selected web tools. Both `openai-chat` and `openai-responses` are supported through `--model-api-protocol`.

## How it works

* **Tools and loop.** `tools` selects the enabled web tools (`search` / `browse` / `visit`); the engine interacts with the model over multiple turns using the function-calling protocol. `max_iterations` caps iterations for the single agent or coordinator, and `sub_agent_max_iterations` caps each child's iterations. `max_tool_calls_per_turn` caps tool calls in a single assistant message, and `max_tool_response_length` truncates an over-long web-tool response (keeping the head and tail). When the model stops issuing tool calls, the answer is considered complete and the content of the last assistant message is taken as the final answer.
* **Agent modes.** `mode: single` is the default. `mode: multi` adds `create_sub_agents` to the coordinator, which delegates research and combines child results into a final answer. Children use the same Model configuration and selected web tools, with independent conversation contexts and no recursive delegation. Each delegation call runs up to four children concurrently by default, without a cumulative child-count limit. When the task has an execution deadline, children stop earlier to leave time for the coordinator's final answer.
* **External services.** `search` depends on [Serper](https://serper.dev), and `browse` / `visit` depend on [Jina Reader](https://jina.ai/reader); keys are supplied via `serper_api_key` / `jina_api_key` (defaulting to the same-named environment variables). `tool_model_name` can set a dedicated web-summary model for `visit`, falling back to the model under test when left empty.

Multi-agent delegation is adapted from [WideSearch](https://github.com/ByteDance-Seed/WideSearch/blob/main/src/agent/multi_agent_tools.py). The coordinator adds complete delegation results and time for final synthesis to NaiveSearchAgent's loop and retry behavior.

## Built-in tools

The agent can call the following three web tools during the search loop. Use the `tools` parameter to choose which to enable (default `["search", "visit"]`); they can be combined as needed.

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'900px', width:'100%'}}>
    <colgroup>
      <col width="12%" />

      <col width="25%" />

      <col width="40%" />

      <col width="23%" />
    </colgroup>

    <thead>
      <tr><th style={{whiteSpace:'nowrap'}}>Tool</th><th>Inputs</th><th>Purpose</th><th>Dependencies</th></tr>
    </thead>

    <tbody>
      <tr>
        <td style={{whiteSpace:'nowrap'}}><code>search</code></td>
        <td><code>query</code> (search terms)</td>
        <td>Runs a single Google search and returns a result list (titles, snippets, links, etc.). Used to discover pages relevant to the question — the entry point of retrieval.</td>
        <td>Serper</td>
      </tr>

      <tr>
        <td style={{whiteSpace:'nowrap'}}><code>visit</code></td>
        <td><code>url</code> (a single link or an array of links), <code>goal</code> (what this visit aims to obtain)</td>
        <td>Fetches one or more pages and returns a **summary** of the content focused on <code>goal</code> (rather than the full text). The summary is generated by the model set in <code>tool\_model\_name</code>, falling back to the model under test. Suited for targeted extraction from long pages.</td>
        <td>Jina Reader + summary model</td>
      </tr>

      <tr>
        <td style={{whiteSpace:'nowrap'}}><code>browse</code></td>
        <td><code>url</code> (a single link)</td>
        <td>Fetches the **full content** of a single page (title, summary, body) and returns it verbatim, without LLM summarization. Suited for cases that need to preserve the page's original detail.</td>
        <td>Jina Reader</td>
      </tr>
    </tbody>
  </table>
</div>

The default combination `search` + `visit` matches the typical deep-search flow: use `search` to find candidate pages, then use `visit` with an explicit `goal` to read closely and extract information. When you need the page's original text rather than a summary (for example, comparing tables, code, or clauses verbatim), switch to or add `browse`. The difference between `visit` and `browse` is that the former returns a **goal-oriented summary** while the latter returns the **full text**.

`create_sub_agents` is a separate Model-callable delegation tool, automatically enabled for the coordinator by `mode: multi`; do not add it to `tools`. Its batch response preserves every child's result and is exempt from `max_tool_response_length`. Children inherit only the selected web tools. The Model calls tools as needed, with no fixed execution order.

## Parameters

Pass a JSON object via `--harness-params '{...}'`, or a `harnesses.naive_search_agent` block in the YAML given to `--config`; the CLI wins on shared keys (deep-merge).

### Parameter reference

<div style={{overflowX:'auto'}}>
  <table style={{minWidth:'1160px', width:'100%'}}>
    <colgroup>
      <col width="23%" />

      <col width="14%" />

      <col width="19%" />

      <col width="14%" />

      <col width="30%" />
    </colgroup>

    <thead>
      <tr><th style={{whiteSpace:'nowrap'}}>Parameter</th><th style={{whiteSpace:'nowrap'}}>Type</th><th style={{whiteSpace:'nowrap'}}>Default</th><th>Choices / values</th><th>Description</th></tr>
    </thead>

    <tbody>
      <tr><td style={{whiteSpace:'nowrap'}}><code>tools</code></td><td>list</td><td><code>\["search", "visit"]</code></td><td><code>search</code> / <code>browse</code> / <code>visit</code></td><td>Enabled web tool list.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>max\_iterations</code></td><td>int</td><td><code>50</code></td><td>≥ 1</td><td>Maximum iterations per agent.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>max\_retry</code></td><td>int</td><td><code>10</code></td><td>≥ 1</td><td>Maximum attempts per call, including the initial attempt.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>retry\_interval</code></td><td>int</td><td><code>5</code></td><td>≥ 1</td><td>Base wait interval between retries, in seconds.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>max\_tool\_calls\_per\_turn</code></td><td>int</td><td><code>5</code></td><td>≥ 1</td><td>Maximum tool calls allowed in one assistant message.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>max\_tool\_response\_length</code></td><td>int</td><td><code>8192</code></td><td>≥ 1</td><td>Maximum printable units retained from a web-tool response (truncated beyond this, keeping head and tail). Does not truncate <code>create\_sub\_agents</code> batch results.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>request\_timeout</code></td><td>int</td><td><code>2000</code></td><td>≥ 1</td><td>Timeout for one Model HTTP request, in seconds.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>tool\_model\_name</code></td><td>string</td><td><code>""</code></td><td>—</td><td>Dedicated web-summary model for the <code>visit</code> tool; falls back to the model under test when empty.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>serper\_api\_key</code></td><td>string</td><td><code>{"${SERPER_API_KEY}"}</code></td><td>—</td><td>Serper search API key.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>jina\_api\_key</code></td><td>string</td><td><code>{"${JINA_API_KEY}"}</code></td><td>—</td><td>Jina Reader API key.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>mode</code></td><td>string</td><td><code>"single"</code></td><td><code>single</code> / <code>multi</code></td><td><code>single</code> runs one agent; <code>multi</code> enables <code>create\_sub\_agents</code> for the coordinator.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>sub\_agent\_max\_iterations</code></td><td>int</td><td><code>50</code></td><td>≥ 1</td><td>Maximum iterations for each child agent.</td></tr>
      <tr><td style={{whiteSpace:'nowrap'}}><code>sub\_agent\_concurrency</code></td><td>int</td><td><code>4</code></td><td>≥ 1</td><td>Applies only in <code>multi</code> mode. Limits concurrent children within each <code>create\_sub\_agents</code> call; additional children wait for a slot.</td></tr>
    </tbody>
  </table>
</div>

At the application level, `max_retry` applies to Model and tool calls; `retry_interval` is the fixed wait between retries. Page-summary Model calls reuse both settings; tools have separate internal HTTP retry rules.

`sub_agent_max_iterations` and `sub_agent_concurrency` apply in `multi` mode. Concurrent delegation calls each have their own concurrency limit; task-level parallelism through `--task-concurrency` adds to the total work.

When a multi-agent task has a total time limit, the coordinator reserves the smaller of `request_timeout` and 10% of that limit for a final answer without further tools, and reserves one turn within `max_iterations` for synthesis. Child research stops before this reserve begins. A delegation batch returns completed results and any available partial results, with an error status for children that timed out. When the research time or turn budget is exhausted, the coordinator synthesizes the collected evidence. This behavior needs no additional configuration and does not change `single` mode.

### Search and parsing API keys

`serper_api_key` / `jina_api_key` default to environment-variable references (`${SERPER_API_KEY}` / `${JINA_API_KEY}`): set the same-named variables in your shell and they are injected automatically, or pass the keys inline in `--harness-params`. When only `search` is enabled you can omit the Jina key; when only `visit` / `browse` are enabled you can omit the Serper key — just provide the key matching the tools actually enabled.

## Run examples

Pass `naive_search_agent` as the second positional argument in this command:

```bash theme={"system"}
agentcompass run <benchmark> naive_search_agent <model>
```

Harness configuration is passed via `--harness-params`. GAIA, DeepSearchQA, and similar benchmarks are all judge-scored and require a judge model `judge_model` via `--benchmark-params`, otherwise tasks cannot be scored (see the respective benchmark docs).

When `mode` is not configured, it defaults to `single`; set it to `multi` to test child-agent delegation. The WideSearch recommended configuration in Quick Start explicitly uses `multi`, matching the Multi-agent example below.

<Tabs>
  <Tab title="Default">
    Pass the Serper / Jina keys directly via `--harness-params`; the default mode is `single`.

    ```bash theme={"system"}
    agentcompass run \
      gaia \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {"id": "Qwen3.6-35B-A3B", "base_url": "https://your-judge-endpoint/v1", "api_key": "sk-…"}
      }' \
      --harness-params '{
        "serper_api_key": "your-serper-key",
        "jina_api_key": "your-jina-key"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat
    ```
  </Tab>

  <Tab title="Custom params">
    Narrow the toolset and iterations, pass keys inline, and set a task wall-clock timeout.

    ```bash theme={"system"}
    agentcompass run \
      gaia \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {"id": "Qwen3.6-35B-A3B", "base_url": "https://your-judge-endpoint/v1", "api_key": "sk-…"}
      }' \
      --harness-params '{"tools": ["search", "visit"], "max_iterations": 40, "serper_api_key": "your-serper-key", "jina_api_key": "your-jina-key"}' \
      --execution-params '{"run_timeout_seconds": 9600}' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 16
    ```
  </Tab>

  <Tab title="Multi-agent">
    Enable parallel delegation for WideSearch while retaining the default `search` and `visit` tools. Install the optional dependencies described on the [WideSearch page](/en/user_guide/modules/benchmarks/widesearch#run-examples) first.

    `sub_agent_concurrency` limits concurrent children within each `create_sub_agents` call and defaults to `4`.

    ```bash theme={"system"}
    agentcompass run \
      widesearch \
      naive_search_agent \
      "$MODEL_NAME" \
      --env host_process \
      --benchmark-params '{
        "judge_model": {"id": "your-judge-model", "base_url": "https://your-judge-endpoint/v1", "api_key": "your-judge-key", "api_protocol": "openai-chat"}
      }' \
      --harness-params '{
        "mode": "multi",
        "sub_agent_concurrency": 4,
        "serper_api_key": "your-serper-key",
        "jina_api_key": "your-jina-key"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 1
    ```
  </Tab>
</Tabs>

## Output

The Harness returns a `RunResult` per task: the final answer (`final_answer`), the trajectory, the execution status, and telemetry such as the iteration count and engine status. Execution failures are reported as [`issues`](/en/user_guide/using_agentcompass/run_controls#error-handling-and-score-validity): model authentication, quota, or server failures and search or Jina service failures (including a rejected Jina credential or quota) are FATAL and rerun the attempt within `execution.max_retries`; model timeouts or no effective output are ERROR. In `multi` mode, an unresolved FATAL issue from a sub-agent is raised to the task result; sub-agent ERROR outcomes such as timeouts stay in that sub-agent's record, and the coordinator can still answer from partial results. Per-task details and aggregate metrics are written by the Benchmark under the [run directory](/en/user_guide/other_features/results/overview#directory-layout) (see [Results](/en/user_guide/other_features/results/overview)).

In `multi` mode, the main trajectory contains only the coordinator. Under `artifacts`, `sub_agents` preserves each child's complete response, messages, ACTF trajectory, and usage, while `sub_agent_runtime` records child execution counters. Partial child results remain available after failure or cancellation. The telemetry fields `agent_prompt_tokens` and `agent_completion_tokens` cover the coordinator and child agent loops; they exclude Model calls made inside `visit` for page summaries.

Set the task execution deadline through `--execution-params` with `run_timeout_seconds` and `run_timeout_multiplier`. See [phase timeouts](/en/user_guide/using_agentcompass/run_controls#set-an-appropriate-timeout).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.