naive_search_agent harness runs the AgentCompass built-in deep-search agent, having the model under test complete research benchmarks such as GAIA, SealQA, DeepSearchQA, FrontierScience, and WideSearch with one agent or a coordinator and parallel child agents.
The Harness supports only host_process and runs the agent and its tools directly in the AgentCompass Python process. It drives multi-turn retrieval and collects the final answer and search trajectory. Provide Model credentials through --model-base-url / --model-api-key, plus the API keys for the selected web tools. Both openai-chat and openai-responses are supported through --model-api-protocol.
How it works
- Tools and loop.
toolsselects the enabled web tools (search/browse/visit); the engine interacts with the model over multiple turns using the function-calling protocol.max_iterationscaps iterations for the single agent or coordinator, andsub_agent_max_iterationscaps each child’s iterations.max_tool_calls_per_turncaps tool calls in a single assistant message, andmax_tool_response_lengthtruncates an over-long web-tool response (keeping the head and tail). When the model stops issuing tool calls, the answer is considered complete and the content of the last assistant message is taken as the final answer. - Agent modes.
mode: singleis the default.mode: multiaddscreate_sub_agentsto the coordinator, which delegates research and combines child results into a final answer. Children use the same Model configuration and selected web tools, with independent conversation contexts and no recursive delegation. Each delegation call runs up to four children concurrently by default, without a cumulative child-count limit. When the task has an execution deadline, children stop earlier to leave time for the coordinator’s final answer. - External services.
searchdepends on Serper, andbrowse/visitdepend on Jina Reader; keys are supplied viaserper_api_key/jina_api_key(defaulting to the same-named environment variables).tool_model_namecan set a dedicated web-summary model forvisit, falling back to the model under test when left empty.
Built-in tools
The agent can call the following three web tools during the search loop. Use thetools parameter to choose which to enable (default ["search", "visit"]); they can be combined as needed.
| Tool | Inputs | Purpose | Dependencies |
|---|---|---|---|
search | query (search terms) | Runs a single Google search and returns a result list (titles, snippets, links, etc.). Used to discover pages relevant to the question — the entry point of retrieval. | Serper |
visit | url (a single link or an array of links), goal (what this visit aims to obtain) | Fetches one or more pages and returns a summary of the content focused on goal (rather than the full text). The summary is generated by the model set in tool_model_name, falling back to the model under test. Suited for targeted extraction from long pages. | Jina Reader + summary model |
browse | url (a single link) | Fetches the full content of a single page (title, summary, body) and returns it verbatim, without LLM summarization. Suited for cases that need to preserve the page’s original detail. | Jina Reader |
search + visit matches the typical deep-search flow: use search to find candidate pages, then use visit with an explicit goal to read closely and extract information. When you need the page’s original text rather than a summary (for example, comparing tables, code, or clauses verbatim), switch to or add browse. The difference between visit and browse is that the former returns a goal-oriented summary while the latter returns the full text.
create_sub_agents is a separate Model-callable delegation tool, automatically enabled for the coordinator by mode: multi; do not add it to tools. Its batch response preserves every child’s result and is exempt from max_tool_response_length. Children inherit only the selected web tools. The Model calls tools as needed, with no fixed execution order.
Parameters
Pass a JSON object via--harness-params '{...}', or a harnesses.naive_search_agent block in the YAML given to --config; the CLI wins on shared keys (deep-merge).
Parameter reference
| Parameter | Type | Default | Choices / values | Description |
|---|---|---|---|---|
tools | list | [“search”, “visit”] | search / browse / visit | Enabled web tool list. |
max_iterations | int | 50 | ≥ 1 | Maximum iterations per agent. |
max_retry | int | 10 | ≥ 1 | Maximum attempts per call, including the initial attempt. |
retry_interval | int | 5 | ≥ 1 | Base wait interval between retries, in seconds. |
max_tool_calls_per_turn | int | 5 | ≥ 1 | Maximum tool calls allowed in one assistant message. |
max_tool_response_length | int | 8192 | ≥ 1 | Maximum printable units retained from a web-tool response (truncated beyond this, keeping head and tail). Does not truncate create_sub_agents batch results. |
request_timeout | int | 2000 | ≥ 1 | Timeout for one Model HTTP request, in seconds. |
tool_model_name | string | "" | — | Dedicated web-summary model for the visit tool; falls back to the model under test when empty. |
serper_api_key | string | — | Serper search API key. | |
jina_api_key | string | — | Jina Reader API key. | |
mode | string | ”single” | single / multi | single runs one agent; multi enables create_sub_agents for the coordinator. |
sub_agent_max_iterations | int | 50 | ≥ 1 | Maximum iterations for each child agent. |
sub_agent_concurrency | int | 4 | ≥ 1 | Applies only in multi mode. Limits concurrent children within each create_sub_agents call; additional children wait for a slot. |
max_retry applies to Model and tool calls; retry_interval is the fixed wait between retries. Page-summary Model calls reuse both settings; tools have separate internal HTTP retry rules.
sub_agent_max_iterations and sub_agent_concurrency apply in multi mode. Concurrent delegation calls each have their own concurrency limit; task-level parallelism through --task-concurrency adds to the total work.
When a multi-agent task has a total time limit, the coordinator reserves the smaller of request_timeout and 10% of that limit for a final answer without further tools, and reserves one turn within max_iterations for synthesis. Child research stops before this reserve begins. A delegation batch returns completed results and any available partial results, with an error status for children that timed out. When the research time or turn budget is exhausted, the coordinator synthesizes the collected evidence. This behavior needs no additional configuration and does not change single mode.
Search and parsing API keys
serper_api_key / jina_api_key default to environment-variable references (${SERPER_API_KEY} / ${JINA_API_KEY}): set the same-named variables in your shell and they are injected automatically, or pass the keys inline in --harness-params. When only search is enabled you can omit the Jina key; when only visit / browse are enabled you can omit the Serper key — just provide the key matching the tools actually enabled.
Run examples
Passnaive_search_agent as the second positional argument in this command:
--harness-params. GAIA, DeepSearchQA, and similar benchmarks are all judge-scored and require a judge model judge_model via --benchmark-params, otherwise tasks cannot be scored (see the respective benchmark docs).
When mode is not configured, it defaults to single; set it to multi to test child-agent delegation. The WideSearch recommended configuration in Quick Start explicitly uses multi, matching the Multi-agent example below.
- Default
- Custom params
- Multi-agent
Pass the Serper / Jina keys directly via
--harness-params; the default mode is single.Output
The Harness returns aRunResult per task: the final answer (final_answer), the trajectory, the execution status, and telemetry such as the iteration count and engine status. Execution failures are reported as issues: model authentication, quota, or server failures and search or Jina service failures (including a rejected Jina credential or quota) are FATAL and rerun the attempt within execution.max_retries; model timeouts or no effective output are ERROR. In multi mode, an unresolved FATAL issue from a sub-agent is raised to the task result; sub-agent ERROR outcomes such as timeouts stay in that sub-agent’s record, and the coordinator can still answer from partial results. Per-task details and aggregate metrics are written by the Benchmark under the run directory (see Results).
In multi mode, the main trajectory contains only the coordinator. Under artifacts, sub_agents preserves each child’s complete response, messages, ACTF trajectory, and usage, while sub_agent_runtime records child execution counters. Partial child results remain available after failure or cancellation. The telemetry fields agent_prompt_tokens and agent_completion_tokens cover the coordinator and child agent loops; they exclude Model calls made inside visit for page summaries.
Set the task execution deadline through --execution-params with run_timeout_seconds and run_timeout_multiplier. See phase timeouts.