Skip to main content
Choose, configure, verify, and troubleshoot baseline, run, and evaluation network policies. AgentCompass can control outbound network access separately for trusted environment setup, the complete untrusted run boundary, and formal verification. Use these controls to reproduce an official benchmark policy, prevent an agent from retrieving external solutions, or allow only the endpoints required by a controlled evaluation. Start with the policy documented by the selected benchmark. Changing network access can change both task difficulty and result comparability, so an alignment run should not silently broaden or narrow the official setting.

Choose a Policy for Each Phase

Network policy is resolved independently for every task. A Benchmark loader can declare the policy fields on each TaskSpec, so two samples in the same run can use different policies. Values passed through --env-params have run-wide scope and explicitly override the corresponding phase for every selected sample: These are not four successive phases. Shared evaluation has one environment baseline plus run/evaluation overrides. Fresh evaluation has two environments, each with its own baseline, and still only two execution-phase overrides. Harbor maps [environment].network_mode to the agent baseline and [verifier.environment].network_mode to the separate verifier baseline; [agent].network_mode and [verifier].network_mode remain the phase overrides. Request-level evaluation_baseline_network_policy requires fresh mode, and overrides a common request baseline before falling back to the task’s verifier baseline or the inherited agent baseline. An explicit empty Harbor verifier environment uses its own default public baseline. The existing phase fields are resolved with this precedence: Omitting a field means “continue to the next source”; if no source supplies that phase, it becomes public. Explicitly passing public and leaving a phase unset therefore produce the same effective policy when no other source supplies one. Use --env-params when intentionally forcing one policy across a run. To preserve official per-sample behavior, omit those overrides and let the Benchmark load its task policy. A compatible Recipe may validate execution-required endpoints, but it does not change the resolved policy; the effective policy and applied_recipes record the resulting decision.
Harness setup happens before run_network_policy is applied. This lets a trusted harness install its runtime under the baseline policy and then execute the untrusted agent under a stricter policy. Network phases are function-level trust boundaries, not a separate policy for every progress stage.

Understand Function-Level Boundaries

AgentCompass applies the same evaluation rule to both evaluation Environment modes: only the complete benchmark.evaluate() call uses evaluation_network_policy. The modes differ because fresh creates another Environment, while reuse evaluates in the agent Environment. The evaluate_environment progress phase in a fresh run is not a separate Benchmark hook; its setup boundary is the evaluation provider’s open() call. The tables below show the active policy for Environment operations. Provider control-plane requests made by the AgentCompass host remain outside sandbox network enforcement. A function can contain several internal operations; the runtime does not split those operations into finer policy scopes.

Shared Environment

With reuse, the same Environment remains active while its lifecycle functions move through the baseline, run, and evaluation policies: Immediately before step 7, the runtime switches directly from run_network_policy to evaluation_network_policy, calls the entire benchmark.evaluate() function, and restores baseline_network_policy in finally. This keeps the evaluation policy scoped to the evaluation function without an intermediate baseline transition.

Separate Environment

fresh completes the agent Environment lifecycle first, then creates a separate evaluation Environment: Step 8 is the conditional evaluation Environment setup introduced by selecting fresh; it is not an optional Benchmark hook. The runtime switches to the evaluation policy only after evaluation_provider.open() completes, restores the evaluation Environment baseline immediately after benchmark.evaluate(), and then releases or retains that Environment. Runtime executes declared collection commands directly, then collects existing files when storage is enabled. Without commands, it skips preparation. Commands and downloads remain inside the run trust boundary; they do not create another network phase.

Network Modes

Each phase accepts one of four modes: Use a string for public or no-network:
Use an object for an allowlist or denylist:
Host values must be hostnames, leading-wildcard hostnames, IP addresses, or canonical CIDR ranges. Do not include a URL scheme, path, whitespace, or an embedded wildcard such as api.*.example.com. Both allowlists and denylists accept a bare host such as www.example.com, a host:port shorthand such as www.example.com:443, or a target object such as {"host": "www.example.com", "port": 443}. Bare hosts retain all-port semantics; port values must be integers from 1 through 65535. Use [2001:db8::1]:443 when adding a port to an IPv6 address. Providers that cannot enforce a port-specific target reject it.

Tell the Agent About Rollout Restrictions

By default, each Harness appends an English network restriction statement to its final user instruction when the effective run_network_policy is no-network, allowlist, or denylist. This prevents an agent from treating an intentional restriction as a transient network failure and repeatedly retrying blocked operations. A public policy does not add a statement. The statement is based on the provider-resolved policy, not only the requested value. Recipe-added targets therefore appear in an allowlist statement, and allowlist or denylist targets are expanded from the effective policy. A provider that cannot enforce the requested policy rejects the run before rollout instead of injecting a misleading statement.
no-network uses the same wrapper without a target list:
For denylist, the effective denied targets replace the list:
The Harness injects this block only into the rollout copy of the prepared input. Artifact collection and benchmark.evaluate() continue to receive the original input. Standard Harness paths support plain prompts, structured user messages, multimodal user messages, and the JSON-encoded messages accepted by openai_chat. The harness-free TauBench inference path does not currently consume this notice. Disable injection through the selected Harness configuration:
For a persistent default, set inject_network_restriction_notice: false under the selected harnesses.<id> entry in a YAML or JSON configuration file. Python callers pass the same field in harness_params. This setting changes only the Harness input notice; it does not change or disable network-policy enforcement.

Pass Phase Policies from the CLI

Pass the three runtime policies as JSON fields in --env-params:
The runtime applies those CLI values as follows: The same values can be supplied as environment_params through the Python API, or by a Benchmark loader on TaskSpec for sample-level policy.

Select the Narrowest Practical Policy

Use this decision sequence:
  1. Check the benchmark page for an official or recommended policy.
  2. Identify where the harness is installed and where it calls the model API.
  3. Keep the baseline public if the sandbox must install a package or executable; otherwise prefer an allowlist or a prebuilt image.
  4. Set the run phase to no-network when the task should use only local evidence.
  5. Include any endpoints needed by Harness close or artifact collection in the run policy; do not broaden it between rollout and verification.
  6. Add only the exact model, search, judge, package, or artifact hosts required by a network-dependent phase.
  7. Run one task and inspect the resolved execution plan before scaling.
The Python packages used by the AgentCompass driver are installed outside the task sandbox and are not controlled by these policies. Packages or CLI tools installed by harness.start_session run inside the environment and therefore use the baseline policy. If the baseline must also be no-network, put those dependencies in the task image or snapshot first. Whether a model endpoint needs to be allowlisted depends on where the harness makes its request:
  • A local harness process calls the model from the AgentCompass host, outside the task environment policy.
  • A harness running inside the sandbox needs the model endpoint in the run-phase allowlist.
  • Some benchmark recipes, including DeepSWE recipes, infer the resolved model endpoint. Do not assume every custom benchmark or external recipe does so; inspect the resolved plan.
  • The meaning of run_network_policy is independent of where it came from. DeepSWE validates that its required model endpoints are permitted and fails during planning if the effective task or CLI policy blocks them; it never broadens that policy. A remote Harness therefore needs an explicit CLI allowlist containing the model host; a local Harness can keep no-network.
The same distinction applies to judge and search services. A request made by the AgentCompass driver is outside the sandbox policy; a request made by a process inside the task or verifier environment must be allowed in that phase.

Provider Support

Daytona accepts at most 20 domain entries or 10 IPv4 network entries. Docker cannot use dynamic phase policies with network set to none, host, or container:<id>. Its default egress proxy image is python:3.12-alpine; make sure the Docker daemon can pull it or pre-pull it on an offline host. Outbound policy accepts only the shared fields above. Provider-native block-all and allowlist parameters are rejected; adapters generate those API options from the resolved shared policy.
Docker is currently the only documented public provider that enforces denylist. Other public providers reject that policy instead of silently broadening it to public.

Verify the Effective Policy

Run one known task with persistent debug logs:
The run log records baseline_network_mode, run_network_mode, and evaluation_network_mode when each task execution plan is built. Per-task details retain the resolved policies and applied_recipes. Verify those effective values rather than relying only on the original command, because a Benchmark Recipe may add an inferred endpoint or provider adaptation. For an adversarial isolation test, ask the agent to access a known external URL and confirm both outcomes:
  • the request fails during the restricted run phase; and
  • the same environment can still perform the trusted setup work allowed by its baseline policy.
For supported terminal trajectories, NetworkOperationAnalyzer can summarize commands such as curl, wget, package installation, or git clone. It observes agent behavior but does not enforce the policy and cannot replace provider transition logs.
A failed application request is not sufficient evidence by itself. It may be caused by DNS, credentials, or an unavailable service. Confirm the resolved policy and provider transition logs as well.

Troubleshoot Network Failures