> ## Documentation Index
> Fetch the complete documentation index at: https://opencompass-docs-preview-pr-335-0.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenEvolve

`openevolve` harness 在准备好的程序演化任务上运行
[OpenEvolve](https://github.com/algorithmicsuperintelligence/openevolve)。它从可运行的初始程序出发，让被测模型不断提出
改进方案，使用任务冻结的 evaluator 评测每个候选程序，并提交搜索到的最佳程序。内置的
[Frontier Engineering](/zh/user_guide/modules/benchmarks/frontier_engineering) benchmark 提供了这套任务契约。

该 harness 支持 `openai-chat` model protocol。Model id 是命令的第三个位置参数，endpoint 与 API key 通过标准
`--model-*` 参数提供。

## 工作原理

* **校验任务契约**：准备后的任务必须提供 `agentcompass.program_evolution.v1` spec、初始程序、evaluator 文件和唯一的
  candidate output path。路径缺失或不一致时，harness 会在开始演化前报错。
* **准备 runner**：AgentCompass 把最小化的 runner 与 evaluator 源码上传到所选 environment。只要迭代次数大于
  0，该 environment 就必须提供准确版本的 `openevolve==0.2.26`。
* **演化并评测**：OpenEvolve 生成候选程序，并在每轮迭代后调用 benchmark 自有的 evaluator。`iterations` 控制
  演化预算，`max_code_length` 限制候选程序长度，`timeout` 默认将整个 harness task 的运行时间限制为 8 小时。
* **回收最佳程序**：harness 把 OpenEvolve 的紧凑演化历史转换为 AgentCompass trajectory，并在 `RunResult` 中返回
  最佳程序、对应指标和执行诊断信息。

## 参数

通过 `--harness-params '{...}'` 传入 JSON；也可以写入 `--config` 所指定 YAML 的 `harness.params`，同名项以命令行
为准。合并与优先级见 [Harness 概览](/zh/user_guide/modules/harnesses/overview)。

| 参数 | 类型 | 默认值 | 可选值 / 取值 | 说明 |
| - | - | - | - | - |
| `python` | string | `""` | Python 可执行文件路径 | `host_process` 下可选的 runner Python 覆盖值。空值使用 environment 或 recipe 默认值；其他 environment 不接受此字段。 |
| `iterations` | int | `100` | 整数 ≥ 0 | 演化迭代次数。设为 `0` 时不调用 LLM 演化，只评测并返回初始程序。 |
| `max_code_length` | int | `20000` | 整数 ≥ 1 | OpenEvolve 接受的候选程序最大长度。 |

## 兼容性与运行要求

### Execution environment

harness 不限制 environment id，但所选 environment 必须支持 POSIX 命令执行、可写 task workspace，并且能读取任务
资源。内置的 Frontier Engineering 集成可直接使用 `host_process`，并提供 Docker recipe。其他 environment 只有在
能访问准备好的程序演化路径并满足下述依赖时才能运行。

当 `iterations > 0` 时，实际运行 harness 的 Python 必须安装准确版本的 `openevolve==0.2.26`。依赖检查发生在所选
environment 内部，而不只是 host process 中：

* 使用 `host_process` 时，通过 `uv pip install -e ".[frontier-engineering]"` 安装项目 extra。如果 OpenEvolve 位于
  另一个解释器中，可用 `python` 参数指定它。
* 使用 Docker 或其他托管 environment 时，应选择已包含 `openevolve==0.2.26` 的 image 或 snapshot。只在 host
  Python 中安装 extra 不会使该依赖出现在目标 environment 内。

将 `iterations` 设为 `0` 会跳过 OpenEvolve 依赖检查，并把随任务提供的初始程序作为 baseline 进行评测。

### Model protocol 与凭据

只支持 `openai-chat`；compatibility validation 会拒绝 `openai-responses` 和 `anthropic`。harness 把
`--model-base-url`、`--model-api-key` 和位置参数中的 model id 作为 `OPENAI_API_BASE`、`OPENAI_API_KEY` 和
`OPENAI_MODEL` 传入所选 environment，由 OpenEvolve 的 OpenAI-compatible Chat Completions client 消费。
`iterations > 0` 时必须提供 API key，且 model endpoint 必须能从所选 environment 访问。Environment 网络策略见
[网络访问](/zh/user_guide/modules/environments/configuration/network)。

### Model 参数

Provider 请求配置通过 `--model-params` 传入，与 `--harness-params` 相互独立。Harness 会把 `temperature`、`top_p`、
`max_tokens`、`timeout`（或 `request_timeout`）、`retries`、`retry_delay`、`reasoning_effort` 和 `extra_body` 映射到
OpenEvolve 的 OpenAI-compatible client；其他请求行为使用 OpenEvolve 默认值。

### Workspace、超时与重试

harness 使用程序演化 spec，而不是通用的 prompt/tool loop。每条任务开始时都会重新创建
`<workspace>/.agentcompass/openevolve`，因此不会从上一次 attempt 的 OpenEvolve checkpoint 续跑。初始程序、
evaluator 命令、evaluator timeout 和最终验证归 benchmark 管理；演化循环与候选程序回收归 harness 管理。

`execution.run_timeout_seconds` 限制整个 harness task。单次模型请求的 `timeout` 或 `request_timeout` 应放在 `--model-params` 中；
evaluator timeout 则属于 benchmark。`--model-params` 中的 `retries` 和 `retry_delay` 控制 OpenEvolve model client
的重试，harness 本身不增加 task-level retry。AgentCompass 的任务重试通过通用
[运行控制](/zh/user_guide/using_agentcompass/run_controls#只重试瞬时失败)配置。

## 运行示例

<Tabs>
  <Tab title="冒烟测试">
    用一次演化迭代运行一条 Frontier Engineering 任务。

    ```bash theme={"system"}
    agentcompass run \
      frontier_engineering \
      openevolve \
      "$MODEL_NAME" \
      --env docker \
      --benchmark-params '{
        "task_set": "v1_lite",
        "sample_ids": ["InventoryOptimization/disruption_eoqd"]
      }' \
      --harness-params '{"iterations": 1}' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 1
    ```
  </Tab>

  <Tab title="自定义演化预算">
    执行更长的演化，并同时配置 harness 预算和单次模型请求参数。

    ```bash theme={"system"}
    agentcompass run \
      frontier_engineering \
      openevolve \
      "$MODEL_NAME" \
      --env docker \
      --benchmark-params '{
        "task_set": "v1_lite",
        "sample_ids": ["InventoryOptimization/disruption_eoqd"]
      }' \
      --harness-params '{"iterations": 100, "max_code_length": 30000}' \
      --execution-params '{"run_timeout_seconds": 14400}' \
      --model-params '{
        "max_tokens": 32768,
        "request_timeout": 1200,
        "reasoning_effort": "high"
      }' \
      --model-base-url "$MODEL_BASE_URL" \
      --model-api-key "$MODEL_API_KEY" \
      --model-api-protocol openai-chat \
      --task-concurrency 1
    ```
  </Tab>
</Tabs>

## 输出

harness 为每个任务返回一个 `RunResult`。`final_answer` 与 `file` artifact 包含最佳程序；metrics 包含 exit code、
是否超时、配置的迭代次数和最佳 evaluator 指标。标准化 trajectory 记录 OpenEvolve 紧凑历史中的程序及其 evaluator
observation；`openevolve` artifact 还保留最佳程序元数据、执行命令和 stdout/stderr 尾部。

Runner 非正常退出、达到 harness timeout，或没有产生最佳程序时，任务返回 `RUN_ERROR`。Benchmark 将单任务详情和
聚合指标写入 [运行目录](/zh/user_guide/other_features/results/overview#目录布局)，详见[结果](/zh/user_guide/other_features/results/overview)。

## 故障排查

| 现象 | 处理方式 |
| - | - |
| 依赖检查提示 OpenEvolve 缺失或版本不符 | 在所选 environment 的 Python 中安装准确版本的 `openevolve==0.2.26`。使用 `host_process` 时，检查 `python` 参数指定的解释器。 |
| Compatibility validation 拒绝 model protocol | 设置 `--model-api-protocol openai-chat`；该 harness 不支持 Responses 或 Anthropic protocol。 |
| Runner 提示缺少 API key 或无法访问 endpoint | 传入 `--model-api-key` 和 `--model-base-url`，并确认所选 environment 能访问该 endpoint。 |
| 任务因没有最佳程序而返回 `RUN_ERROR` | 查看 `openevolve` artifact 中的 stdout/stderr 尾部和最佳指标。只有诊断信息表明超时时，才增大 harness 或单次模型请求 timeout。 |

任务执行 deadline 统一通过 `--execution-params` 中的 `run_timeout_seconds` 和 `run_timeout_multiplier` 设置。详见[阶段超时](/zh/user_guide/using_agentcompass/run_controls#设置合适的超时)。


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.