Skip to content

Commit 80d769b

Browse files
committed
feat(examples): add reproducible eval+optimize closed-loop example
1 parent 8080800 commit 80d769b

12 files changed

Lines changed: 1833 additions & 0 deletions

File tree

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,4 @@
1+
# Runtime side-effects (regenerated on every run); the audited report and the
2+
# runs/latest prompt snapshots are kept in VCS as example deliverables.
3+
__pycache__/
4+
_sdk_eval_metrics.json
Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,71 @@
1+
# eval_optimize_loop
2+
3+
> Part of the `examples/optimization` series. Where [`quickstart`](../quickstart)
4+
> drives a live `AgentOptimizer` (GEPA) against a real agent, this example focuses
5+
> on the **closed loop around** optimization — baseline evaluation, failure
6+
> attribution, candidate validation, an acceptance gate, and an auditable report —
7+
> and runs fully reproducibly in fake/trace mode **without an API key**.
8+
9+
This example implements a reproducible Evaluation + Optimization closed loop:
10+
11+
```text
12+
baseline evaluation
13+
-> failure attribution
14+
-> prompt candidate generation
15+
-> validation regression
16+
-> acceptance gate
17+
-> auditable JSON/Markdown report
18+
```
19+
20+
Run without API keys:
21+
22+
```bash
23+
# First run may spend time on a one-off `uv sync`; the loop itself is ~seconds.
24+
uv run python examples/optimization/eval_optimize_loop/run.py
25+
```
26+
27+
Set `YUN_LOG_LEVEL=DEBUG` for more verbose logs (default `INFO`).
28+
29+
Inputs:
30+
31+
- `prompts/baseline_system.md` — target prompt being optimized.
32+
- `train.evalset.json` / `val.evalset.json` — SDK-clean evalsets (trace mode).
33+
- `case_meta.json` — per-case `key` / `rubric` / `tool_intent` (kept out of the
34+
evalset so `EvalSet` stays schema-clean).
35+
- `optimizer.json` — metric weights, scripted candidate patch, and gate thresholds.
36+
37+
Outputs:
38+
39+
- `optimization_report.json` / `optimization_report.md`
40+
- `runs/latest/baseline_prompt.md` / `runs/latest/candidate_prompt.md`
41+
42+
The sample has 6 cases:
43+
44+
- train: two optimizable failures and one optimization-ineffective format case.
45+
- validation: one new pass, one hard regression, and one soft degradation.
46+
47+
## 设计说明(四支柱)
48+
49+
**失败归因(阶段 2)。** 归因完全基于结构化评测信号,不依赖 case 命名。每条 case
50+
记录三项 metric 子分(final_response / tool_trajectory / rubric)与关键轨迹(query、
51+
expected/actual 工具与回复)。`classify_tool_failure` 据「期望轨迹 vs 实际轨迹」判类:
52+
期望调用权威检索工具却没调用 → `knowledge_recall_insufficient`;调全了期望工具又多调
53+
`spurious_tool_call`;首工具名对但参数不同 → `parameter_error`;否则 `tool_call_error`
54+
rubric 维度由 `case_meta.json` 显式声明(`json_format`/`no_tool`/`single_tool`),失败时
55+
映射为 `format_error``llm_rubric_not_met`。归因只统计 baseline 失败。
56+
57+
**接受策略(阶段 5)。** Gate 以验证集为先,五项可配置约束全过才 ACCEPT:① 验证集均分
58+
提升 ≥ `min_val_score_gain`;② 无新增 hard fail;③ 无「关键 case」退化(关键性由
59+
`case_meta.json``key=true` 标记,而非把所有验证 case 一概视为关键);④ 非过拟合
60+
(训练涨而验证跌);⑤ 优化成本 ≤ `max_cost_usd`。各检查相互独立,便于定位拒绝原因。
61+
62+
**防过拟合。** 第 ④ 项专门拦截「训练大幅提升、验证退化」:本例候选给 baseline 注入
63+
激进检索行为,训练集 +0.53 但验证集回落,gate 据此 REJECT。关键 case 退化(③)与新增
64+
hard fail(②)提供正交的二次保险,即便总分变化很小也能拦住有害候选。
65+
66+
**产物审计(阶段 6)。** `optimization_report.json` 持久化 baseline/candidate 逐 case 分数与
67+
轨迹、逐 case delta、失败归因、gate 各检查、决策理由、成本/耗时/seed、prompt 的 SHA-256 与
68+
config 快照;`runs/latest/` 留存 baseline 与候选 prompt 全文。`.md` 顶部由 `narrate_report`
69+
依据 gate/delta 数据生成「人话总结」,确定性、无需模型,换输入也不会失真。SDK 桥接通过
70+
`EvalSet.model_validate_json` 校验评测集,并用 trace-only `AgentEvaluator` 跑一次冒烟,证明
71+
管线确实接到真实 SDK 评测器;fake/trace 模式仅在评分/优化处用确定性替身以保证无 key 可复现。
Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,9 @@
1+
{
2+
"_comment": "Per-case metadata kept out of the SDK evalset so EvalSet stays schema-clean. 'key' marks cases that must not regress (gate uses it for critical-regression). 'rubric' selects the rubric dimension scored by run.score_case. 'tool_intent' is a trace-level attribution hint: 'authoritative_search' means the expected trajectory relies on the authoritative search tool, so a wrong tool is attributed to knowledge-recall insufficiency rather than a generic tool error.",
3+
"train_ip_lookup_optimizable": {"key": false, "rubric": "none"},
4+
"train_calendar_optimizable": {"key": false, "rubric": "none"},
5+
"train_strict_json_ineffective": {"key": false, "rubric": "json_format"},
6+
"val_search_fallback_new_pass": {"key": false, "rubric": "none", "tool_intent": "authoritative_search"},
7+
"val_smalltalk_no_tool_regression": {"key": true, "rubric": "no_tool"},
8+
"val_weather_soft_degradation": {"key": true, "rubric": "single_tool"}
9+
}

0 commit comments

Comments
 (0)