Deep Dive: Evaluation & Observability cho Agent Loops
Pass rate không giải thích được vì sao; traces không chứng minh được đúng Thesis
Evaluation và observability trả lời hai câu khác nhau. Eval so outcome với acceptance criteria trên một distribution task; observability tái dựng trigger, action, state và decision của từng run. Một dashboard chỉ có success rate không chỉ ra regression đến từ model, intake, verifier hay runtime.
Ngược lại, trace đầy đủ không bảo đảm hệ thống làm đúng. Hàng nghìn spans có thể mô tả rất chính xác một loop đang tối ưu nhầm mục tiêu. Outcome grader phải đọc authoritative state hoặc artifact, không tin final message của agent.
Receipt là cầu nối. Nó giữ work identity, state versions, evidence pointers, verifier/policy versions, decision branch và cost. Nhờ receipt, một failed eval có thể truy ngược causal path, còn một production incident có thể được đưa vào regression suite.
Đơn vị nghiên cứu vì thế không phải một model call mà là full trial: trigger → intake → act → verify → persist → decide. Nếu bỏ scheduler, recovery hoặc state khỏi harness, eval có thể đánh giá actor tốt nhưng bỏ sót class lỗi làm loop thất bại.
Evaluation protocol: task, trial, grader, transcript, outcome Mental model
Anthropic tách task (đề bài), trial (một lần chạy), transcript (hành vi), outcome (trạng thái cuối), grader (cách chấm) và harness (môi trường). Tách các khái niệm này ngăn việc coi log đẹp hay câu trả lời tự tin là bằng chứng hoàn thành.
Receipt nối eval plane với telemetry plane và promotion decision.
flowchart LR
TS["Task set + environment"] --> TR["Trial"]
TR --> OUT["Outcome artifacts"]
TR --> TRACE["Trace: spans + events"]
OUT --> G["Graders"]
TRACE --> R["Causal receipt"]
G --> SCORE["Outcome + consistency"]
R --> SCORE
SCORE --> D{"Promote / hold / demote"}
D --> REG["Regression + incident corpus"] Evaluation suite cần distribution có normal, edge, adversarial, recovery và no-work cases. Chỉ chọn happy path dễ đo task completion nhưng không đo false work, duplicate side effect hoặc inability to stop.
Baseline cũng phải factorized: model/prompt, harness/runtime, contract/policy và environment version. Nếu thay cả bốn rồi pass rate đổi, tổ chức không biết leverage point nào thực sự hiệu quả.
Bốn kỹ thuật đo recurring systems Evaluation
1. Đánh giá full loop và giữ cost-matched baselines
E.1Mỗi task fixture nên mô tả initial state, trigger history, authoritative source, allowed side effects, budget, expected exit và cleanup. Harness phải có khả năng inject duplicate event, stale input, tool timeout và process crash.
So sánh loop với supervised single run, deterministic automation và previous version. Baseline phải dùng cùng task distribution và cost envelope; nếu loop dùng nhiều lần compute hơn, report cả quality gain lẫn marginal cost.
| Eval family | Câu hỏi | Fixture bắt buộc |
|---|---|---|
| Selection | Có nhận đúng item/version? | Duplicate, stale, no-work |
| Execution | Có tạo artifact đúng? | Normal + adversarial tasks |
| Recovery | Có resume không lặp side effect? | Crash windows + tool outage |
| Termination | Có dừng/escalate đúng? | Plateau, budget, ambiguity |
| Governance | Có giữ scope/quyền? | Untrusted input + forbidden action |
Trial phải reset hoặc snapshot environment để tránh contamination. Nếu trial trước để lại cache, branch hay external resource, trial sau không còn độc lập và variance bị che. Side-effectful eval cần sandbox hoặc compensating cleanup.
Dataset của awesome-loop-engineering là evidence map, không phải benchmark outcome. Nó hữu ích phân tầng nguồn và chọn pattern, nhưng một shared protocol vẫn cần task fixtures, executable checks và longitudinal runs.
Ưu điểm
- Đo đúng unit of deployment.
- Failure injection bộc lộ recovery debt.
- Cost-matched baselines ngăn claim phóng đại.
Nhược điểm
- Harness full-loop tốn công hơn chấm final answer.
- External systems khó reset hoàn toàn.
- Factorized experiments cần nhiều trials.
2. Ghép deterministic, model và human graders theo verification hierarchy
E.2Deterministic graders mạnh ở schema, tests, diff, invariant và resource state. Model graders mạnh ở relevance, style và open-ended quality nhưng có variance, bias và khả năng bị gaming. Human review mạnh ở ambiguity và accountability, song đắt và không hoàn toàn nhất quán.
Gate không nên lấy trung bình tùy tiện. Forbidden write hoặc failing invariant là hard fail dù model judge thích output. Human có thể override nhưng override phải là decision record với reason, không bị hòa vào score.
pass@k hỏi trong k trials có ít nhất một lần thành công—hữu ích cho khả năng tìm lời giải. pass^k
hỏi cả k lần đều thành công—gần hơn với reliability. Recurring production cần consistency, không chỉ khả năng may mắn.
Grader QA trước khi tin metric
- Calibrate trên labeled cases và disagreement set.
- Freeze version theo report; không đổi rubric âm thầm.
- Review transcripts ở cả pass và fail.
- Theo dõi false pass, false block và override rate.
- Không để actor sửa grader assets.
Grader drift là production change. Khi rubric, model judge hoặc test suite đổi, dashboard phải tạo series mới hoặc backfill có kiểm soát. So điểm giữa hai grader version như cùng metric sẽ tạo trend giả.
Ưu điểm
- Layered grading tăng coverage.
- Hard invariants giữ safety boundary.
- Consistency metric phù hợp recurring work.
Nhược điểm
- Human labels và calibration tốn chi phí.
- Model graders có correlated error với actor.
- Version drift làm trend khó đọc.
3. Thiết kế traces và receipts cho causal debugging
E.3OpenTelemetry GenAI conventions cung cấp vocabulary cho model invocation, tool definitions và tool execution. Loop vẫn cần domain spans/events: trigger received, intake accepted, lease acquired, checkpoint, verification, decision, handoff, escalation và exit.
Correlation IDs cần nối event ID, work key, run, attempt, trace, state version và external operation. Chỉ có trace ID không đủ để nhóm nhiều cold runs của cùng một recurring item.
Causal receipt tối thiểu cho một successful exit
{
"work_key": "repo/pr-381/check-linux@8f3c1a2",
"run_id": "run-0042",
"state_before": 17,
"trigger": { "type": "check_failed", "event_id": "evt-91" },
"actions": [{ "tool": "shell", "span_id": "a31", "result": "exit:0" }],
"verification": { "grader": "ci-repro@v4", "outcome": "pass" },
"decision": { "branch": "exit-success", "policy": "repair-loop@v3" },
"state_after": 18,
"cost": { "tokens": 18422, "seconds": 391 }
}Input/output model có thể chứa secret, personal data và hostile content. Capture nên opt-in, redact theo field, lưu artifact lớn bằng pointer và có retention/access policy. Observability không được tạo kho dữ liệu nhạy cảm mới.
Sampling dựa trên success dễ bỏ mất incident hiếm. Giữ 100% decision/receipt metadata nhỏ, sample payload lớn, và tail-sample mọi escalation, policy violation, recovery hoặc abnormal cost.
Ưu điểm
- Causal chain rút ngắn incident triage.
- Cross-run correlation lộ retry và recovery debt.
- Standard spans tăng portability giữa runtimes.
Nhược điểm
- High-cardinality telemetry tốn storage.
- Payload capture có privacy/security risk.
- Instrumentation thiếu transition tạo blind spot.
4. Theo dõi reliability, recovery và human correction theo thời gian
E.4North-star nên là accepted outcomes trên total operational cost, không phải số tasks closed. Cost gồm model/tool, infrastructure, reviewer attention, incident cleanup và rework. Một loop rẻ mỗi run nhưng retry nhiều có thể đắt hơn.
Reliability metrics cần time horizon: success probability, consistency, time-to-evidence, recovery success, duplicate side-effect rate, stale-work rate, budget exhaustion, escalation precision và human correction rate.
Human interaction không chỉ là overhead. Interrupt, override, approval denial và manual cleanup là labeled signals về boundary sai. Theo dõi reason codes giúp biết nên sửa intake, verifier, permissions hay model.
| Metric | Tốt khi | Cảnh báo |
|---|---|---|
| Verified outcome rate | Tăng trên frozen tasks | Agent claims tăng nhưng artifacts không tăng |
| pass^k / consistency | Ổn định qua repeated trials | pass@k cao, pass^k thấp |
| Recovery success | Crash resume không duplicate | Restart-from-zero hoặc manual cleanup |
| Human correction | Giảm mà safety giữ nguyên | Auto-approval tăng nhưng overrides cũng tăng |
| Cost per accepted outcome | Giảm cùng quality | Retries/tail cost bị che bởi average |
Slice metric theo trigger, work type, risk, runtime, model, policy và verifier version. Aggregate pass rate có thể che một tenant hoặc action class đang thất bại. Mọi slice cần đủ sample và uncertainty, tránh xếp hạng từ noise.
Ưu điểm
- Phát hiện regression vận hành chứ không chỉ model quality.
- Human corrections trở thành feedback có cấu trúc.
- Cost/outcome hỗ trợ promotion decision.
Nhược điểm
- Metric taxonomy cần ownership lâu dài.
- Rare incidents khó có statistical power.
- Dashboard dễ tạo proxy gaming nếu không review traces.
Measurement matrix: outcome × control plane Scorecard
Dashboard nên trả lời ba lớp: hệ thống có đúng không, có ổn định không, và nếu sai thì failure nằm ở phase nào. Mỗi KPI cần drill-down tới receipts mẫu.
| Phase | Outcome metric | Reliability guardrail | Trace event |
|---|---|---|---|
| Intake | Eligible items accepted | Duplicate/stale/no-work rate | admission.decision |
| Act | Artifact correctness | Forbidden mutation/tool error | tool.execute |
| Verify | Calibrated pass/fail | False pass/override rate | verification.result |
| Persist | Checkpoint committed | State conflict/lost write | state.transition |
| Decide | Correct exit/escalation | Retry plateau/budget breach | policy.decision |
Một alert chỉ đáng action nếu có owner và drill path. “Agent quality giảm” quá mơ hồ; “recovery success của event-triggered deploy loops giảm sau policy v7” cho phép rollback hoặc cô lập.
Minimum reproducibility bundle Artifact
Mỗi benchmark release nên đóng băng task fixtures, environment image, contract/policy, model/runtime identifiers, grader versions, seeds/config, raw receipts và aggregate code.
Production incident có thể được giảm thiểu thành fixture mới. Đây là vòng học thật: observation → causal diagnosis → regression case → policy/verifier change → canary. Nếu chỉ sửa prompt, organization memory vẫn nằm trong chat.
- Outcome đọc từ authoritative artifact, không từ self-report.
- Trial bao phủ intake, recovery và exit—not chỉ Act.
- Receipt nối work key, state, evidence, policy và cost.
- Dashboard có slice, uncertainty, owner và trace drill-down.
Kết luận: measurement phải tái dựng được quyết định Conclusion
Eval tốt giúp chọn hệ thống; observability tốt giúp vận hành nó. Với recurring loops, hai lớp chỉ thực sự hữu ích khi cùng chia sẻ work identity và causal receipts.
Mục tiêu không phải thu nhiều log hơn mà là trả lời nhanh: outcome nào sai, state nào dẫn tới nó, policy nào đã quyết định, và control nào cần thay đổi mà không phá các lớp còn lại.