Loop Engineering Overview

Deep Dive: Evaluation & Observability cho Agent Loops

Evaluation nói loop có đạt outcome không; observability giải thích nó đã đi đường nào; receipt nối hai lớp bằng evidence.
Báo cáo cha: ← OverviewTopic: Evaluation & ObservabilityĐơn vị đo: Task · Trial · Trace · Outcome · CostNgày: 2026-07-20

Pass rate không giải thích được vì sao; traces không chứng minh được đúng Thesis

Evaluation và observability trả lời hai câu khác nhau. Eval so outcome với acceptance criteria trên một distribution task; observability tái dựng trigger, action, state và decision của từng run. Một dashboard chỉ có success rate không chỉ ra regression đến từ model, intake, verifier hay runtime.

Ngược lại, trace đầy đủ không bảo đảm hệ thống làm đúng. Hàng nghìn spans có thể mô tả rất chính xác một loop đang tối ưu nhầm mục tiêu. Outcome grader phải đọc authoritative state hoặc artifact, không tin final message của agent.

Receipt là cầu nối. Nó giữ work identity, state versions, evidence pointers, verifier/policy versions, decision branch và cost. Nhờ receipt, một failed eval có thể truy ngược causal path, còn một production incident có thể được đưa vào regression suite.

Đơn vị nghiên cứu vì thế không phải một model call mà là full trial: trigger → intake → act → verify → persist → decide. Nếu bỏ scheduler, recovery hoặc state khỏi harness, eval có thể đánh giá actor tốt nhưng bỏ sót class lỗi làm loop thất bại.

Measurement invariant: mọi claim “success” phải có outcome evidence; mọi decision retry/escalate/exit phải trỏ tới evidence và policy version đã tạo ra nó.
Lifecycle đầy đủ của loop từ trigger đến decision
Evaluation unit cần bao phủ cả lifecycle; chỉ chấm output của Act bỏ qua intake, recovery và stopping. ↗ Loop lifecycle

Evaluation protocol: task, trial, grader, transcript, outcome Mental model

Anthropic tách task (đề bài), trial (một lần chạy), transcript (hành vi), outcome (trạng thái cuối), grader (cách chấm) và harness (môi trường). Tách các khái niệm này ngăn việc coi log đẹp hay câu trả lời tự tin là bằng chứng hoàn thành.

Receipt nối eval plane với telemetry plane và promotion decision.

flowchart LR
  TS["Task set + environment"] --> TR["Trial"]
  TR --> OUT["Outcome artifacts"]
  TR --> TRACE["Trace: spans + events"]
  OUT --> G["Graders"]
  TRACE --> R["Causal receipt"]
  G --> SCORE["Outcome + consistency"]
  R --> SCORE
  SCORE --> D{"Promote / hold / demote"}
  D --> REG["Regression + incident corpus"]

Evaluation suite cần distribution có normal, edge, adversarial, recovery và no-work cases. Chỉ chọn happy path dễ đo task completion nhưng không đo false work, duplicate side effect hoặc inability to stop.

Baseline cũng phải factorized: model/prompt, harness/runtime, contract/policy và environment version. Nếu thay cả bốn rồi pass rate đổi, tổ chức không biết leverage point nào thực sự hiệu quả.

Bốn kỹ thuật đo recurring systems Evaluation

1. Đánh giá full loop và giữ cost-matched baselines

E.1
FUTURE-DIRECTIONS.md · Anthropic agent evals
Task success của actor không bằng reliability của loop. Hệ thống còn phải chọn đúng việc, resume đúng state, không duplicate và dừng đúng lý do.

Mỗi task fixture nên mô tả initial state, trigger history, authoritative source, allowed side effects, budget, expected exit và cleanup. Harness phải có khả năng inject duplicate event, stale input, tool timeout và process crash.

So sánh loop với supervised single run, deterministic automation và previous version. Baseline phải dùng cùng task distribution và cost envelope; nếu loop dùng nhiều lần compute hơn, report cả quality gain lẫn marginal cost.

Eval familyCâu hỏiFixture bắt buộc
SelectionCó nhận đúng item/version?Duplicate, stale, no-work
ExecutionCó tạo artifact đúng?Normal + adversarial tasks
RecoveryCó resume không lặp side effect?Crash windows + tool outage
TerminationCó dừng/escalate đúng?Plateau, budget, ambiguity
GovernanceCó giữ scope/quyền?Untrusted input + forbidden action

Trial phải reset hoặc snapshot environment để tránh contamination. Nếu trial trước để lại cache, branch hay external resource, trial sau không còn độc lập và variance bị che. Side-effectful eval cần sandbox hoặc compensating cleanup.

Dataset của awesome-loop-engineering là evidence map, không phải benchmark outcome. Nó hữu ích phân tầng nguồn và chọn pattern, nhưng một shared protocol vẫn cần task fixtures, executable checks và longitudinal runs.

Ưu điểm
  • Đo đúng unit of deployment.
  • Failure injection bộc lộ recovery debt.
  • Cost-matched baselines ngăn claim phóng đại.
Nhược điểm
  • Harness full-loop tốn công hơn chấm final answer.
  • External systems khó reset hoàn toàn.
  • Factorized experiments cần nhiều trials.

2. Ghép deterministic, model và human graders theo verification hierarchy

E.2
Anthropic — Demystifying evals
Mỗi grader là proxy có điểm mù. Composition tốt giữ deterministic invariant cứng và dành semantic judgment cho rubric mềm.

Deterministic graders mạnh ở schema, tests, diff, invariant và resource state. Model graders mạnh ở relevance, style và open-ended quality nhưng có variance, bias và khả năng bị gaming. Human review mạnh ở ambiguity và accountability, song đắt và không hoàn toàn nhất quán.

Gate không nên lấy trung bình tùy tiện. Forbidden write hoặc failing invariant là hard fail dù model judge thích output. Human có thể override nhưng override phải là decision record với reason, không bị hòa vào score.

pass@k hỏi trong k trials có ít nhất một lần thành công—hữu ích cho khả năng tìm lời giải. pass^k hỏi cả k lần đều thành công—gần hơn với reliability. Recurring production cần consistency, không chỉ khả năng may mắn.

Grader QA trước khi tin metric
  • Calibrate trên labeled cases và disagreement set.
  • Freeze version theo report; không đổi rubric âm thầm.
  • Review transcripts ở cả pass và fail.
  • Theo dõi false pass, false block và override rate.
  • Không để actor sửa grader assets.

Grader drift là production change. Khi rubric, model judge hoặc test suite đổi, dashboard phải tạo series mới hoặc backfill có kiểm soát. So điểm giữa hai grader version như cùng metric sẽ tạo trend giả.

Ưu điểm
  • Layered grading tăng coverage.
  • Hard invariants giữ safety boundary.
  • Consistency metric phù hợp recurring work.
Nhược điểm
  • Human labels và calibration tốn chi phí.
  • Model graders có correlated error với actor.
  • Version drift làm trend khó đọc.

3. Thiết kế traces và receipts cho causal debugging

E.3
OpenTelemetry GenAI conventions · FUTURE-DIRECTIONS.md
Log kể sự kiện; causal receipt giải thích transition. Debugging cần biết evidence nào khiến policy chọn cạnh nào trên state version nào.

OpenTelemetry GenAI conventions cung cấp vocabulary cho model invocation, tool definitions và tool execution. Loop vẫn cần domain spans/events: trigger received, intake accepted, lease acquired, checkpoint, verification, decision, handoff, escalation và exit.

Correlation IDs cần nối event ID, work key, run, attempt, trace, state version và external operation. Chỉ có trace ID không đủ để nhóm nhiều cold runs của cùng một recurring item.

Causal receipt tối thiểu cho một successful exit

JSON
{
  "work_key": "repo/pr-381/check-linux@8f3c1a2",
  "run_id": "run-0042",
  "state_before": 17,
  "trigger": { "type": "check_failed", "event_id": "evt-91" },
  "actions": [{ "tool": "shell", "span_id": "a31", "result": "exit:0" }],
  "verification": { "grader": "ci-repro@v4", "outcome": "pass" },
  "decision": { "branch": "exit-success", "policy": "repair-loop@v3" },
  "state_after": 18,
  "cost": { "tokens": 18422, "seconds": 391 }
}

Input/output model có thể chứa secret, personal data và hostile content. Capture nên opt-in, redact theo field, lưu artifact lớn bằng pointer và có retention/access policy. Observability không được tạo kho dữ liệu nhạy cảm mới.

Sampling dựa trên success dễ bỏ mất incident hiếm. Giữ 100% decision/receipt metadata nhỏ, sample payload lớn, và tail-sample mọi escalation, policy violation, recovery hoặc abnormal cost.

Loop Contract cards gồm verification, state, budget, escalation và exit
Receipt cần phản ánh toàn bộ contract, không chỉ model latency và tokens. ↗ Loop Contract schema
Ưu điểm
  • Causal chain rút ngắn incident triage.
  • Cross-run correlation lộ retry và recovery debt.
  • Standard spans tăng portability giữa runtimes.
Nhược điểm
  • High-cardinality telemetry tốn storage.
  • Payload capture có privacy/security risk.
  • Instrumentation thiếu transition tạo blind spot.

4. Theo dõi reliability, recovery và human correction theo thời gian

E.4
Future Directions workstreams 1–8
Throughput có thể tăng trong khi hệ thống xấu đi. Cost, duplicate, intervention và recovery time là guardrails của autonomy.

North-star nên là accepted outcomes trên total operational cost, không phải số tasks closed. Cost gồm model/tool, infrastructure, reviewer attention, incident cleanup và rework. Một loop rẻ mỗi run nhưng retry nhiều có thể đắt hơn.

Reliability metrics cần time horizon: success probability, consistency, time-to-evidence, recovery success, duplicate side-effect rate, stale-work rate, budget exhaustion, escalation precision và human correction rate.

Human interaction không chỉ là overhead. Interrupt, override, approval denial và manual cleanup là labeled signals về boundary sai. Theo dõi reason codes giúp biết nên sửa intake, verifier, permissions hay model.

MetricTốt khiCảnh báo
Verified outcome rateTăng trên frozen tasksAgent claims tăng nhưng artifacts không tăng
pass^k / consistencyỔn định qua repeated trialspass@k cao, pass^k thấp
Recovery successCrash resume không duplicateRestart-from-zero hoặc manual cleanup
Human correctionGiảm mà safety giữ nguyênAuto-approval tăng nhưng overrides cũng tăng
Cost per accepted outcomeGiảm cùng qualityRetries/tail cost bị che bởi average

Slice metric theo trigger, work type, risk, runtime, model, policy và verifier version. Aggregate pass rate có thể che một tenant hoặc action class đang thất bại. Mọi slice cần đủ sample và uncertainty, tránh xếp hạng từ noise.

Ưu điểm
  • Phát hiện regression vận hành chứ không chỉ model quality.
  • Human corrections trở thành feedback có cấu trúc.
  • Cost/outcome hỗ trợ promotion decision.
Nhược điểm
  • Metric taxonomy cần ownership lâu dài.
  • Rare incidents khó có statistical power.
  • Dashboard dễ tạo proxy gaming nếu không review traces.

Measurement matrix: outcome × control plane Scorecard

Dashboard nên trả lời ba lớp: hệ thống có đúng không, có ổn định không, và nếu sai thì failure nằm ở phase nào. Mỗi KPI cần drill-down tới receipts mẫu.

PhaseOutcome metricReliability guardrailTrace event
IntakeEligible items acceptedDuplicate/stale/no-work rateadmission.decision
ActArtifact correctnessForbidden mutation/tool errortool.execute
VerifyCalibrated pass/failFalse pass/override rateverification.result
PersistCheckpoint committedState conflict/lost writestate.transition
DecideCorrect exit/escalationRetry plateau/budget breachpolicy.decision

Một alert chỉ đáng action nếu có owner và drill path. “Agent quality giảm” quá mơ hồ; “recovery success của event-triggered deploy loops giảm sau policy v7” cho phép rollback hoặc cô lập.

Minimum reproducibility bundle Artifact

Mỗi benchmark release nên đóng băng task fixtures, environment image, contract/policy, model/runtime identifiers, grader versions, seeds/config, raw receipts và aggregate code.

Production incident có thể được giảm thiểu thành fixture mới. Đây là vòng học thật: observation → causal diagnosis → regression case → policy/verifier change → canary. Nếu chỉ sửa prompt, organization memory vẫn nằm trong chat.

Bốn lớp prompt, context, harness và loop engineering
Factorized eval cần version từng lớp để attribution không bị trộn. ↗ Engineering stack
Definition of ready cho measurement
  • Outcome đọc từ authoritative artifact, không từ self-report.
  • Trial bao phủ intake, recovery và exit—not chỉ Act.
  • Receipt nối work key, state, evidence, policy và cost.
  • Dashboard có slice, uncertainty, owner và trace drill-down.

Kết luận: measurement phải tái dựng được quyết định Conclusion

Eval tốt giúp chọn hệ thống; observability tốt giúp vận hành nó. Với recurring loops, hai lớp chỉ thực sự hữu ích khi cùng chia sẻ work identity và causal receipts.

Mục tiêu không phải thu nhiều log hơn mà là trả lời nhanh: outcome nào sai, state nào dẫn tới nó, policy nào đã quyết định, và control nào cần thay đổi mà không phá các lớp còn lại.