Loop Engineering Overview

Deep Dive: Runtime, Pattern Selection & Loop Maturity

Từ báo cáo Loop Engineering Overview: chọn operating model nhỏ nhất có thể thực hiện recurring job an toàn và có bằng chứng.
Báo cáo cha: ← OverviewTopic: Runtime & MaturityCoverage: 20 patterns · 8 starters · 7 levelsNgày: 2026-07-19

Loop là design; runtime là nơi design được thực thi Thesis

Một CI repair contract có thể chạy trong interactive session, scheduled desktop task, Codex automation, GitHub Actions, shell/cron hoặc durable workflow engine. Điều cần giữ ổn định là objective, permissions, evidence, state, budget và handoff; điều thay đổi là wiring của trigger, workspace, credentials, checkpoint và notification.

Nhầm design với runtime tạo hai dạng lock-in. Technical lock-in xảy ra khi state và transition chỉ tồn tại trong API vendor; conceptual lock-in xảy ra khi team mô tả job bằng feature của tool thay vì objective và evidence. Khi đó, đổi execution surface không chỉ là migration mà buộc team tái khám phá chính operating policy của mình.

Contract-first không có nghĩa runtime là commodity. Runtime quyết định policy có được enforce thật hay chỉ được nhắc trong prompt: schedule có chống overlap không, worktree có sạch không, credential có scope theo job không, checkpoint có atomic không, approval có block side effect không. Cùng contract trên hai runtime có thể cho reliability rất khác nếu enforcement primitives khác nhau.

Vì vậy, runtime selection là một bài toán fit giữa job shape và failure cost. Local session tối ưu cho tốc độ học khi human nearby; CI tối ưu cho repo event và required checks; durable engine tối ưu cho crash recovery và long-lived side effects. Không có runtime “mature nhất” cho mọi loop — chỉ có runtime nhỏ nhất đáp ứng những semantics bắt buộc.

20 Operating patterns 4 domain groups
8 Runtime starters 3 executable + 5 templates
7 Maturity levels 0 through 6
1 Stable contract Migrate wiring, not intent
Selection order: recurring symptom → verified finish → pattern → adapted contract → runtime. Chọn tool trước thường làm team copy permission và scheduling defaults không phù hợp với risk của job.

Thứ tự lựa chọn còn bảo vệ economics. Khi bắt đầu từ tool, team dễ sử dụng schedule, multi-agent và durable storage vì chúng có sẵn, trước khi chứng minh recurrence hoặc verified value. Khi bắt đầu từ symptom và finish condition, mỗi capability mới phải trả lời một failure hoặc constraint cụ thể, nhờ đó complexity có business case rõ.

Runtime map Map

Runtime map nên được đọc như một cây loại trừ. Persistence loại các surface không sống qua process; file locality loại cloud clone khi input chỉ có trên máy; event source ưu tiên CI khi work gắn với repository; crash sensitivity đẩy job sang durable execution. Feature comparison chỉ hữu ích sau khi các constraint bắt buộc đã được thỏa.

Bản đồ runtime cho Loop Engineering từ session local tới cloud, CI và durable execution
Runtime được phân biệt chủ yếu bởi persistence, file access, isolation, permission model và escalation surface. ↗ Runtime Selection Guide

Persistence có ít nhất ba mức. Session persistence giữ context khi process sống. Cross-run persistence giữ checkpoint giữa các invocation. Durable execution còn đảm bảo timer, retry và transition có thể replay sau crash. Một file progress đáp ứng mức hai cho job nhỏ nhưng không tự cung cấp transaction semantics của mức ba.

Decision flow: persistence và blast radius thường loại runtime nhanh hơn feature checklist.

flowchart TD
  A["Recurring job"] --> P{"Must survive app / machine / process?"}
  P -->|"no"| L["Interactive session or local starter"]
  P -->|"across runs"| F{"Needs local non-Git files?"}
  F -->|"yes"| D["Desktop scheduled task or controlled local cron"]
  F -->|"no"| G{"GitHub event / repo-centric?"}
  G -->|"yes"| C["CI / agentic workflow / cloud automation"]
  G -->|"no"| X{"Crash-sensitive or long-lived?"}
  X -->|"yes"| W["Custom durable workflow runtime"]
  X -->|"no"| B["Background cloud automation"]
  L --> S["Add worktree / sandbox + external gate"]
  D --> S
  C --> S
  W --> S
  B --> S

File access và isolation thường xung đột. Local runtime thấy uncommitted files, private corpus và desktop apps nhưng dễ inherit môi trường bẩn cùng privilege người dùng. Cloud/CI tạo fresh clone và credential tách biệt nhưng mất local context. Contract phải chỉ rõ source-of-truth; runtime không nên bí mật bù thiếu context bằng cách đọc rộng hơn.

Escalation surface cũng là capability runtime. Report trong inbox phù hợp với maintenance không khẩn cấp; PR comment phù hợp với code review; pager phù hợp với incident. Nếu handoff cần action trong vài phút nhưng runtime chỉ tạo artifact không ai theo dõi, contract có escalation field nhưng operating system chưa có escalation path thật.

Bốn quyết định vận hành Deep Dive

1. Chọn pattern từ symptom và verified finish

O.1
patterns/README.md · patterns/MATRIX.md
Pattern là failure-mode playbook, không phải agent persona. Cùng một model xử lý PR, deploy hay red team nhưng mỗi job có intake, action boundary, evidence và human owner khác nhau.

Symptom mô tả pain lặp lại trong hệ thống, còn pattern mô tả control structure có khả năng xử lý pain đó. “PR bị kẹt” có frequency, owner và cost; “manager-worker agents” chỉ mô tả topology. Bắt đầu từ symptom giữ thiết kế gắn với outcome và cho phép so sánh loop với giải pháp đơn giản hơn như alert, checklist hoặc deterministic workflow.

Recurring symptomPatternLoop có thể kết thúc khi
PR bị kẹtPR babysitterRequired checks pass, review threads resolved, merge state current
CI fail lặp lạiCI repairOriginal failing command passes với scoped patch
Docs có thể staleDocs driftVerified mismatches fixed và examples còn chạy
Deploy cần theo dõiDeploy verifierSynthetic checks/thresholds giữ trong policy
Eval score tụtEvaluation regressionTargeted evals về baseline mà không đổi scorer
Corpus knowledge staleKnowledge freshnessVersioned corpus pass provenance/freshness/retrieval/leakage gates
Sensitive changeSecurity reviewFindings có evidence và approval boundary nguyên vẹn
Cần tìm failure mớiAdversarial red teamFinding reproduced, minimized, reported và regression-tested

Verified finish là phép thử xem symptom đã đủ rõ để loop hóa chưa. “Cải thiện docs” không có terminal observation; “mọi API thay đổi trong release có reference tương ứng, examples chạy và diff nằm trong docs” thì có. Nếu finish cần judgment không thể đóng thành gate, pattern vẫn có thể report candidate nhưng không nên tự exit hoặc merge.

Bốn domain của 20-pattern library
  • Build & Maintain (6): PR, CI, docs, dependency, bug hunting, release notes.
  • Operate & Observe (6): deploy, incident, data quality, cost, model routing, performance.
  • Learn & Optimize (4): feedback, eval regression, benchmark optimization, knowledge freshness.
  • Govern & Protect (4): security, enterprise approval, accessibility, adversarial red team.

Domain grouping phản ánh asset và feedback khác nhau. Build loops thường có diff và test; operations loops có metric, SLO và incident owner; learning loops có dataset, baseline và leakage risk; governance loops có evidence bundle và approval authority. Dùng cùng một “generic autonomous agent” cho bốn domain sẽ che các verifier và blast radius khác nhau.

Pattern composition cần một primary owner và terminal policy thống nhất. Ví dụ deploy verifier có thể gọi performance diagnosis, nhưng loop con không được tự thay đổi production nếu contract cha chỉ cấp quyền report. Composition an toàn truyền budget, permission và evidence boundary xuống; nó không cộng dồn quyền của mọi pattern thành một super-agent.

Ưu điểm
  • Symptom-first selection gắn loop vào recurring value thật.
  • Named pattern làm failure boundaries và finish reviewable.
  • Pattern matrix giúp so sánh state/gate/budget/escalation ngang hàng.
Nhược điểm
  • Pattern library không thay domain-specific threat model.
  • Một job lai có thể cần compose patterns và một owner rõ.
  • Copy nguyên default budget có thể sai blast radius thực tế.

2. Chọn runtime từ sáu constraints, không từ brand

O.2
meta/RUNTIME_SELECTION.md
Runtime fit là operational fit. Một surface nhiều feature vẫn sai nếu không truy cập được state source, không isolate workspace, inherit privilege quá rộng hoặc không có escalation channel.

Sáu constraint là cách dịch contract semantics sang infrastructure requirements. “State update after each attempt” trở thành cross-run store và commit point; “clean workspace” trở thành worktree/container; “named human owner” trở thành notification channel có identity và SLO. Selection tốt đánh giá mapping này field-by-field, không dựa vào số logo integration trên trang sản phẩm.

ConstraintCâu hỏiHệ quả lựa chọn
PersistenceSession, across runs hay crash-safe?Long-lived critical loop cần durable engine/checkpoint.
File accessLocal non-Git files hay cloned repo?Cloud fresh clone không thấy local uncommitted state.
IsolationWorktree/container có sẵn?Local runtime phải tự bổ sung isolation.
PermissionsSecrets/network/write scope cap được không?Unattended job ưu tiên per-job least privilege.
VerificationGate có chạy y hệt ở runtime?CI natural cho required check; runtime khác phải gọi same command.
EscalationPR/issue/inbox/Slack/pager nào nhận handoff?Incident loop cần page owner; report-only loop có thể dùng inbox.

Constraints có hard và soft. Không truy cập được source-of-truth hoặc không enforce được production approval là hard blocker. Latency cao hơn hay setup phức tạp hơn là soft trade-off. Tách hai loại ngăn team chọn surface tiện nhất rồi gọi những semantics còn thiếu là “future hardening”. Với unattended writes, hardening là prerequisite chứ không phải backlog.

Runtime evaluation nên dùng một representative failure drill: duplicate trigger, crash sau side effect, verifier timeout, secret expired và escalation unavailable. Feature documentation có thể nói có retry/checkpoint, nhưng drill mới cho biết retry granularity, state visibility và operator recovery thực tế có phù hợp contract không.

Capability drift cũng cần owner. Vendor có thể đổi scheduling, permission inheritance hoặc isolation semantics; local cron có thể đổi environment sau OS update. Contract test và runtime adapter test nên chạy định kỳ để phát hiện execution surface không còn giữ policy đã review.

Ưu điểm
  • Giảm lock-in bằng contract-first design.
  • Làm rõ operational debt trước khi schedule.
  • Cho phép cùng pattern đi từ local prototype tới durable production.
Nhược điểm
  • Vendor capability và permission semantics thay đổi theo thời gian.
  • Tự host cho control cao nhưng tăng gánh isolation/secrets/ops.
  • Cloud isolation tốt có thể đánh đổi local context và latency.

3. Tám starter là executable teaching artifacts, không phải framework magic

O.3
examples/runnable/README.md
Starter tốt giữ control loop visible. Verifier exit code, state file, max attempts, allowed paths và stop reason nằm trong code/config có thể đọc thay vì biến mất sau framework abstraction.

Starter hữu ích vì nó làm control flow nhìn thấy được. Builder có thể chỉ vào dòng đếm attempt, file checkpoint, command verifier và nhánh stop. Điều này tạo mental model trước khi framework trừu tượng hóa chúng thành decorator, queue config hoặc platform UI. Đọc starter là học semantics, không phải nhận production guarantee.

StarterFormStateExternal gate
Test repairExecutable BashMarkdown progress ledgerCheck-command exit code
Threshold monitorExecutable Bash, read-onlySample receiptsNumeric threshold
Queue workerExecutable Python + JSONLPer-item JSONL receiptsVerifier exit code
Claude /loopCopy/paste session templateSession + progress fileProject check command
Claude desktop scheduleCopy/paste local templateLast-run marker + outputLive source checks
Codex automationCopy/paste background templateWorktree receipts + reportDeclared repo checks
GitHub agentic workflowActions event/cron templateIssue/PR/artifact/cacheRequired workflow checks
Shell / cronOS scheduling templateLock + progress fileScript exit code

Ba executable starter đại diện ba feedback shape. Test repair dùng binary command feedback và mutation có scope; threshold monitor là read-only với numeric evidence; queue worker xử lý identity và per-item receipt. Chúng cho thấy lifecycle chung có thể giữ nguyên trong khi intake, verifier và state artifact thay theo job.

Dependency-light starter không tự cung cấp exactly-once execution, transactional checkpoint, durable timer hay replay. Crash-sensitive queue, financial/production side effect hoặc long-lived orchestration cần runtime bền hơn.

Khoảng cách prototype–production chủ yếu nằm ở failure semantics. File lock có thể ngăn hai process đơn giản nhưng không giải network partition; append JSONL giữ audit nhưng không tạo transaction với external API; shell timeout dừng command nhưng không chắc side effect bên ngoài đã rollback. Upgrade trigger nên dựa trên những failure này, không dựa trên số run đã chạy.

Cách dùng starter an toàn là fork tối thiểu rồi viết rõ phần nào đang giả lập. Dry-run trước, một item trước, write allowlist nhỏ và verifier độc lập. Sau mỗi failure drill, quyết định control nào có thể bổ sung tại chỗ và control nào đòi runtime khác. Starter tốt giúp ta phát hiện điểm chuyển đó sớm.

Ưu điểm
  • Dễ đọc, dry-run và failure-test.
  • External verifier và state artifact được nêu rõ.
  • Tốt cho bounded prototype trước khi chọn framework.
Nhược điểm
  • Shell/file state yếu khi nhiều worker hoặc crash giữa side effect.
  • Secret isolation, observability và retries phải tự harden.
  • Template vendor cần xác minh theo docs/capability hiện tại.

4. Maturity 0–6 là thứ tự dependency, không phải leaderboard

M.1
README.md#loop-maturity-model
Các level là cumulative capabilities. Multi-agent không chữa missing state; verifier không chữa trigger/intake mơ hồ; production automation không an toàn nếu chưa có telemetry, approval, rollback và owner.

Maturity level mô tả phần trách nhiệm từng do con người giữ đã được hệ thống hóa đến đâu. Level 0: người dùng giữ trigger, state và judgment. Level 2: scheduler và intake được externalize. Level 3–4: memory và verification được externalize. Level 5 tách role; Level 6 externalize cả production supervision nhưng vẫn giữ human owner ở control plane.

Bảy level Loop Maturity Model từ manual prompting đến production-supervised
Earn autonomy theo thứ tự: persist → verify → separate duties → production controls. Dừng ở level thấp nhất đáp ứng job. ↗ Loop Maturity Model

Dependency order ngăn automation khuếch đại debt. Schedule trước no-work/dedupe tạo duplicate work; state trước identity vẫn lưu nhầm item; verifier trước protected boundary có thể bị actor sửa; multi-agent trước handoff contract tạo coordination noise. Promotion chỉ hợp lý khi failure ở level hiện tại không còn được giải tốt bằng control đơn giản hơn.

LevelOperating modelCapability mớiVí dụ
0Manual promptingHuman giữ state và judgeMột fix được review trực tiếp
1Scripted retryBounded feedback wrapperRerun pytest tối đa 3 lần
2Scheduled/event loopFresh intake + no-work exitNightly docs drift report
3Stateful loopCheckpoint + receipts qua runsProcessed feedback IDs + next action
4Self-verifying loopExternal gate kiểm soát exitCI repair chỉ exit khi original command pass
5Multi-agent loopExplicit specialists + handoffsExplorer/implementer/reviewer
6Production-supervisedTelemetry, least privilege, approval, rollbackDeploy verifier với release owner

Promotion signal phải là evidence vận hành: frequency đủ lớn cho schedule, lost-context incident cho durable state, false completion cho external gate, coordination bottleneck hoặc independent-review need cho multi-agent, và real user/infra impact cho production supervision. “Framework hỗ trợ” không phải signal; capability unused vẫn mang complexity và attack surface.

Demotion và retirement cũng là maturity. Nếu job ít xuất hiện, verifier drift, operator burden tăng hoặc product workflow đổi, loop có thể quay về report-only, giảm cadence hay bị tắt. Một operating system trưởng thành quản lý vòng đời của automation, không coi mọi loop đã tạo phải tiếp tục tồn tại.

Ưu điểm
  • Promotion gắn với recurring failure/operating need cụ thể.
  • Ngăn team nhảy thẳng vào orchestration phức tạp.
  • Tạo shared language cho risk review và investment.
Nhược điểm
  • Level không đo model intelligence hay business value.
  • Hai loop cùng level có reliability rất khác.
  • Có thể bị dùng như checklist theater nếu không test failure/recovery.

Decision matrix: runtime nào cho job nào? Decision

Matrix đưa ra starting point chứ không thay contract review. Hai job cùng “repo triage” có thể chọn runtime khác nếu một job cần local proprietary corpus còn job kia chỉ đọc public GitHub event. Hãy đọc từng hàng như default khi các constraint khác bằng nhau, rồi override bằng evidence về source, risk và recovery requirement.

Job shapeRuntime bắt đầuTại saoUpgrade trigger
Interactive fix, human nearbySession loop + worktreeFast feedback, permission rõ trong sessionCần chạy khi app đóng hoặc lặp thường xuyên
Recurring local non-Git filesDesktop scheduled task / controlled cronCó local filesystemMáy/app availability không đủ SLO
Repo event + required checksCI/agentic workflowFresh clone, event trigger, natural status gateWork kéo dài qua run hoặc cần external systems
Background repo triageCloud automation + isolated worktreeSchedule, inbox, scoped repoCần crash-exact resume hoặc complex dependencies
Read-only threshold monitorShell/cron starterDeterministic numeric gate, low blast radiusNeed HA, alert dedupe hoặc durable timers
Queue có side effects quan trọngDurable workflow runtimeCheckpoint, retry semantics, idempotency, observabilityĐây đã là target state

Upgrade trigger nên là một failure hoặc SLO không thể đáp ứng ở surface hiện tại. App đóng làm mất schedule, file state gây duplicate sau crash, queue cần timer bền, hoặc external action cần transaction/reconciliation — đây là các lý do nâng runtime. “Muốn scale” quá chung nếu chưa nêu scale của item, concurrency, state hay recovery.

Ngược lại, tránh over-provision. Read-only threshold monitor với command deterministic có thể chạy đáng tin bằng cron, lock và receipt; durable orchestration platform sẽ thêm infra và on-call mà không tăng verified value tương ứng. Runtime tối ưu là runtime làm failure quan trọng nhất observable và recoverable với tổng chi phí thấp nhất.

Migration path: giữ contract, thay execution wiring Migration

Migration an toàn bắt đầu bằng contract equivalence: objective, intake rule, permissions, gate, state semantics, budget và terminal outcomes phải giữ nghĩa trước/sau. Nếu runtime mới không biểu diễn trực tiếp một field, adapter hoặc control bổ sung phải được thiết kế; không được âm thầm làm field yếu đi để migration “thành công”.

Loop lifecycle giữ nguyên từ intake đến decision trong quá trình đổi runtime
Lifecycle và policy giữ nguyên; trigger adapter, workspace isolation, state backend và escalation transport thay đổi. ↗ Runtime migration path

State migration là phần rủi ro nhất vì nó mang lịch sử và idempotency. Cần map stable item key, attempt sequence, committed side-effect IDs và active owner; chạy dual-read hoặc shadow mode khi cần; freeze writes trong cutover nếu không reconcile được. Copy progress summary mà bỏ ledger có thể làm runtime mới xử lý lại toàn bộ backlog.

  1. 1
    Prototype một iteration

    Interactive/session hoặc executable starter; verifier và receipt phải nhìn thấy.

  2. 2
    Thêm trigger + no-work path

    Schedule/event chỉ sau khi intake, dedupe và idempotency rõ.

  3. 3
    Externalize state

    Progress ledger → issue/database/checkpoint khi work span nhiều run.

  4. 4
    Cứng hóa isolation + permissions

    Worktree/container, scoped secrets, network/action allowlist.

  5. 5
    Promote external gates

    CI/eval/policy/human approvals kiểm soát transition và exit.

  6. 6
    Durable execution khi cần

    Crash replay, transactional side effects, durable timers và observability.

  7. 7
    Production supervision

    SLO, rollback, incident drill, owner, cost controls và retirement plan.

Shadow run là bridge hữu ích: runtime mới intake và verify cùng item nhưng không write, sau đó so decision và receipt với runtime cũ. Disagreement cho thấy khác biệt context, gate hoặc policy adapter. Chỉ chuyển authority sau khi failure paths — không chỉ happy path — cho kết quả tương đương và rollback procedure đã thử.

Migration cũng là dịp xóa debt thay vì bê nguyên abstraction vendor cũ. Giữ semantics của contract, nhưng chuẩn hóa receipt, tách secret, bảo vệ verifier và thêm no-work path nếu trước đây thiếu. Portability không có nghĩa bảo tồn mọi implementation detail; nó nghĩa bảo tồn intent và observable guarantees.

Economics, multi-agent và giới hạn Critical reading

Anthropic khuyến nghị bắt đầu từ solution đơn giản nhất và chỉ tăng agentic complexity khi outcome chứng minh giá trị; agentic systems thường đổi latency/cost lấy flexibility. Addy Osmani cũng lưu ý subagents đốt thêm tokens và human review bandwidth vẫn là trần. Do đó topology và maturity cần một business case, không phải aesthetic preference.

Có thể mô hình hóa thô giá trị mỗi chu kỳ là tần suất × lợi ích của verified outcome, trừ inference, tool, verification, review, failure recovery và platform operations. Từ “verified” là quan trọng: số patch hoặc số run cao không phải value nếu false completion hoặc rework tăng. Economics phải đo ở outcome đã qua gate và chi phí toàn hệ thống.

Thêm capabilityChi phí mớiChỉ nên thêm khi
Schedule/eventNo-work runs, stale intake, duplicate handlingRecurring frequency đủ cao và dedupe rõ
Durable stateStorage, schema, retention, migrationResume/audit tránh lãng phí hoặc incident
Independent verifierLatency, false positives, maintenanceFalse completion cost lớn hơn gate cost
Multi-agentToken, coordination, merge/review taxDecomposition hoặc separation of duties đo được
Durable engineInfra, ops, observabilityCrash recovery và side-effect semantics là requirement
Production writeApproval, rollback, incident riskRead-only/report-only không đạt outcome và controls đã proven

Multi-agent chỉ đáng thêm vì hai lý do chính: decomposition tạo throughput/quality đo được, hoặc separation of duties giảm một risk cụ thể. Persona đa dạng tự thân không phải benefit. Mỗi handoff thêm serialization loss, token, latency, merge conflict và một boundary cần receipt. Team nên so với single agent có tool/verifier tốt hơn trước khi tăng topology.

Review bandwidth thường là constraint ẩn. Loop có thể tạo candidate nhanh hơn nhưng đẩy hàng chục evidence bundle tới maintainer, làm queue quyết định dài hơn. Metrics nên gồm operator minutes trên mỗi verified outcome, escalation precision và tỷ lệ bundle đủ để quyết ngay. Autonomy tốt nén judgment work; autonomy kém chỉ di chuyển bottleneck.

Không dùng maturity như KPI. Một Level 3 loop nhỏ, stateful và idempotent có thể tạo nhiều giá trị hơn Level 5 agent team tốn 10× nhưng verifier yếu. KPI nên là verified outcome, recovery, false completion, operator burden và total cost.

Giới hạn bằng chứng của repo cũng áp dụng ở đây: pattern matrix, starters và maturity model là design framework mạnh, chưa phải benchmark chứng minh runtime hay level nào tạo ROI phổ quát. Mỗi tổ chức cần chạy pilot, ghi baseline của supervised process và quyết định promotion dựa trên verified value, recovery và risk của chính workload.

Kết luận Wrap

Runtime là nơi policy gặp thực tại. Một contract có thể đẹp trên giấy nhưng chỉ trở thành operating guarantee khi trigger, isolation, credential, checkpoint, verifier và escalation của runtime thực sự enforce được các field tương ứng. Vì vậy selection cần bắt đầu từ semantics bắt buộc và được xác nhận bằng failure drill.

Pattern giữ thiết kế gần recurring value; contract giữ intent qua nhiều run; runtime thực thi; maturity cho biết control nào phải có trước khi tăng autonomy. Chuỗi này ngăn team nhảy từ prompt trực tiếp tới agent platform rồi mới phát hiện không biết success, state hay owner là gì.

Maturity tốt là khả năng chọn đúng điểm dừng. Nhiều loop tạo giá trị tối đa ở Level 2–4 với một actor, state rõ và gate mạnh. Multi-agent hoặc production write chỉ nên xuất hiện khi failure/value case đòi hỏi. Complexity không được xem là bằng chứng của tiến bộ.

Quy tắc thực dụng cuối cùng: prototype trên surface rẻ nhất, externalize contract sớm, đo verified outcome và operator burden, rồi chỉ migrate khi constraint hiện tại đã được chứng minh. Bằng cách đó, runtime phục vụ loop design thay vì định nghĩa ngược operating model của tổ chức.

Điểm chính
  • Chọn pattern từ symptom và verified finish, không từ tên model hay runtime.
  • Chọn runtime theo persistence, access, isolation, permissions, verification và escalation.
  • Starter là teaching artifact tốt; production-critical loop cần durability tương xứng side effects.
  • Maturity là dependency order: state → external verification → multi-agent → production controls.