🌐 Languages: English · Tiếng Việt bên dưới
AI coding agents are getting faster, but speed is not trust
AI coding agents can now inspect repositories, edit multiple files, execute tests, open pull requests, recover from build failures, and interact with cloud infrastructure. For enterprise delivery, however, the key question is no longer, “Which model writes the best code?”
The better question is:
“Which system can prove that this change is correct, authorized, reproducible, and safe to release?”
Recent research points to four recurring problems:
An agent may silently resolve an ambiguous requirement in a way the user never intended.
The action a person approves may not be identical to the action ultimately executed.
The same model can perform very differently when placed inside another harness.
A test or benchmark can report success even when the database, repository, or cluster is in the wrong state.
A production-ready coding agent therefore needs more than a model, prompt, and collection of tools. It needs an independently controlled delivery system around the model.
Think of the agent as a very fast home contractor
Imagine hiring a highly capable contractor to renovate your home. You say, “Make the kitchen larger, more modern, and safer.” The contractor understands quickly, owns powerful tools, and can begin immediately but several questions remain:
How much larger should the kitchen become?
Is the wall they plan to remove load-bearing?
What is the budget limit?
Which rooms may they enter, and which keys should they receive?
Who will inspect the electrical, plumbing, and structural work?
If the contractor guesses every answer and works at maximum speed, the kitchen may look impressive while missing your needs, exceeding the budget, or weakening the house. A coding agent is similar: producing code that runs does not prove that it built the right thing.
The everyday mapping is straightforward:
Everyday renovationAI software deliveryRenovation requestBusiness requirement or Jira ticketMeasured blueprintSpecification and acceptance criteriaContractorAI coding agentPower toolsTools, MCP servers, shell, Git, and cloud APIsKeys for selected roomsLeast-privilege permissionsBuilding permitPolicy and approvalElectrical and structural inspectorTests, security scans, and architecture gatesFinal walkthroughRuntime and state verificationSigned handover recordEvidence-bound release receipt
The goal is not to avoid capable contractors. It is to give them a clear blueprint, the right keys, a bounded scope, and an independent inspection process.
The target architecture
The model should propose a plan, patch, tests, and actions. Independent verifiers should collect evidence; a policy engine should control authority; and a release controller should promote only verified artifacts.
The crucial separation is that the agent does not grade its own work or grant itself release authority.
Clarify only when ambiguity changes behavior
CONTRA generates plausible answers to an underspecified requirement and executes the corresponding alternatives. It asks a question only when those alternatives produce a stable behavioral difference, improving macro-average F1 by 13.88 percentage points over the strongest baseline on ClarifyCodeBench.
Consider this requirement:
“Lock an account after five failed payments.”
It leaves important questions unanswered: Which failures count? Over what time window? Does the lock affect the payment flow or the entire portal? Does a later success reset the counter? What happens under concurrent requests?
If the answers produce different database states, API responses, events, or authorization outcomes, the agent should stop and clarify. If they only affect naming or internal structure, it should continue without interrupting the developer.
For an agent workflow, this suggests an executable clarification gate before implementation: generate competing interpretations, identify the observable difference, and attach every question to a test or state transition.
Approval must be bound to the executed action
Approval Laundering describes six ways an approved action may change before execution: scope, arguments, time, tool, delegation, or semantic meaning. Signed approval tokens prevented some tested failures, but could not detect changes occurring beneath the tool-call fields visible to the mediator.
Suppose a user approves deployment of rite-online:1.8.4 to a test namespace. Before execution, the image tag may resolve to a different digest, the kube-context may point to production, Helm values may introduce cluster-wide authority, or another agent may inherit and broaden the operation.
Approval should therefore bind the principal, agent, session, normalized tool arguments, namespace, commit SHA, image digest, policy version, and expiry. After execution, the system must still verify the resulting resources and runtime state.
Evaluate the model–harness–task combination
Finding the Right Fit evaluated 66 model–harness configurations and released 6,204 scored trajectories. On one benchmark, Claude led GPT by 7.94 points under one harness but trailed it by 30.16 under another; a leaner GPT configuration also achieved a stronger result at less than one-quarter of the cost.
Teams using GPT, Claude, Qwen, DeepSeek, or Kimi should not select a model from a single demo. Evaluate the complete tuple:
model + harness + task type + repository type + tool contract + verification gatesBug fixing, Java modernization, database migration, greenfield implementation, Kubernetes incident response, and cross-repository change each deserve a separately measured profile.
Evidence quality matters more than reviewer size
Groundability, Not Scale Alone shows that smaller reviewers can audit stronger coding agents when they receive decisive execution evidence. When the evidence was incomplete, an automated cascade still caught 76–80% of defects but incorrectly rejected 66–67% of acceptable results.
The lesson is not to automate every review with a small model. It is that reviewers need an abstention path: when mandatory evidence is missing, they should escalate rather than manufacture a binary verdict.
For Spring Boot delivery, the evidence package should include JUnit and integration results, Testcontainers state, ArchUnit boundaries, API compatibility, migration validation, security checks, rendered Kubernetes diffs, and runtime health evidence. A high aggregate score must never compensate for a failed mandatory gate.
Verify state, not success messages
Do Agent Benchmarks Do What They Say? confirmed seven defects across 34 mutating tools in four agent benchmarks. In several cases, the evaluator trusted what a tool claimed instead of verifying the state the tool should have changed.
The same failure appears in enterprise systems: an API returns 200 before a transaction rolls back; kubectl apply succeeds while the deployment never becomes ready; a Maven build omits the affected module; a migration runs against the wrong schema; or a producer emits an event the consumer cannot deserialize.
Tool output is a claim. Database, Git, cluster, and runtime state are evidence.
Evaluate SRE agents under continuous change
SRE-Marathon placed agents in a live two-zone Kubernetes environment with roughly sixty overlapping faults per run. The strongest of ten methods scored only 41.3/100: agents often correlated and localized failures but rarely completed repairs while the faults remained active.
Production platform agents must handle overlapping alerts, recent rollouts, configuration changes, intermittent dependencies, concurrent operators, and state preserved across multiple investigation cycles. A “pod crashed, restart pod” task does not measure that capability.
Score the complete sequence: correlation, localization, safe mitigation, restoration, and recurrence prevention.
Small interface changes can produce large reliability gains
MCP Error Messages Written for Developers Hurt Agents examined 3,001 messages from 150 MCP servers. Many told agents to open a browser, edit configuration, or simply wait even when those actions were unavailable. Naming the correct recovery tool raised credential-error recovery to 84% and rate-limit recovery to 88%.
Machine-facing errors should state a stable code, retryability, recovery tool, retry tool, required state, and correlation ID. Better tool contracts can make cheaper and smaller models more reliable; not every problem requires a larger model.
Nine quick analogies for explaining the system
Technical conceptEveryday analogyEngineering useClarification gateA tailor measures before cutting the fabricClarify irreversible behavioral choices before editing codeApproval bindingA signed check fixes both the recipient and amountBind approval to the exact tool, arguments, environment, and artifactModel–harness fitThe same driver performs differently with another car, tires, and pit crewBenchmark the model and harness togetherEvidence-based reviewA doctor uses lab results instead of relying only on “I feel fine”Give reviewers tests, traces, and verified stateState verificationA delivery app says “delivered,” but the package still needs to be at the doorDo not trust HTTP 200 or tool success without checking stateLeast privilegeA hotel keycard opens one room for a limited timeGrant only the authority required by the current taskReproducible environmentA recipe without oven temperature cannot be repeated reliablyPin the JDK, dependencies, images, and build inputsContinuous SRE evaluationAn emergency room handles overlapping cases, not one practice patientTest agents under concurrent incidents and changing stateMachine-actionable errors“Turn right in 200 meters” is more useful than “find another route”Return a specific recovery tool and retry action
These comparisons make the architecture easier to explain to product owners, managers, and security teams without beginning with LLM, MCP, or harness terminology.
A practical experiment for next week
Select eight completed Spring Boot or Kubernetes tickets and add an Intent-and-Authority Gate:
Generate two plausible interpretations of every underspecified requirement.
Ask for clarification only when they change observable API, database, event, authorization, or cluster behavior.
Let the agent propose the patch inside a least-privilege sandbox.
Run Maven/JUnit, Testcontainers, ArchUnit, API compatibility, SAST, migration, and Kubernetes policy checks outside the agent’s control.
Bind approval to the commit, artifact digest, exact tool arguments, namespace, policy version, and expiry.
Verify final state and emit a reproducible release receipt.
Repeat each task three times and measure
pass³, unnecessary questions, unsupported assumptions, false rejection, approval mismatch, elapsed time, and cost.
Closing thought
A production-ready coding agent is not one that never fails. It is a system in which failures are detected before they cause harm, authority is bounded, evidence is preserved independently, release decisions are reproducible, and final state is verified instead of inferred from confident model narratives.
More capable models will generate software faster. Trustworthy production delivery will come from specification, executable evidence, least privilege, approval binding, and deterministic release gates.
Do not place trust in the model. Build a system that can prove when the model deserves trust.
🇻🇳 Phiên bản tiếng Việt
AI coding agent đang viết code nhanh hơn, nhưng “nhanh” không đồng nghĩa với “đáng tin”
Trong vài tháng gần đây, AI coding agent đã tiến rất xa. Agent có thể đọc repository, sửa nhiều file, chạy test, tạo pull request, sửa lỗi build, thậm chí thao tác với cloud và Kubernetes.
Nhưng khi đưa agent vào quy trình phát triển phần mềm enterprise, câu hỏi quan trọng nhất không còn là:
“Model nào code giỏi nhất?”
Mà phải là:
“Hệ thống nào có thể chứng minh thay đổi này đúng, an toàn, được cấp quyền và có thể tái hiện?”
Các nghiên cứu mới nhất về coding agent cho thấy bốn vấn đề rất thực tế:
Agent có thể tự hiểu một yêu cầu chưa rõ ràng rồi triển khai sai hành vi người dùng mong muốn.
Hành động người dùng phê duyệt có thể không hoàn toàn giống hành động cuối cùng được thực thi.
Cùng một model nhưng chạy qua harness khác nhau có thể cho kết quả khác biệt rất lớn.
Test hoặc benchmark có thể báo thành công dù trạng thái thật của database, Git hay Kubernetes chưa đúng.
Vì vậy, một AI agent production-ready không thể chỉ gồm model + prompt + tools. Nó cần một lớp kiểm soát độc lập bao quanh model.
Hãy hình dung agent như một nhà thầu sửa nhà
Giả sử bạn thuê một nhà thầu rất giỏi và làm việc cực nhanh để sửa lại căn nhà.
Bạn nói: “Hãy làm phòng bếp rộng hơn, hiện đại hơn và an toàn hơn.” Nhà thầu hiểu rất nhanh, biết sử dụng nhiều công cụ và có thể bắt tay vào làm ngay. Nhưng trước khi họ đập tường, vẫn còn hàng loạt câu hỏi:
“Rộng hơn” là thêm bao nhiêu mét?
Bức tường cần đập có phải tường chịu lực không?
Ngân sách tối đa là bao nhiêu?
Nhà thầu được vào phòng nào và được dùng chìa khóa nào?
Ai kiểm tra điện, nước và kết cấu sau khi hoàn thành?
Nếu người thợ tự đoán tất cả rồi làm thật nhanh, căn bếp có thể trông rất đẹp nhưng lại không đúng nhu cầu, vượt ngân sách hoặc làm yếu kết cấu căn nhà. AI coding agent cũng tương tự: viết được code chạy không có nghĩa là đã xây đúng thứ cần xây.
Mapping giữa câu chuyện sửa nhà và hệ thống AI agent:
Đời thườngTrong AI software deliveryYêu cầu sửa nhàBusiness requirement hoặc Jira ticketBản thiết kế có kích thướcSpecification và acceptance criteriaNhà thầuAI coding agentBộ dụng cụTools, MCP server, shell, Git, cloud APIChìa khóa chỉ mở một số phòngLeast-privilege permissionsGiấy phép xây dựngPolicy và approvalThanh tra điện, nước, kết cấuTest, security scan và architecture gateNghiệm thu căn nhàRuntime/state verificationBiên bản bàn giaoEvidence-bound release receipt
Điểm mấu chốt: chúng ta không từ chối dùng một nhà thầu giỏi. Chúng ta cho họ bản thiết kế rõ ràng, đúng chìa khóa, phạm vi công việc cụ thể và một quy trình nghiệm thu độc lập.
Kiến trúc nên hướng tới
```mermaid
flowchart TD
A["Yêu cầu + tiêu chí chấp nhận"] --> B{"Có nhiều cách hiểu làm đổi hành vi?"}
B -->|Có| C["Làm rõ yêu cầu"]
B -->|Không| D["Agent tạo kế hoạch và bản vá"]
C --> D
D --> E["Sandbox quyền tối thiểu"]
E --> F["Build, test, security và policy gates"]
F --> G{"Bằng chứng đầy đủ?"}
G -->|Không| H["Sửa lại hoặc chuyển người review"]
H --> D
G -->|Có| I["Approval ràng buộc đúng hành động"]
I --> J["Deploy có kiểm soát"]
J --> K["Xác minh trạng thái thật + release receipt"]
```Kiến trúc này tách bốn trách nhiệm:
Agent tạo ứng viên: plan, code, test và đề xuất hành động.
Verifier thu thập bằng chứng độc lập từ build, test, static analysis và trạng thái runtime.
Policy engine quyết định agent được phép làm gì trong phạm vi nào.
Release controller chỉ triển khai artifact đã được xác minh và lưu lại receipt có thể audit.
Điểm quan trọng là agent không tự chấm bài của chính nó và cũng không tự cấp quyền release cho chính nó.
1. Chỉ hỏi lại khi sự mơ hồ làm thay đổi hành vi
Nghiên cứu CONTRA đưa ra một cách tiếp cận rất thực tế: tạo hai câu trả lời hợp lý cho điểm chưa rõ, sinh chương trình tương ứng rồi kiểm tra xem hành vi có thực sự khác nhau hay không. Phương pháp này cải thiện F1 trung bình 13,88 điểm phần trăm so với baseline tốt nhất trên ClarifyCodeBench.
Ví dụ, ticket ghi:
“Khóa account sau 5 lần payment thất bại.”
Một developer có kinh nghiệm sẽ hỏi thêm:
Năm lần trong bao lâu: một ngày, năm ngày hay toàn bộ lịch sử?
Chỉ tính hard reject hay cả soft reject?
Khóa payment flow hay khóa toàn bộ portal?
Thành công ở lần tiếp theo có reset bộ đếm không?
Hai request đồng thời có được tính chính xác không?
Nếu mỗi cách trả lời tạo ra trạng thái database hoặc API response khác nhau, agent phải dừng và yêu cầu làm rõ. Nếu khác biệt chỉ là tên biến hoặc cách tổ chức code, agent không cần làm phiền người dùng.
Áp dụng vào AI Agent Kit: thêm một clarification gate trước implement-feature. Gate phải tạo được ít nhất hai interpretation, chỉ ra hành vi nào thay đổi, và gắn mỗi câu hỏi với test hoặc state transition cụ thể.
2. Approval phải ràng buộc với đúng hành động được thực thi
Approval Laundering mô tả sáu cách hành động được duyệt có thể bị thay đổi trước khi thực thi: thay đổi scope, argument, thời điểm, tool, delegation hoặc ý nghĩa thực tế. Trong 118 lần replay, approval token có chữ ký ngăn được một số nhóm lỗi nhưng không thể phát hiện thay đổi nằm dưới lớp tool call mà mediator không nhìn thấy.
Ví dụ, người dùng duyệt:
Deploy image rite-online:1.8.4 vào namespace testNhưng giữa lúc duyệt và lúc chạy, một trong các yếu tố sau thay đổi:
tag
1.8.4trỏ tới digest khác;kube-context chuyển từ
testsangprod;Helm values bổ sung quyền cluster-wide;
một sub-agent khác nhận lại quyền và chạy lệnh tương đương nhưng scope rộng hơn.
Một nút Approve chung chung không đủ an toàn. Approval nên ràng buộc tối thiểu với:
principal + agent_id + session_id + tool + normalized_arguments
+ namespace + commit_sha + image_digest + policy_version + expirySau khi chạy, hệ thống vẫn phải kiểm tra kết quả thực tế: resource nào được tạo, image digest nào đang chạy, database migration nào đã commit và quyền nào đã thay đổi.
3. Đừng chọn model riêng lẻ hãy chọn cả model và harness
Finding the Right Fit đánh giá 66 cấu hình model–harness và công bố 6.204 trajectory. Trên cùng một benchmark, Claude dẫn GPT 7,94 điểm với một harness nhưng lại thua 30,16 điểm khi đổi sang harness khác. Một cấu hình GPT gọn hơn còn đạt kết quả cao hơn với chi phí dưới một phần tư.
Điều này đặc biệt quan trọng nếu đội ngũ đang dùng nhiều model như GPT, Claude, Qwen, DeepSeek hoặc Kimi. Không nên kết luận “model A tốt hơn model B” từ một lần demo.
Hãy đo theo tuple đầy đủ:
model + harness + task type + repository type + tool contract + verification gatesVí dụ, một profile mạnh cho sửa bug Spring Boot chưa chắc tốt cho:
migration Java 8 lên Java 21;
sửa stored procedure thành service;
tạo module mới từ specification;
xử lý incident Kubernetes;
thay đổi đồng thời nhiều repository.
Trong .ai/, nên có profile riêng cho fix-bug, implement-feature, database-change, production-incident và repository-intelligence. Mỗi profile cần benchmark, ngân sách token, tools và gates riêng.
4. Reviewer nhỏ vẫn hữu ích nếu có bằng chứng tốt
Groundability, Not Scale Alone cho thấy reviewer nhỏ hơn có thể đánh giá agent mạnh hơn nếu được cung cấp bằng chứng thực thi có tính quyết định. Nhưng khi bằng chứng không đầy đủ, hệ thống review tự động bắt được 76–80% lỗi đồng thời reject nhầm tới 66–67%.
Thông điệp thực tế không phải là “dùng model nhỏ để review mọi thứ”. Thông điệp là:
Chất lượng của bằng chứng quan trọng hơn kích thước reviewer, và reviewer phải được phép trả lời “chưa đủ bằng chứng”.
Với Spring Boot, evidence package nên gồm:
Khu vựcBằng chứng nên cóFunctional behaviorJUnit, integration test, contract testDatabaseTestcontainers, migration validation, before/after stateArchitectureArchUnit, dependency-boundary checksAPIOpenAPI diff, backward-compatibility checkSecuritySAST, dependency scan, authorization-path testsKubernetesSchema validation, policy checks, rendered manifest diffRuntimeTrace, log assertion, health/readiness evidence
Nếu một mandatory gate chưa có kết quả, tổng điểm cao không được phép bù cho nó. Hệ thống phải retry, chuyển người review hoặc dừng.
5. Đừng tin tool báo “success”; hãy kiểm tra state thật
Do Agent Benchmarks Do What They Say? tìm thấy bảy lỗi tool trong 34 mutating tools thuộc bốn benchmark. Có evaluator cho điểm thành công dựa trên message của tool mặc dù trạng thái đáng lẽ phải thay đổi lại không thay đổi đúng.
Đây là lỗi rất dễ xuất hiện trong hệ thống enterprise:
API trả HTTP 200 nhưng transaction sau đó rollback;
agent chạy
kubectl applythành công nhưng deployment không ready;Maven build pass vì module cần kiểm tra không được include;
migration tool báo success nhưng chạy nhầm schema;
event publish thành công nhưng consumer không thể deserialize payload.
Quy tắc đơn giản là:
Tool output chỉ là một claim. Database, Git, cluster và runtime state mới là evidence.
6. Kubernetes agent phải được test trong incident kéo dài, không phải một ticket cô lập
SRE-Marathon cho agent vận hành Kubernetes hai zone với khoảng 60 lỗi chồng chéo trong mỗi lần chạy. Phương pháp tốt nhất chỉ đạt 41,3/100; agent thường tìm ra nguyên nhân nhưng hiếm khi hoàn tất repair khi lỗi còn đang xảy ra.
Một platform agent thực tế cần xử lý đồng thời:
alert cũ và mới;
rollout vừa diễn ra;
config thay đổi;
dependency bên ngoài chập chờn;
nhiều người hoặc agent khác đang thao tác;
trạng thái được giữ lại qua nhiều vòng điều tra.
Vì vậy, bài test “pod crash → agent restart pod” chưa đủ. Cần đo toàn bộ chuỗi:
```mermaid
flowchart LR
A["Alert correlation"] --> B["Fault localization"]
B --> C["Safe mitigation"]
C --> D["Service restoration"]
D --> E["Recurrence prevention"]
```7. Những chi tiết nhỏ của interface có thể tạo khác biệt lớn
MCP Error Messages Written for Developers Hurt Agents kiểm tra 3.001 error message từ 150 MCP server. Nhiều message yêu cầu agent mở browser, sửa config hoặc “đợi rồi thử lại” dù agent không có khả năng đó. Khi error trả về đúng recovery tool, tỷ lệ phục hồi lỗi credential tăng lên 84% và rate limit lên 88%.
Thay vì trả:
{"error": "Authentication failed. Run login in your terminal."}Nên trả một contract máy có thể xử lý:
{
"code": "AUTH_EXPIRED",
"retryable": true,
"recoveryTool": "refresh_session",
"retryTool": "get_account",
"correlationId": "req-8f31"
}Thiết kế tool tốt giúp model rẻ hơn và nhỏ hơn hoạt động ổn định hơn; không phải mọi vấn đề đều cần nâng cấp model.
Chín cách liên tưởng nhanh để giải thích cho team
Khái niệm kỹ thuậtVí dụ đời thường dễ hình dungCách áp dụngClarification gateThợ may phải hỏi số đo trước khi cắt vảiLàm rõ yêu cầu trước khi agent sửa code khó hoàn tácApproval bindingTờ séc phải khóa đúng người nhận và số tiềnApproval phải khóa đúng tool, argument, môi trường và artifactModel–harness fitCùng một tay đua nhưng xe, lốp và pit crew khác nhau sẽ cho kết quả khácBenchmark cả model lẫn harness, không chỉ modelEvidence-based reviewBác sĩ dùng xét nghiệm, không chỉ nghe bệnh nhân nói “tôi thấy khỏe”Reviewer cần test, trace và state thậtState verificationỨng dụng giao hàng báo “đã giao” nhưng phải kiểm tra gói hàng có ở cửa khôngKhông tin HTTP 200 hoặc tool success nếu state chưa đúngLeast privilegeThẻ khách sạn chỉ mở đúng phòng và đúng thời gian lưu trúAgent chỉ nhận quyền cần thiết cho task hiện tạiReproducible environmentCông thức nấu ăn thiếu nhiệt độ lò sẽ không thể lặp lại món ănKhóa JDK, dependency, container image và build inputsContinuous SRE evaluationCấp cứu bệnh viện phải xử lý nhiều ca cùng lúc, không phải một bệnh nhân luyện tậpTest agent với incident chồng chéo và trạng thái thay đổiMachine-actionable errorGPS nói “rẽ phải sau 200 m” hữu ích hơn “hãy tìm đường khác”Error phải chỉ rõ recovery tool và hành động retry
Những phép so sánh này giúp giải thích với product owner, manager hoặc security team mà không cần bắt đầu bằng thuật ngữ LLM, MCP hay agent harness.
Một engineering practice đáng thử ngay tuần tới
Chọn tám ticket Spring Boot hoặc Kubernetes đã hoàn thành và chạy thử Intent-and-Authority Gate:
Tạo hai interpretation hợp lý cho mỗi yêu cầu chưa rõ.
Chỉ hỏi lại nếu hai interpretation làm thay đổi API, database, event, authorization hoặc cluster state.
Cho agent tạo patch nhưng chạy trong sandbox quyền tối thiểu.
Dùng gate độc lập để chạy Maven/JUnit, Testcontainers, ArchUnit, OpenAPI diff, SAST, migration và Kubernetes policy checks.
Approval phải gắn với commit SHA, artifact/image digest, tool arguments, namespace, policy version và thời hạn.
Sau khi chạy, kiểm tra state thật và tạo một release receipt không cần gọi lại model vẫn có thể tái hiện quyết định.
Chạy mỗi task ba lần và đo
pass³, câu hỏi thừa, assumption không được xác nhận, false rejection, approval mismatch, thời gian và chi phí.
Kết luận
AI coding agent production-ready không phải là agent không bao giờ sai. Đó là một hệ thống mà khi agent sai:
sai lệch được phát hiện trước khi tạo hậu quả;
quyền hạn bị giới hạn;
bằng chứng được lưu độc lập;
quyết định release có thể audit và tái hiện;
trạng thái thật được kiểm tra thay vì tin vào lời model hoặc tool.
Model ngày càng mạnh sẽ giúp tạo code nhanh hơn. Nhưng khả năng đưa code vào production một cách đáng tin sẽ đến từ specification, executable evidence, least privilege, approval binding và deterministic release gates.
Nói ngắn gọn:
Đừng trao niềm tin cho model. Hãy xây dựng một hệ thống có thể chứng minh khi nào model đáng được tin.
Sources
- Approval Laundering: Approval–Execution Binding Failures
- CONTRA: Selective Clarification in LLM Code Generation
- SRE-Marathon: Continuous Evaluation of Autonomous SRE Agents
- Groundability, Not Scale Alone
- Finding the Right Fit: Model–Harness Interactions
- Do Agent Benchmarks Do What They Say?
- Aletheia: Permission-Minimality Testing for Coding-Agent Rules
- Code That Works, Environments That Don’t
- MCP Error Messages Written for Developers Hurt Agents
- Path2Spec: Path-Aware Specification Generation
Discussion
Newest firstLoading comments…
Comments appear after review.