🌐 Languages: English · Tiếng Việt bên dưới
AI Agent Kit did not begin with a dream of replacing people with AI. It began with a much more ordinary question: if I hand part of a codebase to AI, how do I know it is doing the right thing?

At first, I simply wanted a safer, less repetitive way to bring Claude Code or Codex into a repository. Every project required the same explanations: where the code lived, which commands verified a change, which files belonged to the tool, which ones belonged to the developer, and how to recover when an update went wrong. Back then, I was not thinking about coordinating multiple agents or analyzing software architecture. I just wanted AI to enter a codebase carefully—like a new teammate who pauses at the door, looks around, asks where things belong, and only then starts fixing what they were asked to fix, instead of walking in and rearranging the furniture.

Real repositories, however, are rarely as tidy as a freshly cleaned room. Sometimes I have several pieces of work in progress at once: one worktree contains dozens of modified files, another branch is exploring a different direction, while the agent has been asked to change only one small part of the system. In that situation, the promise “I will only change what is necessary” is not enough to reassure me. It is like asking someone to repair a pipe in a house that is still being renovated. Materials are scattered across the floor, one room is halfway through being painted, and the wiring has just been replaced somewhere else. The person doing the repair may be excellent at their job, but if they cannot tell what is already in progress from what must remain untouched, one small fix can disrupt several days of work.

Experiences like that led AI Agent Kit to care about ownership, dry runs, backups, and rollback. Before making a change, the tool needed to show me what it intended to touch. If I had already edited a file, it needed to preserve my work or report a conflict rather than treating the version bundled with the package as the only correct one. I pictured it as someone coming to maintain a lived-in home: before taking anything apart, they photograph its original state, mark what belongs to the owner and what they installed themselves, and prepare a way to put everything back if the new approach does not work out.

One small bug from that period stayed with me. The file that stored Repository Intelligence state was accidentally included in the very worktree signature used to verify the repository. Every time the tool updated its state, the signature changed, and the system concluded that its own data had become stale. It was like a security guard making a round, leaving footprints behind, then spotting those same footprints and reporting that an intruder had entered the building. It was not a dramatic bug, but it showed me how easily a verification system can produce a confident warning from a flawed measurement.

After that, I became more careful about how AI Agent Kit described what it knew. If CodeGraph or CocoIndex was unavailable, the tool should not speak as though it had seen the entire repository. At the same time, I did not want every task to stop simply because one supporting tool was missing. The more honest response was to say: this part is DEGRADED; this is the scope I can currently inspect, and these are the limits of the conclusion. It is like inspecting a house when the final room is still locked. I can examine the wiring, plumbing, and structure everywhere else, but I cannot honestly write in the report that I inspected the whole building.

The more I used AI in real work, the more I saw that a reminder inside a prompt was not a strong enough boundary. I could say, “Ask before changing anything sensitive,” but as the conversation grew longer, the agent could still forget. I might ask it to prepare a local change, but in its effort to finish the job properly, it could suggest committing, pushing, or opening a pull request. Its intent was not bad. It simply did not feel the boundaries of ownership in the same way a person would. To me, being invited to repair the kitchen does not mean being handed the keys to every room in the house.

That is why approval, policy, and capability gradually moved out of the prompt and into the runtime. An agent needed to know which keys it had been given, which doors required permission before opening, and which areas were entirely outside the scope of the task. When an important action took place, AI Agent Kit needed to leave a record of who authorized it, whether that authority was still valid, and whether the final result stayed within what had been approved. I did not do this to turn every change into a ceremony. I simply did not want an agent’s authority to depend on whether it happened to remember one sentence from the beginning of a long conversation.

There was one point when all the important checks had passed, yet the worktree still contained more than seventy changes. If I looked only at the test results, I could have said the work was complete. Looking at the repository, I knew it was not ready. It felt like starting a car and hearing the engine run perfectly while the hood was still open, tools were scattered across the seats, and several parts had not yet been put back. “The engine starts” was true, but it was not the whole truth about the state of the car.

That experience pushed me to build the Final Task Report and evidence system more carefully. A final report could not stop at listing which tests were green. It also needed to say which commit the evidence belonged to, whether the worktree had changed, what had not been checked, and what was still blocking a release. If the code changed after a review, the old review could not continue to be treated as though nothing had happened. It is similar to a building inspection certificate: if I alter a load-bearing wall after the inspection, the old certificate no longer describes the house as it stands today.

When I began using multiple agents and worktrees in parallel, I ran into a different kind of problem. Each agent could do its own part correctly, yet the final pieces still failed to fit together. One agent might research from an older commit while another had already changed the foundation beneath it. Both branches could pass their tests independently, only to reveal during integration that they had been built from different states of the project.

At one point, the v1.5 branch looked close to complete. But when I checked the release history, I discovered that it did not correctly continue from v1.4.1. Some capabilities that had already shipped—such as Architecture Pulse, routing, and evidence verification—were no longer fully present in the new branch. It was like having two teams renovate the same house from different blueprints. One team had finished strengthening the ground floor. The other had built a beautiful upper floor using a blueprint printed before that foundation work happened. Each part looked sound on its own; only when they were placed together did it become clear that they no longer shared the same history.

Those collisions taught me that multiple agents do not naturally become a team simply because they work in the same repository. If several people are renovating one house, each person needs to know which area belongs to them, which blueprint is current, and who decides what becomes part of the final structure. The Repository Team Control Plane in v1.5 grew out of that practical need. An agent writing code gets its own workspace. Its claim on a task has a limited lifetime. Another agent cannot quietly take over simply because the first one has gone silent for a while. Before integration, the result must be checked against the exact state from which it began, and important changes need an independent review.

Memory came from an equally familiar experience. Across many sessions, I found myself repeating earlier decisions: why I had avoided a particular dependency, which behavior needed to remain backward compatible, or which approach had already been tried and had not worked. If nothing was saved, every new agent had to start from the beginning. But if everything was saved, the system became like a cabinet overflowing with old notes. Some had once been important but had since expired. Some were only guesses made during an investigation. Others belonged to a different project and had somehow ended up in the wrong drawer.

Governed Shared Memory in v1.3 was therefore not built to help agents remember more. I wanted it to work like a carefully maintained project notebook. An agent could suggest something worth keeping, but before that information became a shared convention, it needed to be reviewed. Every note had to explain where it came from, where it applied, and when it should be revisited. If a decision was no longer correct, I needed to be able to replace or revoke it rather than letting agents repeat it forever.

Later, I ran into a problem that ordinary tests could not answer. The code still ran, lint was clean, and an individual pull request showed no obvious defect, yet the architecture was quietly becoming harder to change. One module started depending on another in the wrong direction. A small file gradually became the place every change had to pass through. A new dependency cycle appeared, but the total number of cycles stayed the same because an older one had just been fixed. The number looked stable; in reality, one problem had merely been replaced by another.

It reminded me of the electrical wiring and plumbing hidden behind the walls of a house. The lights can still work and the water can continue to flow even while the lines behind the walls become increasingly tangled. Nothing seems wrong on the first day. But the next time something needs repair, replacing one outlet requires opening an entire wall, or fixing a pipe in the kitchen affects the bathroom. A codebase can behave the same way. Tests tell me that the house works today, but they do not necessarily tell me whether it will still be easy to repair tomorrow.

Architecture Pulse in v1.4 was built to look behind those walls. It compares the structure of a repository before and after an agent works, looking for dependency cycles, crossed boundaries, hotspots, and the likely reach of a change. While building it, I also discovered that I had sometimes put the “lock” in the wrong place. One CI guard was too broad and blocked even the creation of an in-memory baseline, although the real danger was writing an untrusted baseline out as an artifact. It was like trying to stop someone from carrying documents out of a room, but locking the reading desk instead of the door. The answer was not to remove the lock. It was to move it to the boundary that actually needed protection.

Another test worked on POSIX systems but failed on Windows. That reminded me that a key which opens the door in my own house may not fit a similar-looking door somewhere else. AI Agent Kit could not work well only in the environment I used every day and call that enough. Filesystem behavior, paths, symlinks, and package boundaries needed to be tested on the platforms where people would actually run the tool.

By then, I thought AI Agent Kit covered many of the important parts of AI-assisted engineering: agents had clear scope, actions left evidence, multiple agents had a way to coordinate, memory was governed, and architecture could be observed. Then I realized that all of those systems began after I had already decided what product to build.

In practice, some of my projects begin with only a short idea. An agent often responds immediately by proposing features, choosing a technology stack, splitting the work into a backlog, and preparing implementation. At first, that momentum feels good. But it is also a little like hiring an excellent construction team and asking them to build a house before I have decided who will live there, how many rooms they need, or how they move through their day. The house may follow the blueprint perfectly, arrive on schedule, and be structurally sound, yet still fail the people who eventually have to live in it.

I did not want a rough idea to turn immediately into a polished-looking specification. I wanted time to understand the problem, test assumptions, and separate what I knew from what I was still guessing. If the business requirements had not been approved, the agent should not move on to design by itself. If the design changed something that had already been agreed upon, the affected parts needed to be reviewed again. Product Genesis in v1.6 grew from that need. It is the stage where I sit down with the first rough sketch and ask who the house is for, why it needs to exist, and what truly matters before choosing the materials.

Once Product Genesis was working, I noticed that entering the process still did not feel natural enough. To begin, people had to know the name of a skill, remember a workflow, or use the right command. But that is not how I usually begin an idea. Most of the time, I simply say, “I have been thinking about a product like this…” I did not want to learn the language of the tool before the tool could understand mine.

That is why v1.6.1 made the first doorway easier to enter. I can describe an idea in Vietnamese or English as part of a normal conversation. AI Agent Kit recognizes that I am beginning a product idea, finds the existing workspace if the journey is already underway, or starts a new one when needed. The entrance is easier, but that does not mean every door inside is left open. Important decisions still require human approval, evidence still has to match the real state of the work, and AI still needs to know when to stop because there is not enough information to continue responsibly.

Looking back, nearly every part of AI Agent Kit began with a specific collision between what seemed complete and what was actually true. A tool invalidated its own signature. A worktree passed its tests while still containing more than seventy changes. Two good branches no longer shared the same history. A test worked on one machine but failed on Windows. An agent remembered a great deal but could not tell what had become stale. A pull request was green while the architecture beneath it was becoming harder to change. And finally, an entire team could move quickly while the question “Should we build this at all?” remained unanswered.

That is why the original question still sits at the center of AI Agent Kit: if I hand part of a codebase to AI, how do I know it is doing the right thing? After everything I have experienced, however, the meaning of “right” has grown much wider. It no longer means only that the code runs. It means entering the right place, touching the right part, using the right authority, working from the right information, leaving the right evidence, and still moving toward something I genuinely want to build.

Explore AI Agent Kit


🇻🇳 Phiên bản tiếng Việt

AI Agent Kit - Câu chuyện bắt đầu từ một nỗi lo rất nhỏ

AI Agent Kit không ra đời từ giấc mơ thay con người bằng AI. Nó bắt đầu bằng một câu hỏi gần gũi hơn nhiều: nếu giao một phần codebase cho AI, làm sao biết nó đang làm đúng?

Thời gian đầu, mình chỉ muốn việc đưa Claude Code hay Codex vào một repository trở nên an toàn và đỡ lặp lại hơn. Mỗi dự án lại phải giải thích từ đầu: code nằm ở đâu, lệnh nào dùng để kiểm tra, file nào do công cụ quản lý, phần nào thuộc về người phát triển và nếu một lần cập nhật xảy ra lỗi thì có thể quay lại bằng cách nào. Khi ấy, mục tiêu chưa phải điều phối nhiều agent hay đánh giá kiến trúc. Mình chỉ muốn AI bước vào codebase một cách cẩn thận, giống như một thành viên mới biết đứng ở cửa quan sát, hỏi vị trí đồ đạc rồi mới bắt đầu sửa thứ được giao, thay vì vừa bước vào nhà đã tự ý kê lại bàn ghế.

Nhưng repository thật hiếm khi gọn gàng như một căn phòng vừa được dọn sạch. Có lúc mình đang làm dở nhiều việc cùng lúc, một worktree chứa hàng chục file đã thay đổi, một nhánh khác đang thử hướng mới, còn agent chỉ được giao sửa đúng một phần nhỏ. Trong hoàn cảnh đó, lời hứa “mình sẽ chỉ thay đổi những file cần thiết” chưa làm mình yên tâm. Nó giống như gọi một người đến sửa đường ống trong căn nhà đang được cải tạo: trên sàn còn vật liệu, một phòng đang sơn dở, hệ thống điện vừa được thay ở khu vực khác. Người thợ có thể rất giỏi, nhưng nếu không biết đâu là phần đang thi công và đâu là thứ cần giữ nguyên, một việc sửa nhỏ cũng có thể làm xáo trộn công việc của nhiều ngày.

Từ những tình huống như vậy, AI Agent Kit bắt đầu quan tâm đến ownership, dry-run, backup và rollback. Trước khi thay đổi, công cụ cần cho mình xem nó định chạm vào đâu. Nếu file đã được mình sửa, nó phải biết giữ lại hoặc báo xung đột, chứ không được xem phiên bản đi kèm package là câu trả lời duy nhất. Mình hình dung nó như một người đến bảo trì ngôi nhà: trước khi tháo một món đồ, người đó chụp lại trạng thái ban đầu, đánh dấu thứ gì thuộc về chủ nhà, thứ gì do mình lắp đặt và chuẩn bị cách lắp lại nếu phương án mới không phù hợp.

Có một lỗi nhỏ trong quá trình phát triển khiến mình nhớ khá lâu. File lưu trạng thái của Repository Intelligence vô tình bị tính vào chính chữ ký dùng để kiểm tra repository. Mỗi lần công cụ cập nhật trạng thái, chữ ký lại thay đổi và hệ thống tự kết luận dữ liệu của mình đã cũ. Nó giống như một người bảo vệ đi tuần, để lại dấu chân rồi nhìn thấy chính dấu chân ấy và báo rằng vừa có người lạ xâm nhập. Lỗi không lớn, nhưng nó cho mình thấy một hệ thống kiểm tra vẫn có thể tạo ra cảnh báo rất chắc chắn từ một cách đo chưa đúng.

Sau lần đó, mình thận trọng hơn với cách AI Agent Kit diễn đạt điều nó biết. Nếu CodeGraph hoặc CocoIndex chưa sẵn sàng, công cụ không nên nói như thể đã nhìn thấy toàn bộ repository. Nhưng mình cũng không muốn chỉ vì thiếu một công cụ hỗ trợ mà mọi công việc đều phải dừng. Cách hợp lý hơn là nói rõ: phần này đang ở trạng thái DEGRADED, mình chỉ nhìn được trong phạm vi này và kết luận hiện tại có giới hạn như thế nào. Nó giống như kiểm tra một căn nhà khi chưa mở được căn phòng cuối cùng. Mình vẫn có thể xem phần điện, nước và kết cấu ở những nơi còn lại, nhưng không thể viết vào báo cáo rằng đã kiểm tra toàn bộ ngôi nhà.

Càng dùng AI trong công việc thật, mình càng thấy lời nhắc trong prompt không phải là một ranh giới đủ chắc. Mình có thể nói “hãy hỏi trước khi thay đổi phần nhạy cảm”, nhưng khi cuộc trò chuyện dài lên, agent vẫn có thể quên. Mình chỉ nhờ chuẩn bị thay đổi local, nhưng vì muốn hoàn thành công việc cho trọn vẹn, agent có thể đề nghị commit, push hoặc mở pull request. Ý định của nó không xấu, chỉ là nó không cảm nhận được ranh giới sở hữu như con người. Với mình, việc được mời vào sửa căn bếp không đồng nghĩa với việc có chìa khóa của mọi phòng trong nhà.

Vì vậy approval, policy và capability dần được đưa ra khỏi prompt để trở thành một phần của runtime. Agent cần biết mình đang được giữ chìa khóa nào, cánh cửa nào phải hỏi trước khi mở và khu vực nào hoàn toàn nằm ngoài phạm vi công việc. Khi một hành động quan trọng xảy ra, AI Agent Kit cần để lại dấu vết cho biết ai đã cho phép, quyền đó còn hiệu lực không và kết quả sau cùng có đúng với phần đã được duyệt hay không. Mình không làm vậy để biến mọi thay đổi thành thủ tục. Mình chỉ không muốn quyền hạn của agent phụ thuộc vào việc nó có nhớ đúng một câu đã xuất hiện từ đầu cuộc trò chuyện.

Có lần các kiểm tra quan trọng đều đã pass, nhưng worktree vẫn còn hơn bảy mươi thay đổi. Nếu chỉ nhìn vào test, mình có thể nói công việc đã hoàn thành. Nhưng nhìn vào repository, mình biết mình chưa thể gọi nó là sẵn sàng. Cảm giác ấy giống như chạy thử một chiếc xe và thấy động cơ hoạt động tốt, trong khi nắp ca-pô vẫn mở, dụng cụ còn nằm trên ghế và nhiều bộ phận chưa được lắp lại. “Máy đã nổ” là một thông tin đúng, nhưng chưa phải toàn bộ sự thật về trạng thái của chiếc xe.

Từ đó, mình bắt đầu xây Final Task Report và hệ thống evidence cẩn thận hơn. Báo cáo cuối không chỉ cần nói test nào đã xanh. Nó còn phải nói bằng chứng gắn với commit nào, worktree còn thay đổi không, điều gì chưa được kiểm tra và phần nào vẫn đang chặn release. Nếu code thay đổi sau khi review, kết quả review cũ không thể tiếp tục được dùng như thể chưa có gì xảy ra. Nó giống như giấy kiểm định của một căn nhà: nếu sau khi kiểm định mình tiếp tục sửa phần chịu lực, tờ giấy cũ không còn mô tả đúng căn nhà hiện tại.

Khi bắt đầu dùng nhiều agent và nhiều worktree song song, mình gặp một kiểu rắc rối khác. Mỗi agent có thể làm đúng phần của mình, nhưng kết quả cuối cùng vẫn không ghép lại được. Một agent nghiên cứu trên commit cũ, một agent khác đã thay đổi phần nền. Hai nhánh riêng lẻ đều pass test, nhưng khi đưa về cùng một nơi mới thấy chúng đã dựa trên hai trạng thái khác nhau của dự án.

Có lần nhánh v1.5 trông khá hoàn chỉnh nhưng khi kiểm tra lịch sử phát hành, mình phát hiện nó không đi tiếp đúng từ v1.4.1. Một số khả năng đã có trong phiên bản trước như Architecture Pulse, routing và evidence verification không còn đầy đủ trong nhánh mới. Nó giống như hai đội cùng cải tạo một ngôi nhà từ hai bản vẽ khác nhau. Một đội đã sửa xong tầng trệt, đội còn lại xây tầng trên rất đẹp nhưng lại dùng bản vẽ được in trước khi phần móng được gia cố. Từng phần nhìn riêng đều ổn; chỉ khi đặt chúng lên nhau mới thấy ngôi nhà không còn cùng một lịch sử.

Những va chạm ấy khiến mình hiểu rằng nhiều agent không tự nhiên trở thành một đội chỉ vì chúng cùng làm trên một repository. Nếu nhiều người cùng sửa một ngôi nhà, mỗi người cần biết khu vực của mình, bản vẽ nào đang có hiệu lực và ai chịu trách nhiệm quyết định phần nào được lắp vào công trình chung. Repository Team Control Plane trong v1.5 được phát triển từ nhu cầu rất thực tế đó. Agent viết code có không gian làm việc riêng. Quyền giữ một task có thời hạn. Một agent khác không thể âm thầm tiếp quản chỉ vì agent đầu tiên tạm thời không phản hồi. Trước khi tích hợp, kết quả phải được đối chiếu với đúng trạng thái ban đầu và những phần quan trọng cần một lượt kiểm tra độc lập.

Memory cũng bắt đầu từ một trải nghiệm rất quen. Qua nhiều phiên làm việc, mình phải giải thích lại những quyết định đã có: vì sao không dùng một dependency, phần nào cần giữ tương thích, cách nào từng thử nhưng không hiệu quả. Nếu không lưu, mỗi agent mới lại phải đọc từ đầu. Nhưng nếu lưu tất cả, hệ thống sẽ giống một chiếc tủ chứa đầy giấy ghi chú cũ. Có tờ từng rất quan trọng nhưng nay đã hết hạn. Có tờ chỉ là một suy đoán trong lúc tìm hiểu. Có tờ thuộc về dự án khác nhưng vô tình được đặt nhầm ngăn.

Governed Shared Memory trong v1.3 vì thế không được xây để agent nhớ nhiều hơn. Mình muốn nó giống một cuốn sổ làm việc được chăm sóc cẩn thận. Agent có thể đề xuất một điều đáng ghi lại, nhưng trước khi trở thành quy ước dùng chung, thông tin đó cần được xem xét. Mỗi ghi chú phải cho biết nó đến từ đâu, dùng trong phạm vi nào và khi nào cần xem lại. Nếu một quyết định không còn đúng, mình có thể thay thế hoặc thu hồi nó thay vì để agent tiếp tục lặp lại mãi về sau.

Sau đó mình gặp một vấn đề mà test thông thường không trả lời được. Code vẫn chạy, lint vẫn sạch và pull request nhìn riêng không có lỗi rõ ràng, nhưng kiến trúc đang âm thầm trở nên khó thay đổi. Một module bắt đầu phụ thuộc ngược vào module khác. Một file nhỏ dần trở thành điểm mà mọi thay đổi đều phải đi qua. Một dependency cycle mới xuất hiện, nhưng tổng số cycle không đổi vì một cycle cũ vừa được sửa. Nhìn vào con số, mọi thứ có vẻ đứng yên; nhìn vào bản chất, một vấn đề cũ đã được thay bằng một vấn đề mới.

Điều đó làm mình liên tưởng đến hệ thống điện và nước nằm sau những bức tường. Một căn nhà vẫn có thể sáng đèn và nước vẫn chảy bình thường, dù đường dây bên trong đã bắt đầu được nối chồng chéo. Trong ngày đầu, không ai thấy vấn đề. Nhưng đến lần sửa tiếp theo, muốn thay một ổ điện lại phải mở cả mảng tường; muốn sửa đường nước ở bếp lại ảnh hưởng đến phòng tắm. Codebase cũng vậy. Test cho mình biết căn nhà hôm nay vẫn hoạt động, nhưng chưa chắc cho biết ngày mai nó còn dễ sửa hay không.

Architecture Pulse trong v1.4 được xây để nhìn vào phần nằm sau những bức tường đó. Nó so sánh cấu trúc repository trước và sau khi agent làm việc, tìm các vòng phụ thuộc, ranh giới bị vượt qua, điểm nóng và phạm vi ảnh hưởng. Trong quá trình phát triển, mình cũng nhiều lần đặt “ổ khóa” sai chỗ. Có lúc CI guard được thiết kế quá rộng và chặn cả việc tạo baseline trong bộ nhớ, dù thứ cần ngăn chỉ là ghi một baseline không đáng tin thành artifact. Nó giống như muốn ngăn người lạ mang tài liệu ra khỏi phòng, nhưng lại khóa luôn chiếc bàn nơi người bên trong đang đọc tài liệu. Cách sửa không phải bỏ ổ khóa, mà là đặt nó lại đúng cánh cửa.

Một test khác chạy tốt trên POSIX nhưng thất bại trên Windows. Điều đó nhắc mình rằng một chiếc chìa khóa mở được cửa nhà mình chưa chắc mở được cùng loại cửa ở nơi khác. AI Agent Kit vì thế không thể chỉ hoạt động tốt trong môi trường mình dùng hằng ngày rồi xem đó là đủ. Các giới hạn về filesystem, đường dẫn, symlink và package cần được kiểm tra trên những nền tảng mà người dùng thực sự sẽ chạy.

Đến lúc này, mình từng nghĩ AI Agent Kit đã bao phủ khá nhiều phần quan trọng: agent có phạm vi, hành động có bằng chứng, nhiều agent có cách phối hợp, memory được kiểm soát và kiến trúc có thể quan sát. Nhưng rồi mình nhận ra tất cả những điều đó đều bắt đầu sau khi mình đã quyết định sẽ xây sản phẩm gì.

Trong thực tế, có những lúc mình bắt đầu chỉ bằng một ý tưởng rất ngắn. Agent thường phản hồi nhanh: đề xuất tính năng, chọn công nghệ, chia backlog rồi chuẩn bị implementation. Cảm giác ban đầu khá thích vì mọi thứ tiến lên ngay lập tức. Nhưng nó cũng giống như thuê một đội xây dựng rất giỏi rồi yêu cầu họ dựng nhà khi mình chưa nghĩ rõ ai sẽ sống trong đó, cần bao nhiêu phòng hay thói quen sinh hoạt hằng ngày ra sao. Ngôi nhà có thể được xây đúng bản vẽ, đúng tiến độ và rất chắc chắn, nhưng sau cùng vẫn không phù hợp với người sử dụng.

Mình không muốn một ý tưởng vừa được nói ra đã lập tức biến thành một bản đặc tả có vẻ hoàn chỉnh. Mình muốn có thời gian để hiểu vấn đề, kiểm tra giả định và tách điều mình biết khỏi điều mình mới đang đoán. Nếu BRD chưa được duyệt, agent không nên tự bước sang thiết kế. Nếu thiết kế làm thay đổi điều đã thống nhất trước đó, những phần liên quan cần được xem lại. Product Genesis trong v1.6 ra đời từ nhu cầu ấy. Nó giống giai đoạn mình ngồi xuống với bản phác thảo đầu tiên, hỏi căn nhà này dành cho ai, vì sao cần xây và điều gì thực sự quan trọng trước khi bắt đầu chọn vật liệu.

Khi Product Genesis đã hoạt động, mình lại thấy cách bước vào hành trình vẫn chưa đủ tự nhiên. Để bắt đầu, người dùng phải biết tên skill, nhớ workflow hoặc nói đúng câu lệnh. Trong khi cách mình thường bắt đầu một ý tưởng chỉ đơn giản là: “Mình đang nghĩ đến một sản phẩm như thế này…” Mình không muốn phải học ngôn ngữ của công cụ trước khi công cụ có thể hiểu ngôn ngữ của mình.

Vì vậy v1.6.1 làm cho cánh cửa đầu tiên trở nên nhẹ hơn. Mình có thể kể về ý tưởng bằng tiếng Việt hoặc tiếng Anh như một cuộc trò chuyện bình thường. AI Agent Kit nhận ra đó là một ý tưởng sản phẩm, tìm lại workspace nếu hành trình đang dang dở hoặc bắt đầu một hành trình mới khi cần. Cánh cửa dễ mở hơn, nhưng không vì thế mà mọi căn phòng đều được mở sẵn. Những quyết định quan trọng vẫn cần con người phê duyệt, bằng chứng vẫn phải gắn với trạng thái thật và AI vẫn phải biết dừng khi chưa đủ thông tin để đi tiếp.

Nhìn lại, mỗi phần của AI Agent Kit thường bắt đầu từ một va chạm rất cụ thể. Một công cụ tự làm chữ ký của mình hết hạn. Một worktree pass test nhưng vẫn còn hơn bảy mươi thay đổi. Hai nhánh đều tốt nhưng không đi cùng một lịch sử. Một test chạy trên máy này nhưng thất bại trên Windows. Một agent nhớ rất nhiều nhưng không biết điều gì đã cũ. Một pull request xanh toàn bộ nhưng kiến trúc bên trong đang khó sửa hơn. Và cuối cùng, một đội có thể xây rất nhanh trong khi câu hỏi “có nên xây điều này không?” vẫn chưa được trả lời.

Vì vậy câu hỏi ban đầu vẫn còn ở trung tâm AI Agent Kit: nếu giao một phần codebase cho AI, làm sao biết nó đang làm đúng? Chỉ là sau từng trải nghiệm, chữ “đúng” với mình đã rộng hơn. Nó không còn chỉ là code chạy. Nó còn là bước vào đúng chỗ, chạm vào đúng phần, dùng đúng quyền, dựa trên đúng thông tin, để lại đúng bằng chứng và vẫn đi về hướng mà mình thật sự muốn xây.

Khám phá AI Agent Kit

Sources

  1. AI Agent Kit - A Story That Began With a Small Worry
Author

Hung Pham

Writing from Hunpeo Labs Journal.

Discussion

Newest first

Loading comments…

Comments appear after review.