AI agents can research, code, test, and review at remarkable speed. The harder question is whether anyone still owns the outcome.

🌐 Languages: English · Tiếng Việt bên dưới

Scroll through your feed and you can feel that something has shifted. One founder has built a product over the weekend; another is running a small army of AI agents from a laptop. Research, design, code, testing, launch copy work that once required a team now seems to happen somewhere between Friday night and Monday morning. The one-person startup no longer sounds like a thought experiment. It sounds like something you could begin tonight with an idea, a laptop, and a few terminal windows.

The excitement makes sense. For years, ideas moved faster than the people available to build them. You could see the product clearly but had no designer. You found a designer but still needed an engineer. By the time the first version was ready, you were already looking for someone to test it, explain it, sell it, or support the people using it. There were always more ideas than hands. Then AI arrived. At first, it helped us think; soon it could write code, inspect systems, plan features, run tests, and review the work of other agents. After a long drought, it felt like rain. We were so relieved to have water that few of us stopped to ask where it would all go.

Imagine a founder building a booking app for small hair salons. The idea is straightforward: customers choose a service, pick a time, and make an appointment; the salon sees the booking and assigns a stylist. A few years ago, someone with this idea might have spent months finding the right people before writing a single line of code. Now the founder can divide the work among agents. One researches the market, another designs the interface, a third creates the database, while others handle authentication, scheduling, notifications, and testing. Within a week, the product is running.

Every morning, the founder opens his laptop and finds that it has grown overnight. One day there is a calendar for managing appointments; the next, reminder emails. Soon there is a revenue report, followed by suggestions for discount codes, a loyalty programme, and a dashboard filled with charts. A steady stream of “completed” statuses makes the product feel as though it is moving at extraordinary speed. Keep assigning tasks, and perhaps it will continue to grow, polish itself, and eventually find its own way into the market.

That feeling lasts until a real salon tries it. A customer books an appointment for three o’clock, then cancels because something comes up. The appointment disappears from the customer’s phone, but the three o’clock slot remains blocked in the salon’s system. Another customer tries to book the same time and cannot. It sounds like a small bug, so the founder asks an agent to investigate.

The first agent decides that the problem comes from state synchronisation. It changes the API, adjusts the frontend data, and adds a test. A review agent examines the work and concludes that the appointment state model is not robust enough, then proposes reorganising the booking flow. The database agent notices that the current schema may be difficult to extend and creates another table. Finally, the testing agent runs the suite and returns a reassuring wall of green. By the end of the afternoon, a minor cancellation bug has produced changes across more than forty files.

Every explanation sounds reasonable, and each agent has done its part according to the way it understood the task. Yet somewhere among the analyses, new code, schema changes, and passing tests, one simple question remains unanswered: after the first customer cancelled, could a second customer actually book the three o’clock slot?

This is the quiet gap between work being produced and a problem being solved. The frontend agent cares whether the interface updates. The backend agent looks at the API response. The database agent checks whether the records remain consistent. The testing agent knows only the cases that someone thought to encode as tests. Every part has someone working on it, yet the customer’s complete experience belongs to no one.

The same thing happens in an ordinary restaurant during the lunch rush. Customers are arriving faster than the tables can turn. Someone takes orders near the door, another calls them into the kitchen, cooks move between hot pans, and servers weave through the room carrying plates. Finished dishes cover the counter. Everyone looks busy, and the restaurant appears to be operating at full capacity. Yet table three has been waiting for almost half an hour, table seven receives the same dish twice, and a plate of fried rice sits cooling at the edge of the kitchen because nobody remembers who ordered it.

From inside the kitchen, judged by the number of dishes prepared, this looks like an exceptionally productive lunch service. From table three, where no food has arrived, it looks completely broken. Software built with AI can create the same illusion. We count tasks closed, lines of code generated, tests executed, and agents kept busy, while the user cares about something much simpler: did the meal they ordered reach the right table? When an agent reports that a task is complete, it may only mean the dish has left the pan. It does not mean the dish reached the customer, that it was the right dish, or that the customer could actually eat it.

We understand this distinction instinctively in everyday life. Nobody evaluates a restaurant by counting how many times the cooks moved their spatulas; we care whether customers received the right food, how long they waited, and how often a dish had to be sent back. We would not judge a plumber by the number of metres of pipe replaced either. We want to know whether the tap still leaks, whether anything else was damaged, and who will return if there is water on the floor again tomorrow.

Yet in the world of AI, volume is remarkably persuasive. Five thousand lines of code sound more impressive than five. Ten agents feel more powerful than one. A three-page report appears more trustworthy than a tiny change with almost nothing to explain. But if five lines fix the cause while five thousand leave a human reviewing code for two days and repairing side effects for another week, which result was more efficient? If ten agents produce ten interpretations of the same task, do we have a stronger team or simply ten new directions in which to get lost?

An honest measure of AI’s effectiveness should not begin with how much it produces. It should begin with what remains after the agent says it is done. How much cleanup is left for a human? How often does the work need to be redone? Does anyone understand why the system changed? When something breaks, do we know where to begin looking? Most importantly, did the completed work help the user do what they came to do, or did it merely make the repository busier?

This is the less glamorous side of the one-person startup. One founder may be able to operate many agents, but that does not make the different roles inside a company disappear. Consider the owner of a small restaurant. She may buy the ingredients, work the register, and step into the kitchen when the lunch rush begins. The restaurant has one owner, but the responsibilities remain distinct. At the register, she needs to know which tables have paid. In the kitchen, she needs to know which orders are waiting. Before a plate leaves the counter, she still checks that it is going to the right table. One person can perform many roles; that does not make those roles interchangeable.

A founder working with AI faces the same reality. Agents can research, code, test, and write content, but at some point the founder must step out of the role of the person trying to move faster and look at the product as a whole. AI can perform the work, but it cannot own the consequences. When the system fails in the middle of the night, the agent will not answer the customer’s call. When data is changed incorrectly, the agent will not have to explain what happened. When a technical decision costs the company weeks of recovery work, the agent will not live with the consequences. Approval may require only one click, but the responsibility behind that click still belongs to a person.

None of this means dragging AI into the same heavy processes, endless meetings, and paperwork that have slowed teams down for years. A restaurant does not need a twenty-page operating manual to deliver the right meal to the right table. It needs an order with a table number, the name of the dish, and someone who checks the plate before it leaves the kitchen. Calling a plumber does not require a project plan either. A clear conversation is enough: fix the leak; if you need to break through the wall or replace the entire pipe, call me first; turn the water on when you are finished; and if it leaks again tomorrow, come back and make it right.

These expectations feel ordinary because, in everyday life, we understand that even a skilled and fast worker must know what the job is, how far they are allowed to go, and what must be true before the work can honestly be called finished. It is only when we enter the world of AI carried away by its speed and apparent limitlessness that we forget how useful ordinary clarity can be.

Return to the booking bug. The founder does not need another elaborate process. He only needs to keep the task anchored to the original story: a customer cancelled, but the slot did not reopen. Reproduce that exact failure first. Change only what is necessary to release the time slot. If the fix requires altering the database or touching the payment flow, stop and ask before proceeding. When the change is ready, cancel the appointment with one account and use another account to book the same time. A second agent may review the work, but it should not use the review as an opportunity to redesign the entire system.

This time, the agent may change only a few files, and its report may be too short to show off. But another customer can now book the three o’clock appointment. That is what the user needed, and it is the moment when the work can truthfully be called complete. The result is no longer judged from the perspective of the person or agent who produced it. It is judged from the perspective of the person waiting to receive it.

This is where the image of water becomes clearer. AI resembles a mountain stream running down a steep slope: fast, forceful, and carrying more energy than we have ever had at our disposal. Every agent opens another current. Every prompt creates another tributary. Code, documentation, tests, analyses, and ideas keep flowing downhill until the repository slowly becomes a lake.

At first, watching the waterline rise feels like progress. Every day brings more features, more files, and more reports. Once enough water has gathered, however, it becomes difficult to tell which streams are clean and which are carrying mud, where the lake is safely deep and where the pressure is quietly weakening the bank. The problem is not that we have too much water. It is that we have not built channels to carry it to the right fields, installed gates that can close before it spills over, or placed anyone at the edge of the lake to watch the waterline and say, “This field has enough. Send the rest somewhere else.”

A farmer does not boast about how many cubic metres of water passed through the land that day. The farmer waits for the harvest. That may be the most important lesson in working with AI: do not mistake flow for results, activity for progress, or an agent’s declaration of “completed” for a user receiving what they actually needed.

The age of AI may indeed produce very small companies capable of work that once required large teams. That is worth being excited about. But the people who go furthest may not be those running the most agents, writing the longest prompts, or generating the most code. They will be the ones who know what is worth delegating, when to stop, which results can be trusted, and how to keep speed from becoming an illusion of productivity.

Fast-moving water has never created a harvest on its own. It matters only when it reaches the right field, at the right time, in the right amount.

🇻🇳 Phiên bản tiếng Việt

Dạo này, chỉ cần lướt một vòng mạng xã hội là có thể cảm nhận rất rõ một cơn sốt đang lan ra. Người này kể chuyện làm xong một sản phẩm chỉ trong cuối tuần, người kia khoe đội ngũ AI agent có thể nghiên cứu thị trường, thiết kế giao diện, viết code, chạy test và chuẩn bị nội dung ra mắt gần như không cần thêm ai. “Startup một thành viên” từ một ý tưởng nghe có phần viển vông bỗng trở thành thứ rất nhiều người tin rằng mình có thể bắt đầu ngay tối nay, chỉ với một chiếc laptop và vài cửa sổ terminal.

Sự hào hứng ấy hoàn toàn dễ hiểu. Đã có một thời, thứ giữ chân chúng ta không phải là thiếu ý tưởng mà là thiếu người. Có ý tưởng nhưng không biết thiết kế; có bản thiết kế lại không có người viết code; sản phẩm vừa thành hình thì tiếp tục thiếu người kiểm thử, viết nội dung, nghiên cứu thị trường hoặc nói chuyện với khách hàng. Ý tưởng lúc nào cũng nhiều hơn số đôi tay có thể biến chúng thành hiện thực. Rồi AI xuất hiện, giống như một trận mưa lớn đến sau quãng hạn kéo dài. Chúng ta vui vì cuối cùng cũng có nước, nên chẳng mấy ai nghĩ đến chuyện nếu mưa cứ tiếp tục thì nước sẽ chảy về đâu.

Hãy thử hình dung một người đang làm ứng dụng đặt lịch cho các tiệm tóc nhỏ. Ý tưởng ban đầu rất gọn: khách chọn dịch vụ, chọn giờ, đặt lịch; chủ tiệm nhìn thấy lịch hẹn và sắp xếp nhân viên. Trước đây, một người có ý tưởng như vậy có thể mất vài tháng chỉ để tìm đủ người bắt đầu. Bây giờ, anh ta giao việc cho từng agent. Một agent nghiên cứu thị trường, một agent dựng giao diện, một agent thiết kế database; những agent khác tiếp tục lo đăng nhập, lịch hẹn, thông báo và kiểm thử. Chưa đầy một tuần, sản phẩm đã có thể chạy.

Mỗi sáng mở máy là một lần anh ta thấy sản phẩm lớn thêm. Hôm nay có trang quản lý lịch, ngày mai có email nhắc hẹn, hôm sau nữa đã xuất hiện báo cáo doanh thu. AI còn rất nhiệt tình đề xuất thêm mã giảm giá, chương trình khách hàng thân thiết và một dashboard đầy biểu đồ. Những dòng trạng thái “completed” nối tiếp nhau tạo ra cảm giác mọi thứ đang tiến lên rất nhanh. Chỉ cần tiếp tục giao việc, dường như sản phẩm sẽ tự lớn, tự hoàn thiện rồi tự tìm được đường ra thị trường.

Cảm giác ấy kéo dài cho đến khi một tiệm tóc thật sự dùng thử. Một khách đặt lịch lúc ba giờ chiều, sau đó huỷ vì có việc bận. Trên điện thoại của khách, lịch hẹn đã biến mất; nhưng trong hệ thống của tiệm, khung giờ ba giờ vẫn bị giữ. Người khác muốn đặt đúng giờ đó thì không thể. Một lỗi rất nhỏ, ít nhất là khi nghe qua.

Founder giao cho agent kiểm tra. Agent đầu tiên cho rằng nguyên nhân nằm ở trạng thái đồng bộ nên sửa API, điều chỉnh dữ liệu phía giao diện và bổ sung một bài test. Agent review nhìn vào thay đổi rồi nhận xét mô hình trạng thái hiện tại chưa đủ tốt, từ đó đề xuất tổ chức lại luồng đặt lịch. Agent phụ trách database thấy cấu trúc cũ sẽ khó mở rộng nên tạo thêm một bảng mới. Agent kiểm thử chạy lại toàn bộ test và trả về một bản báo cáo xanh mướt. Chỉ trong một buổi chiều, một lỗi huỷ lịch đã tạo ra hơn bốn mươi thay đổi nằm rải rác khắp hệ thống.

Mọi lời giải thích đều có vẻ hợp lý. Agent nào cũng làm đúng phần việc theo cách nó hiểu. Thế nhưng giữa những bản phân tích, những đoạn code mới và những bài test vừa được bổ sung, vẫn còn một câu hỏi rất đơn giản chưa ai trả lời: sau khi khách thứ nhất huỷ lịch, khách thứ hai đã thật sự đặt được khung giờ ba giờ hay chưa?

Chính ở đó, sự khác biệt giữa “có nhiều việc được làm” và “có một vấn đề được giải quyết” bắt đầu lộ ra. Agent frontend quan tâm giao diện đã cập nhật chưa. Agent backend nhìn vào phản hồi của API. Agent database kiểm tra tính nhất quán của dữ liệu, còn agent kiểm thử chỉ biết những trường hợp đã được viết thành test. Mỗi phần đều có người chăm sóc, nhưng câu chuyện trọn vẹn của người khách cần đặt lịch lại không thực sự thuộc về ai.

Nó giống một quán cơm vào giờ trưa. Ngoài cửa, khách bắt đầu đông; người ghi món gọi liên tục vào bếp, đầu bếp không ngừng đảo chảo, nhân viên bưng bê chạy qua chạy lại. Đĩa thức ăn làm xong phủ kín mặt bàn, nhìn ai cũng bận và có vẻ quán đang hoạt động hết công suất. Thế nhưng bàn số ba đã chờ gần nửa tiếng vẫn chưa có món, bàn số bảy lại nhận hai phần giống nhau, còn một đĩa cơm rang nằm nguội ở góc bếp vì chẳng ai nhớ nó thuộc về bàn nào.

Nếu đứng trong bếp và đếm số đĩa đã nấu, đó là một buổi trưa vô cùng năng suất. Nếu đang ngồi ở bàn số ba, câu chuyện hoàn toàn khác. Phần mềm làm bằng AI cũng vậy: chúng ta đếm số task được đóng, số dòng code được tạo, số test đã chạy và số agent đang hoạt động, trong khi người dùng chỉ quan tâm món họ gọi có được mang đến đúng bàn hay không. Một agent báo “đã hoàn thành” nhiều khi mới chỉ có nghĩa món ăn đã rời khỏi chảo. Nó chưa chắc đã đến đúng người, càng chưa chắc người đó có thể dùng được.

Ngoài đời, chúng ta vốn hiểu chuyện này rất tự nhiên. Không ai đánh giá một quán ăn bằng số lần đầu bếp đảo chảo; điều đáng quan tâm là khách có nhận đúng món không, phải chờ bao lâu và món có bị trả lại không. Cũng chẳng ai đánh giá người thợ sửa nước bằng số mét ống anh ấy đã thay. Chủ nhà chỉ muốn biết chiếc vòi còn rò hay không, những thứ xung quanh có bị làm hỏng không và nếu ngày mai nước lại chảy ra sàn thì ai sẽ quay lại xử lý.

Vậy mà bước sang thế giới AI, chúng ta rất dễ bị khối lượng làm cho choáng ngợp. Năm nghìn dòng code nghe ấn tượng hơn năm dòng. Mười agent tạo cảm giác mạnh hơn một agent. Một bản báo cáo dài ba trang trông đáng tin hơn một thay đổi nhỏ đến mức gần như không có gì để kể. Nhưng nếu năm dòng code xử lý đúng nguyên nhân, còn năm nghìn dòng khiến con người mất hai ngày để đọc lại và thêm một tuần sửa lỗi, đâu mới là hiệu quả thật sự? Nếu mười agent tạo ra mười cách hiểu khác nhau về cùng một công việc, ta đang có một đội ngũ mạnh hơn hay chỉ có thêm mười hướng để đi lạc?

Nếu cần một KPI thực tế cho AI, có lẽ đừng bắt đầu từ việc nó tạo ra được bao nhiêu. Hãy nhìn vào phần còn lại sau khi nó nói “xong”. Con người có phải dọn dẹp nhiều không, kết quả có phải làm lại không, có ai hiểu vì sao hệ thống thay đổi và khi sự cố xảy ra, chúng ta có biết nên bắt đầu tìm từ đâu không? Quan trọng hơn cả, thứ vừa được hoàn thành có thật sự giúp người dùng làm được điều họ cần hay chỉ khiến repository trở nên bận rộn hơn?

Đây cũng là phần ít hào nhoáng nhất trong câu chuyện startup một thành viên. Một người có thể vận hành nhiều agent, nhưng điều đó không khiến các vai trò trong một công ty tự nhiên biến mất. Ở một quán ăn nhỏ, chủ quán có thể vừa nhập hàng, vừa thu ngân, vừa vào bếp khi đông khách. Quán chỉ có một người chủ, nhưng khi đứng ở quầy, cô ấy phải biết bàn nào đã thanh toán; khi bước vào bếp, cô ấy phải biết món nào đang chờ; trước khi đưa đồ ăn ra, cô ấy vẫn nhìn lại xem có đúng bàn hay không. Một người đảm nhiệm nhiều vai trò không có nghĩa những vai trò ấy trở thành một.

Founder dùng AI cũng vậy. Anh ta có thể nhờ agent nghiên cứu, viết code, kiểm thử và làm nội dung, nhưng vẫn cần có lúc bước ra khỏi vai trò của người đang muốn làm thật nhanh để nhìn lại toàn bộ sản phẩm. AI có thể giúp thực hiện công việc, nhưng nó không thể làm chủ hậu quả thay con người. Khi hệ thống gặp lỗi lúc nửa đêm, agent không phải người nghe điện thoại của khách hàng. Khi dữ liệu bị thay đổi sai, agent không phải người giải thích chuyện gì đã xảy ra. Và khi một quyết định kỹ thuật khiến cả sản phẩm phải mất nhiều tuần sửa lại, agent cũng không phải người sống cùng hậu quả của quyết định ấy.

Điều đó không có nghĩa chúng ta phải kéo AI trở lại những quy trình nặng nề, đầy họp hành và giấy tờ. Một quán cơm không cần viết tài liệu hai mươi trang để nhân viên mang đúng món đến đúng bàn; họ chỉ cần một tờ order ghi rõ số bàn, tên món và một người nhìn lại trước khi món rời khỏi bếp. Người gọi thợ sửa vòi cũng không cần lập kế hoạch dự án. Họ chỉ cần nói rõ: hãy sửa chỗ đang rò; nếu phải đục tường hay thay cả đường ống thì gọi lại trước; sửa xong mở nước kiểm tra; nếu ngày mai vẫn còn rò thì quay lại xử lý.

Những điều ấy nghe bình thường vì trong đời sống, chúng ta hiểu rằng một người dù giỏi và nhanh đến đâu cũng cần biết mình đang làm việc gì, được phép đi đến đâu và khi nào công việc mới thật sự kết thúc. Chỉ khi bước vào thế giới AI, bị cuốn theo tốc độ và cảm giác vô hạn, chúng ta mới quên mất sự bình thường ấy.

Quay lại lỗi đặt lịch, founder không cần tạo thêm một tầng quy trình phức tạp. Anh ta chỉ cần giữ công việc ở đúng với câu chuyện ban đầu: khách đã huỷ nhưng khung giờ chưa được mở lại; trước hết hãy tái hiện đúng lỗi đó; chỉ sửa phần liên quan đến việc giải phóng khung giờ; nếu buộc phải thay đổi database hoặc luồng thanh toán thì dừng lại và báo trước. Sau khi sửa, hãy dùng một tài khoản huỷ lịch rồi dùng tài khoản khác đặt lại đúng giờ đó. Một agent khác có thể kiểm tra phần thay đổi, nhưng không được nhân cơ hội viết lại toàn bộ giải pháp.

Lần này, có thể agent chỉ sửa vài file. Bản báo cáo cũng chẳng dài đến mức đáng đem đi khoe. Nhưng khung giờ ba giờ chiều đã được đặt lại thành công. Đó mới là điều người dùng cần, và cũng là lúc công việc có thể thật sự được gọi là hoàn thành. Sự khác biệt nằm ở chỗ kết quả không còn được đánh giá từ phía người tạo ra nó, mà từ phía người đang chờ nhận nó.

Đến đây, hình ảnh dòng nước mới hiện ra rõ ràng hơn. AI giống một con suối trên dốc cao, chảy nhanh, mạnh và mang theo nguồn năng lượng mà trước đây chúng ta chưa từng có. Mỗi agent mở thêm một dòng chảy, mỗi prompt tạo thêm một nhánh nước; code, tài liệu, test và ý tưởng cứ thế đổ xuống. Repository dần trở thành một cái hồ.

Ban đầu, nhìn mặt hồ lớn lên khiến chúng ta tin rằng mình đang tiến bộ. Mỗi ngày có thêm tính năng, thêm file, thêm báo cáo. Nhưng khi nước đã quá nhiều, ta bắt đầu không biết dòng nào sạch, dòng nào đang mang theo bùn đất, chỗ nào đủ sâu và chỗ nào đã âm thầm làm bờ hồ yếu đi. Vấn đề không phải chúng ta có quá nhiều nước. Vấn đề là chưa có con mương đưa nước đến đúng thửa ruộng, chưa có chiếc van để đóng lại khi nước bắt đầu tràn và cũng chưa có ai đứng trên bờ để nói rằng chỗ này đã đủ rồi.

Một người nông dân không khoe rằng hôm nay ruộng của mình nhận được bao nhiêu mét khối nước. Điều họ chờ là mùa thu hoạch. Có lẽ đó cũng là bài học quan trọng nhất khi làm việc với AI: đừng nhầm dòng chảy với kết quả, đừng nhầm sự bận rộn với tiến bộ và đừng nhầm lời thông báo “đã hoàn thành” với việc người dùng đã nhận được điều họ cần.

Thời đại AI hoàn toàn có thể tạo ra những công ty rất nhỏ nhưng làm được những việc từng cần cả một đội ngũ lớn. Điều đó đáng để hào hứng. Nhưng người đi xa có lẽ không phải người mở được nhiều agent nhất, viết được prompt dài nhất hay tạo ra nhiều code nhất. Đó sẽ là người biết việc nào đáng giao, lúc nào cần dừng, kết quả nào có thể tin và đủ tỉnh táo để không biến tốc độ thành một ảo giác về năng suất.

Bởi nước chảy nhanh chưa bao giờ tự làm nên mùa màng. Nó chỉ trở nên có ý nghĩa khi đến đúng nơi, vào đúng lúc và vừa đủ cho điều đang cần được nuôi lớn.

Sources

  1. You Can Run a Company With AI Agents. But Who Is Running the Agents?
Author

Hung Pham

Writing from Hunpeo Labs Journal.

Discussion

Newest first

Loading comments…

Comments appear after review.