Lately, my tweets have been almost entirely about LLMs. I'm that obsessed. If I have free time, I immediately ask LLMs for something or ask them questions. Especially having agents write code feels like being engrossed in a new social game, not wanting to waste a single moment.

It's fun to have agents write code. I'm creating an RDBMS, and things like "What should I do next?" "Add a PostgreSQL-compatible interface?" "Huh? Can you do that?" "Done!" happen one after another, and implementations spring up rapidly. Of course, it's not perfect from the start, so while writing unit tests, flaws come out one after another, but even so, the development speed more than makes up for it. More than anything, the fact that nearly 100,000 lines of code can suddenly appear with about the same burden as having one more chat partner alongside my work is awe-inspiring beyond admiration.

As those who have been following my words and actions on social media know, I have long distrusted software development engineering. I strongly felt that it was not something a programmer should learn to improve, but rather a field of study for engineering the behavior of groups when programming is done collectively, and I doubted whether it would be applicable to all the diverse programming sites. However, when using multiple agents, I realized that I was naturally taking a software development engineering approach, requiring a sense of pinpointing the existence of bugs without reading the code or gradually expanding reliable areas. I'm not specifically incorporating things like bug density or convergence curves, but I recognize that I'm not far from them when viewed from a bird's-eye perspective.

Even in situations where a layman might shout, "Create software with no bugs!!" at the agent, most problems can be solved by patiently issuing instructions with the knowledge of software, triaging rejected work through dialogue, and issuing instructions again. Repeating such work reassures me that the senses I have cultivated so far have not been wasted, but at the same time, I strongly feel that it is only a matter of time before agents can solve these kinds of things on their own. In this respect, I have a very pessimistic outlook on the future of programmers' jobs.

The code written by agents is honestly like the uncanny valley, and it doesn't make me want to maintain it very much, but as a premise, software engineers don't want to maintain code in the first place. They only maintain it because the code generates money at work, and reading code is the last resort as a maintenance method for professional programmers (craftsman programmers are welcome to happily tinker with specific code for the purpose of the code itself, and I'm fundamentally more inclined towards that as well). Programming is the work of construction, while software engineering is a continuous act to ensure that it continuously generates value, and, without fear of misunderstanding, it is the culmination of decades of effort to guarantee quality without having to read code that you don't even want to read.

The rules have changed significantly, so some things can be used and some things cannot. I will write down the tips that I currently feel as much as I can think of below.

Context Engineering is Organizational Theory

Whether AI writes it or not, the universal truth of software is that code is always growing in size. I think the context length of LLMs will continue to increase in the future, but the growth rate of software code will always exceed it. Even if LLMs can avoid context problems, the range of context that can be focused on is finite, and the problem of finding a needle in a haystack will always be present. Of course, the entire team cannot push the full picture of the production code into their heads, so they will divide roles without being told to do so. That division of roles is context engineering from a different perspective. In other words, throwing the entire project image into the context is too large and makes it impossible to think about anything, so "I want you to add this function to this module in this job, and you don't need to know anything that is unnecessary for performing that task" is a way of dividing work that is done to new employees in most occupations, and it is not limited to AI agents. As you can see when using AI agents, they are pathologically fast at turning 0 into 1. They were raised on such a benchmark set, and solving bugs that humans struggle with is just a bonus. The reason why it is fast to turn 0 into 1 is that the existing code that must be put into the head to perform that task is zero, so there is extremely little concern about contact points. Programmers with small contexts also tend to want to create microservices, because microservices reduce the context that must be put into the head. Microservices are an extreme example, but the act of separating the inside when constructing software is to limit the range that others have to worry about, and what has been called "separation of concerns" is also the separation of information that should be put into the agent's head as is. The story of what "software architecture" is has not yet reached a complete consensus, but my feeling is that it is what separates the context that should be put into the brains of the people working on it. Even if a value object is introduced with a word from the top, if the context that lower-level programmers have to push into their heads does not change, then it cannot be said that the architect has done the right job. I think that dividing tasks into a form that is easy for agents to work with and dividing the inside of the software into a form that is easy for humans to work with is quite similar. My current intuition is that the things to worry about at that layer are larger than I imagined, so the context is also large and a smart high-end model is used, but if I offload it to an investigation agent or something, a cheap model might work well.

LLM is Expensive

Everyone says that the high-end model outputs better code, but the price of the high-end model is in the world of $10 or $20 per 1M tokens. The work done using that context will take several hours of a programmer's time, so I think it pays off, but if I can leave the work to a cheap model, that's even better. As a result of price competition, the current high-end model equivalent may become incredibly cheap and used like water, but by that time, the high-end model of that time will be released, and it will be advertised that it can solve problems of difficulty that cannot be solved elsewhere. As a result, I think that the state of writing tasks that are not so difficult when writing programs with cheap models will continue in the future. What's more, the local LLM, which is the extreme of cheap models, is attractive as an offload destination, and a group of local agents that subcontract tasks defined using high-end models is a realistic story. I don't know how complex coding tasks the current local LLM can handle, but if I let the cheap one write it and throw up my hands because I can't debug it, the story of escalating to a smart model with that knowledge will probably become common sense. Local and high-end will work together in the form of sub-agent calls, and a team of several people who actually have a large number of local LLMs and continue to run them may be the finished form of a software development company. AMD's Ryzen AI Max series is also likely to increase in price more and more, because the competing product is a blood-carrying human programmer.

AI Agent is Unstable

I don't say anything definitive because I don't know what implementation is bad, but unexplained stops, intermittent errors, and things I don't understand are always happening. It is not uncommon for the kernel to not crash, but for the editor's response speed to be in minutes. Also, the quality of rewriting is very uneven, and the line numbers of the rewriting destination are messed up, so it often breaks in a terrible way. If it breaks, it will spontaneously notice it and go to fix it, so it's not that fatal, but as it is, there is a trick to using unstable things as they are, and as far as I know, Rall Loop and the accompanying progress tracking are effective.

In a nutshell, the Rall Loop is to put the agent's startup itself in a loop. However, instruct the prompt to write the progress somewhere, and load the progress at the beginning of the loop. Then, the agent will perform "this task" after receiving "the progress so far" as context. What's good about this is that even if the agent stops or dies due to OOM, it will run one way or another. It's a problem if the progress memo becomes too large, but you can write a prompt to summarize and compress it as needed. However, if it is a single loop, all tasks will be performed by a single expensive model, so the phase of designing and dividing tasks is left to the expensive model, and tasks that seem easy are done by cheap agents, and so on. The contents of the loop itself become more complex. Since I am programming whether to call a coding agent in that way, it is a fine meta-programming. I have not yet reached this stage... It will be natural to have the loop itself improved by LLM.

The prompt that antirez issued to speed up the code is interesting because it instructs "5. Use this file to track progresses." Since the work continues by referring to the same text even if the work ends in the middle, it is very reasonable that the text contains the progress up to the previous time.

5 Self-Reviews

90% of the code output by the coding agent works well, but when it becomes a large amount, improvements are found each time it is self-reviewed. I want you to write complete code that does not require review from the beginning, but since it is difficult for humans, LLM probably has drawbacks that cannot be found unless it is reviewed in a different context. It may take time for humans to reflect on the code they have written, but overseas engineers have written that it often converges in 5 times when they have the coding agent self-review.

https://steve-yegge.medium.com/six-new-tips-for-better-coding-with-agents-d4e9c86e42a9

Of course, it is very costly, but it is still much cheaper than humans, and if it is included as a mechanism, there will be no room for human intervention. I don't know if the number 5 has any specific meaning, and it may be a different number in the future, but as a matter of experience, the code written by the agent still has improvements even if it is reviewed several times. LLM's special skill is reflection. (There are also review comments that seem to be attached like a fault to make the dialogue work, so it may be necessary to review the review comments themselves). If the code written by the coding agent is used in production, it seems that it will continue for a while that humans should review it at the end, but as it is, it seems that it will become common sense that the machine itself reviews it N times before humans review it. It seems that the creating side and the blaming side raise each other. Reviewing with agents from multiple companies may also bring new perspectives.

It's getting too long, so I'll post it around here for now.