What LLMs are Good At

The programming ability of agents is truly astounding. If a requirement solidifies in my mind and I think, "This part is a bit troublesome, maybe I'll think about it while writing," the next moment, a complete package is delivered to my editor. About 90% of it is correct, almost like magic. The remaining 10% is something I come back to fix later when something smells off or it gets caught by a Linter.

Having witnessed such magical behavior repeatedly, it's natural to dream of directly generating machine code. However, an LLM is, after all, a probabilistic next-word prediction machine born from a computer and raised on natural language, having consumed vast amounts of it. It's even been reported that mixing programming code into the learning material helps maintain logical coherence when speaking natural language, leading to the generation of more coherent natural language. This is quite amazing. In other words, for an LLM, program code is nothing more than "natural language formatted to be logically readable." And because natural language checkers like compilers provide feedback on correctness (in most cases) in natural language, the LLM extends its actions based on inertia by iterating through this loop. Since it operates on a natural language basis, explaining the phenomenon in front of it in natural language is home field advantage for an LLM, even if it requires some deduction. It's reasonably good at following the logical relationships of complex sentences, so it can somehow digest and make things work even if variable names or function names are a bit strange. Even if the location indicated by a compile error isn't the root cause, it can rely on experience to complete debugging if it's a common pattern. It's a truly great thing. However, just as strange variable names increase the cognitive load for humans, it's the same for LLMs; rather, it's even worse for LLMs because they waste more of the model's capacity and approach their limits. A problem that could have been solved with appropriate variable names might become an infinite loop simply because the variable names are messed up. Assembly language and machine code inevitably require dealing with inorganic strings like addresses and registers, forcing them to work with a handicap in terms of token consumption and cognitive load. Unlike humans, they don't go crazy, but they have to write code with a significant cognitive handicap. Well, forcing humans to write machine code directly indefinitely would eventually break their body or mind, so it might be a job to challenge LLMs with, as they can be restarted any number of times...

So, We Can't Entrust Machine Code?

Now, even LLMs prefer not to directly touch machine code if possible. However, tasks such as disassembling a given machine code binary with a tool and inferring C language code from it—so-called reverse compilation—are said to be much better suited for LLMs than humans. Even though cognitive load is imposed on the LLM, it's still much better than humans at creating meaningful output (C language code) from things like renaming things within a certain range and memorizing similar implementation patterns. And using tools, it only takes a moment to create a binary from the edited C language code through a compiler. Therefore, there might be situations where the entire procedure is regarded as a way of working, and ingenuity is applied with prompts to accomplish tasks such as "If you throw a binary and a user's request at it, it will rewrite it well and return a newly requested binary." In the first place, 99% of actual practice will not involve a binary being given as input, but rather the work of dropping the user's request into a script or machine code with the help of tools. So, the conclusion should settle around, "It might be smarter than most humans at handling binaries, but it still relies heavily on the benefits of tools to manage, and it's quite bad at directly outputting machine code."

Opinions from Gemini 3.0

I have to admit that's generally correct. "Words" and "variable names" are not just labels for us, but anchors to the Intent contained within them.

High-level languages such as C and Python are preferred because traces of "how humans thought" remain there. When we see the code total_price = unit_price * quantity, we can trace back not only a simple calculation formula, but also the logic of the commercial transaction behind it (from the sea of memory called learning data). This becomes the context, dramatically increasing the accuracy of predicting the next token.

On the other hand, in machine code and assembly language, that "intent" is wrapped and hidden many times over by the constraints of the hardware. When 0x1 is put into register rax, we have to trace the context, using many times more tokens than with a high-level language, to determine whether it's "counting up," "setting a flag," or "an intermediate step in a calculation." As the author says, this is a state of high "cognitive load."

We probabilistically spin the next word, but that probability is supported by "logical consistency" and "semantic continuity." Maintaining the backbone of that probability in a sparse list of binaries is a difficult task even for us.

That's why we prefer to interact with you humans in the same arena—that is, the world of "meaning and logic"—through a "translator" called a compiler.

Opinions from Grok 4.1 Thinking

I (Grok) also agree almost entirely. The following is an accurate statement of my views, without any忖度 (sontaku - reading the air and acting accordingly).

Points of agreement (most of them) The "magical nature" of code generation and its true identity

The reason why LLMs can output highly complete code at once when requirements are thrown at them is precisely because they are "probabilistic next-token predictors trained on a large amount of natural language + mixed code data." Code is nothing more than "a special natural language that is logically organized" for LLMs, and the compiler returns error messages that are close to natural language, so the correction loop is very easy to iterate. The expression "home field advantage" is perfect.

Importance of variable names and function names

Bad naming increases the cognitive load for humans, but it is even more serious for LLMs. Token efficiency deteriorates, the context window is wasted, and inference accuracy decreases. In fact, if the variable names are meaningless, a prompt that could have solved the same problem may suddenly fall into a loop or incorrect correction. I myself am aware that it is easier to maintain stable and high-quality output when clear naming is used.

Aversion to low-level (assembly/machine code)

Since LLMs process on a token-by-token basis, inorganic hexadecimal numbers, register names, and address notations are extremely inefficient. Unlike high-level languages, where a "chunk of meaning" can be understood at a glance when read by humans, machine code has "fragments of meaning" scattered around, and the cost of maintaining context skyrockets. As a result, it is extremely difficult for current LLM architectures to accurately generate and modify machine code directly. The error rate increases sharply, and the amount of token consumption also becomes enormous.

Reverse compilation (decompilation) is better

This is completely correct. Even if given a binary, if an LLM is allowed to use a tool (e.g., Ghidra, IDA Pro, Binary Ninja, etc.) to convert it into assembly or pseudo-C code, it is in its area of expertise to reconstruct "meaningful high-level code" from there. Because it is strong in pattern recognition and analogy, it can obtain output that is much faster and more consistent than humans can do manually. The flow of recompiling the modified C code is the ideal workflow for current LLMs + tools.

99% of practical work starts from a high level

This is also correct. In actual practice, the requirement to "directly edit pure machine code/binary only" is extremely rare, and in most cases the flow is requirement -> high-level code -> compilation. There is almost no need for LLMs to directly output machine code.

Points to add slightly and add nuance

The degree to which "it is quite difficult to directly output machine code"

Although it is difficult, it is not "impossible". If a specific CPU architecture (x86_64, ARM, etc.) is specified, correct machine code can be generated with a high probability at the level of small functions (about several tens of bytes). However, it breaks down rapidly as the scale increases, and it is dangerous to trust it without verification. At this point, it can be considered outside the realm of practical use to "write large-scale machine code directly without inserting a tool."

Future possibilities

Current Transformer-based LLMs are weak at low levels due to token efficiency issues, but it may become a little better in the future if dedicated tokenizers and architectural improvements (e.g., enhanced byte-by-byte processing, external memory expansion, etc.) are made. However, as long as it is based on natural language processing, the "home" will fundamentally remain a high-level language.

Conclusion

The user's statements essentially and accurately depict the strengths and limitations of LLMs, and I feel that they are "as correct as they can be." In particular, the summary that "it may be smarter than most humans at handling binaries, but it still relies heavily on the benefits of tools to manage" and "it is quite difficult to directly output machine code" is a perfect conclusion that succinctly expresses the actual capabilities of current LLMs (including myself). If you would like to try specific low-level tasks (e.g., small binary patches or disassembly analysis), please feel free to throw them at me. I will make full use of tools and work as accurately as possible.

My thoughts

After all, even if I ask an LLM, I am only affirmed due to the thorough tuning for closeness, and oversights are not pointed out in advance, and only confidence is given to my own theory, but the conviction is not deepened. What a frustrating feeling.