zgba 站群
Coding Is Not Solved

Coding Is Not Solved

Disclaimer: you are about to read a lot of opinions, many of them have references but some are the result of my own experience building with AI and building AI systems in the past 4 years. Regardless, beware of the straw-man fallacy: just because one argument doesn’t map to your belief system, it doesn’t mean the rest are invalid. I should also say upfront that I’m not anti-AI. If you’ve been following my work, you know that I was an early adopter of not only using LLM-powered coding tools, but building my own harness, teaching these topics and building LLM-powered products. It’s not about fear of AI but rather challenging the brain-dead narrative that asserts “coding is solved” and engineering is about “taste” now.

Update: someone put this on Hackernews.

Tell me you don’t understand software without literally using those words!!!

People who claim “LLMs can write decent code” don’t understand how code works. Sure, creation is much cheaper, but anyone who has run software in production at scale knows that maintenance, reliability, security, scalability, etc. is the majority of the cost. These are commonly known as NFR (non-functional requirements).

In my experience even the Functional Requirements (what the code is supposed to do) is NOT a solved problem yet. There’s a bit of Dunning-Kruger effect at place where the people who don’t read the output are more confident in it.

As a veteran developer holding 2 engineering degrees (hardware and systems engineering), I can list 3 types of products that do not strictly require reading the code:

Personal software: scratching an itch, automation, DIY patches, etc.

POC (proof of concept): demonstrating technical feasibility and product viability

Weaponized AI: acknowledge the risk and deliberately point it at a target to cause harm

Notice the commonality: the first 2 have high risk tolerance while the last one weaponizes the inherent risk.

Most software that requires hiring and paying software engineers has low risk tolerance:

…wherever a mistake can cost money, lives or legal consequences you need accountability.

AI cannot be held accountable. It cannot suffer any consequences. The worst thing you can do to AI is to unplug it. And although it mimics human emotions (due to training data), it couldn’t care less. AI doesn’t die either. It cannot suffer a prison sentence or fines. You cannot punish AI, therefore it can never be held accountable.

You cannot be responsible for what you can’t control either. That understanding is key to reasoning about system behavior and fixing it when the AI inevitably fails.

If you’re toying around, LLMs do a great job. That’s why some of the most aggressive proponents of the “coding is solved” narrative have nothing to show for it. Anthropic accidentally leaked Claude Code (which on further study turned out to have many flaws) and their status page shows orange is the new green!

Contrary to common narrative, coding is actually one of the last areas for the current generation of LLMs to take over!!!

Allow me to elaborate:

Coding is about logic. Anyone who has dealt with compiler errors knows that computers don’t give a f*** about how right you think you are. If it’s logically wrong, it doesn’t compile. Even if the syntax is fine, there are runtime errors.

The reason LLMs are successful in writing code is because we’ve made a feedback loop that feeds the syntax/runtime errors back to the LLM and loops until most errors are solved or hidden.

LLMs can wing it for tasks that are related to natural language (e.g. writing social media posts, reports, articles, etc.) but when it comes to code, the same engine that struggles to count number of R’s in “Raspberry” or suggests a walk to the carwash, also exposes other logical fallacies.

LLMs are stochastic and probabilistic. The only way we could even get remotely close to making them logical is to wrap them in traditional code (known as harness), run tests, and a bunch of other techniques (e.g. CoT) but the core issue remains: LLMs struggle with logic and volume (the larger the input and the more the context window is used, the less accurate they get).

I’m not saying LLMs cannot generate code or maintain existing code bases. They have their utility as a tool and their capabilities are increasing in an S-curve. There is a point of diminishing return where more expensive models aren’t necessarily more productive at the rate of the price increase.

Those who claim LLM-generated software is good enough:❌ Haven’t written code in ages❌ Cannot spot if their code figuratively had 6 fingers!❌ Have a low bar for what good looks like❌ Don’t care about quality or NFR❌ Have difficulty understanding an S-curve✅ Are honest: AI genuinely writes better code than them

But to go ahead and extrapolate that to an entire professional industry requires a level of brain-dead thinking that’s only present in people who spend too much time with sycophantic AI.

I’m not here to change anyone’s workflow or toolbox. I couldn’t care less.

What I do care is that the services I’m paying for (looking at you Google and GitHub) are degrading with stupid bugs that could be avoided if we prioritize reliability and accountability over velocity.

Anthropic’s Boris Cherny is one of the most vocal proponents of the “coding is solved” narrative. By many accounts Cloude Code is the epiphany of his ideology:

Running Claude Code returns Bun’s help menu

Claude Code CLI binary installer silently deletes itself after installation

Extra Usage charged despite available plan capacity + false rate limit errors

Reminder: Anthropic controls the model (Claude), the harness (Claude Code), the prompt (see the leaked versions) and the runtime (Bun).

If you’re in leadership position, please don’t stress your [otherwise smart] developers to force AI into every possible surface and workflow.

The tech has some genuine power and is the biggest change in our industry in ages. But AI overuse is a thing, and when it hurts the customer, you are accountable.

Stop repeating the half-baked narratives from token sellers about exaggerating the capabilities of AI because we, the consumers pay the end price.

AI is great for POC (proof of concept), Personal Software (a growing category), Map-reduce on human language (e.g. translation, converting different formats, summation, expansion) and cyber attacks (due to the delta between artificial intelligence and organic one) with varying degrees of success but the current generation of tech has fundamental problems too.

I don’t want to belittle how far we have come with harness, SKILLS, AGENTS-md, MCP, A2A, ACP, RLM, OKF, MoE, MoA, self-healing, and various runtimes, quantizations, optimizations, architectures, and memory techniques.

I’ve written about many of those before:

Those are great pragmatic approaches to work around LLM shortcomings and there are probably more to come.

What I’m trying to elaborate is that I don’t want the services (that I depend on) to degrade just because someone pushed AI where it didn’t belong or skipped their job in quality, security, reliability and verification.

AI overdose is a thing and it directly puts an expiration date on your skill set. Those of you who are in the unfortunate position where your manager is whipping you harder and harder to realize AI value, should fight back.

Don’t sacrifice your long term relevance for short term velocity.

How to spot AI overdose?

You have zero tolerance for disagreement and civil discourse.

You let AI run your life and trust AI vendors with stuff that was unthinkable just a few years ago.

You run to AI for things that are slightly cognitively challenging.

You frame your naïveté and laziness as optimism and think the government can save you if things get bad.

You have stopped reading long form text: books, articles, even long emails.

You spend more time with AI than with other human beings or let AI shield you from raw genuine human interaction.

And a bonus point: you skim. Did you notice number 5? 😄

I believe AI is a bar raiser: if the quality of your output is equal or subpar to AI, upskill.

“You can create a full spec upfront”. If you’re that naive, I know a guy in a white van who gives free ice cream! Let me guess, you also believe software estimates are accurate and Santa is real. Anyone with a few years of industry experience knows that it’s impossible to spec the software meaningfully ahead of time (unless it’s very trivial).

“English is the new programming language”. Human language is vague and conflicting. That’s the primary reason programming languages are created. A compiler or type-checker flags some of those conflicts. How on earth can you be sure that one part of your NL instructions doesn’t conflict with another? With syntax checkers we get some help. While it’s possible to task another LLM to read through the instructions and reason about those conflicts, the safest way to discover those nuances is to ask your agent to build what you asked for. But that’s much more expensive than a linter or compiler.

“I move much faster. Can’t remember the last time I wrote code by hand”. Don’t confuse motion with progress. Don’t measure progress with vanity metrics like SLOC, PR count or features. Measure service levels, ie. service consumer’s happiness. Call me when you can prove a margin between token costs and business value.

“I have stopped writing code by hand. I primarily read code and probably next year I won’t even do that”. First of all, human beings are notorious at understanding the S-curve so it may take longer than a year. But even if AI completely eliminates the need to read or write code, you do understand that you are confessing to being redundant right? If a power user can prompt the AI to get what they need, then what value can you bring to the table? Instead of replacing yourself with AI, you should look at what value you can create on top of AI to stay relevant and worth your money.

“The leverage has shifted to taste”. Yeah, this is the lie retired chefs tell to themselves. Just because there’s a bot in the kitchen doesn’t mean that you should sit in the customer’s area in the restaurant! “Taste” is not as payable as you wish! Everyone got a taste! I say that as someone who has spent a big part of my career in Frontend and UX land. Everyone and their dog has an opinion and taste. I know what you mean: taste == experience. But believe me, AI has lowered the bar for the skills required to create decent looking software and simultaneously raised the bar for what’s payable effort. If you bring up “taste” to a job interview, you’ll learn the hard way that the market doesn’t value it as much as you do.

“AI is an equalizer. It makes creativity (writing, coding, making music, videos, etc.) more approachable”. AI is a multiplier: it gives wings to both stupid and smart people. I’m not here to judge but I’ve seen too many sloppy efforts from social media posts, to blogs, memes, and what not. I’ve also seen good use of AI where it genuinely creates high quality work at speed and fraction of the cost. The main difference is human involvement, iteration and depth of knowledge leading to stronger feedback loops. The latter takes more time and effort to the extent some tasks are genuinely cheaper and faster to do manually (e.g. the other day I ran an experiment and tasked my agent to update 5 npm dependencies, all patch releases. It took 12 minutes and 72 steps. I could do it in less than a minute.) Tools like Lovable make it cheaper than ever to fake credibility. Gone are the days when a polished website meant some craftsmanship or at least a deep pocket. AI is a force multiplier, but the force vector direction is more important!

“Agent is the new compiler”. Ah that one again! Sure! If that’s your reality, I let this meme do the work.

Pssst! Do you want to know an old trick to make your LLM-generated code instantly superior?

Run multiple-agents in parallel! The sheer volume of code makes it humanly impossible/expensive to review and you give up!

The trick is the same as pre-AI era: if you want a PR to be merged, make it massive because ain’t nobody got time for that.

It’ll be merged based on “trust”!

You want another tip? Loop engineering: let the agents prompt each other. Big AI labs find about their rogue agents months after the damage is done! Do you think you’re better than them? Learn from the masters! 🙃

We don’t exactly trust AI but we have to because the alternative (having to read the output) is too hard for some folks! Instead they come to social media and claim that since UAT (user-acceptance testing) passes, the code is “good enough”. Then ship it to me and you to do the rest of the testing.

We’re just lab rats after all. 🙃 Just a friendly advice: have a little AI-free hobby project to keep your coding skills fresh for when you’re thrown back to the job market. Cheers!

When talking about AI (not just LLM), there are 2

View original article