zgba 站群
Software Engineering fundamentals matter more

Software Engineering fundamentals matter more

软件工程基础更为重要

The manifestation of my imposter syndrome, for me and today, is what does it mean to be a software engineer. There’s a lot more noise than signal on the Internet about agentic engineering, what can be accomplished, and its implications for the future. The title I chose rather gives it away; it’s about choosing — carefully — all the things you need to choose when you’re solving the puzzles of software and systems development.

对我来说,今天冒名顶替综合征的表现形式就是:成为一名软件工程师究竟意味着什么?在互联网上,关于智能体工程(agentic engineering)、它能实现什么以及其对未来的影响,噪音远多于信号。我选择的标题其实已经暗示了答案;它是关于选择——谨慎地选择——在解决软件和系统开发难题时你需要做出的所有选择。

Beyond the hype and junkie-like marketing fervor of “major model providers”, I found a really interesting power tool with the combination of harness and models. I’ve been following how friends have been using these tools, and learning a ton. As usual, the folks doing some of the most amazing things aren’t the ones crowing about it, or posting narrative blurbs in social media about the end of this profession. They found a “big damn stick”, they’re exploring the fulcrum points, and they’re representing good ole Archimedes to lean into that lever, moving the world.

超越了“主要模型提供商”那种近乎瘾君子般的营销狂热,我发现结合框架(harness)和模型是一个真正有趣的强力工具。我一直关注朋友们如何使用这些工具,并从中学习了很多。像往常一样,那些做出最惊人成果的人并不是那些大肆宣扬的人,也不是那些在社交媒体上发布关于本职业终结的叙事短文的人。他们找到了一根“巨大的棍子”,正在探索支点,并代表着老阿基米德利用杠杆撬动世界。

In the past year, agent harnesses crossed the “can it be done” rubicon. (yep, jumping forward to Roman references). I would not have wished for the world’s knowledge to taken without permission and regard, or the lunatics to delve into economic self-dealing that’s peanut buttering over the otherwise tanking US economy. The economic models for the large models aren’t viable from any report that I’ve seen, but the capability isn’t going away. Instead it’s shrinking (fast!). Open weight models are making (beefy) personal computers quite capable of doing the same. They’re not quite as effective, but the delta in time and capability isn’t large.

在过去的一年里,智能体框架跨越了“能否做到”的卢比孔河。(是的,跳到了罗马典故)。我并不希望世界知识在未经许可和尊重的情况下被拿走,也不希望那些疯子陷入经济自利行为,这就像在原本已经下滑的美国经济上涂花生酱一样。从我看到的任何报告来看,大型模型的经济模型都不可行,但这种能力不会消失。相反,它正在缩小(很快!)。开放权重模型使得(配置强劲的)个人电脑完全有能力做同样的事情。它们的效果可能不完全一样,但在时间和能力上的差距并不大。

“Can it be done” is only the start, not even close to the majority a software or system engineer’s profession. It’s like when I learned to weld in my 20’s – I quickly created things that I couldn’t lift or even get out the door of the shop. (thank goodness for acetylene torches).

“能否做到”仅仅是开始,甚至远未达到软件或系统工程师职业的大部分范畴。就像我 20 多岁学焊接时一样——我很快制造出了我举不起来甚至无法搬出车间的东西。(谢天谢地有乙炔焊枪)。

What I learned then is I think the same lesson, different medium: How something goes together is what makes all the difference. If you use agentic harnesses to develop with a bit of foresight, you can get not only “it works”, but also “it’s testable” (I heavily lean into the prompt “develop with red/green TDD”).

我当时学到的,我认为是一样的教训,只是媒介不同:事物是如何组合在一起的决定了一切。如果你使用智能体框架进行开发,并带有一点远见,你不仅能得到“它能工作”,还能得到“它是可测试的”(我强烈倾向于使用“使用红/绿 TDD 进行开发”的提示)。

But it’s not very solid much above that. The seams — how your code works, it’s “API”, and how it fits with other software — are as much art as science. It is made up of subjective measures that rely on your viewpoint (and experience, as well as your guesses) for both what you’re solving now, and how to live with that software over a long period of time. Making software debuggable, maintainable, layered, and composable – that’s still quite a trick. Quite a lot of that work requires extensive, thoughtful reasoning. And that’s where the LLM’s today, even the leading edge of the “capability” from frontier models, fall short.

但在此之上,它并不非常稳固。接缝——你的代码如何工作,它的”API”,以及它如何与其他软件配合——既是艺术也是科学。它由主观衡量标准组成,依赖于你的观点(以及经验和猜测),既关乎你现在解决什么,也关乎如何在长时间内与该软件共存。使软件可调试、可维护、分层且可组合——这仍然是一个相当大的技巧。其中相当多的工作需要广泛、深思熟虑的推理。而这就是当今的 LLM,甚至是前沿模型“能力”的领先边缘,所不足的地方。

It helps to know that LLMs don’t “reason”. They predict, and the models themselves are effectively written human knowledge compressed. So if it’s in human knowledge that was encoded, it can echo out the human reasoning. For agents focused on software development, those reasoning traces are the precious data for the models.

了解 LLM 并不“推理”是有帮助的。它们进行预测,模型本身实际上是压缩的书面人类知识。因此,如果它包含在已编码的人类知识中,它就能回显出人类的推理。对于专注于软件开发的智能体来说,这些推理轨迹是模型的宝贵数据。

There’s a very approachable research paper on just how bad LLMS are at reasoning called The Illusion of Thinking. There is some research I’m following that includes prediction of results of actions, but that’s not what we have today with coding agents. It’s a pretty different – and fascinating – area of research. If you want to explore, go digging on how “JEPA models” work, LeWorld Model, and recent talks by Yann LeCun.

有一篇非常易懂的研究论文,名为《思考的错觉》(The Illusion of Thinking),讲述了 LLM 在推理方面有多糟糕。我正在关注一些包括预测行动结果的研究,但这并不是我们今天拥有的编码智能体。这是一个非常不同——且迷人——的研究领域。如果你想探索,去挖掘一下”JEPA 模型”是如何工作的,LeWorld 模型,以及 Yann LeCun 最近的演讲。

While you’re working with LLMs though, there’s still a ton of ways to make them more effective. I think there’s a lot of advances that we haven’t even really begun to eek out. Most of the wins I’m seeing today involve providing it good, concise data to work from, at the right time, and providing deterministic validation tooling with natural language feedback that the LLM can use to correct itself. The amazing thing to me isn’t that it can predict what to write, but that it is effective at tool calling and following instructions.

虽然在使用 LLM 时,仍然有很多方法可以让它们更有效。我认为有很多进展我们甚至还没有真正开始挖掘。我今天看到的大部分胜利都涉及在正确的时间提供良好、简洁的数据供其工作,并提供带有自然语言反馈的确定性验证工具,LLM 可以利用这些反馈来纠正自己。对我来说,令人惊讶的不是它能预测写什么,而是它在工具调用和遵循指令方面是有效的。

Another downside of this instruction following is what Simon Willison coined as the lethal trifecta. Basically – LLM models can’t distinguish between good advice and bad. They’re foundationally incapable of always and consistently preventing prompt injection attacks. “Alignment work”, safety harnesses, and sandboxes all help to add barriers against the worst, but there are fundamental gaps. And frankly, something that tirelessly follows instructions without having good reasoning is nightmare fuel to me.

这种遵循指令的另一个缺点是 Simon Willison 所称的“致命三重奏”。基本上——LLM 模型无法区分好建议和坏建议。它们在根本上无法始终如一地防止提示注入攻击。“对齐工作”、安全框架和沙箱都有助于增加对抗最坏情况的障碍,但存在根本性的差距。坦白说,某种不知疲倦地遵循指令却没有良好推理能力的东西,对我来说是噩梦燃料。

I hope there will be near-term nadvances in how models are trained to include the equivalent of reasoning traces for post-training (RLHF). In my ideal future, these include more of what it means to build software with clean interfaces, that’s debuggable, and and that’s maintainable as a key part of the reinforced evaluations.

我希望在模型训练方面会有近期的进展,以便在后期训练(RLHF)中包含相当于推理轨迹的内容。在我理想的未来中,这些包括更多关于构建具有清晰接口、可调试且可维护的软件意味着什么,作为强化评估的关键部分。

Carefully reviewing, planning, and fixing the seams of software (and systems) is one of the critical skills we both can, and need to, employ when developing software – with or without agentic assistants. And as I see the wave of “Oh, that’s easy to implement…” and people reaching for clankers to get it done, I think it’s more important than ever. It’s a great time to be following folks who write, talk, and share about the craft of software, and how we can be better artisans.

仔细审查、规划和修复软件(和系统)的接缝,是我们开发软件时(无论是否有智能体助手)都能且需要运用的关键技能之一。当我看到“哦,这很容易实现……”的浪潮,以及人们伸手拿旧机器(clankers)来完成工作时,我认为这比以往任何时候都更重要。这是一个关注那些撰写、谈论和分享软件工艺以及如何成为更好工匠的人们的绝佳时机。

Hopefully it’s obvious, but there’s never a single answer — a panacea. It’s always about tradeoffs, choosing what makes sense for the problem at hand. With the help of a lot of great minds sharing their thoughts — both now and going back decades — we have a great tool chest for this work. It’s about picking, or reworking to move to a better choice, the right abstractions.

希望这显而易见,但从来没有单一的答案——没有灵丹妙药。它总是关于权衡,选择对当前问题有意义的事情。在许多伟大头脑分享他们想法的帮助下——无论是现在还是追溯到几十年前——我们拥有了一套伟大的工具箱用于这项工作。它是关于挑选,或者重新加工以转向更好的选择,正确的抽象。

It’s core is managing the cognitive load, learning which pieces we need to be stable, and where we want our work to flex and bend (and how). And yes, I w

其核心是管理认知负荷,学习哪些部分我们需要保持稳定,以及我们希望我们的工作在哪里灵活和弯曲(以及如何)。是的,我 w