Marketing vs Reality

Skepticism about claims without showing the product, comparison to productivity porn, buzzword chasing around 'harness,' suspicion of OpenAI's motives

← Back to Harness engineering: Leveraging Codex in an agent-first world

Skeptics dismiss the "agent-first" narrative as marketing-driven hype, noting that many of the touted breakthroughs are already standard developer practices or "productivity porn" that prioritizes complex system-building over actual shipping. Critics are particularly frustrated by the lack of tangible products or public repositories to back up these grand claims, arguing that a perfect architectural "harness" is meaningless if it merely creates an "empty safe" with no functional code inside. The reliance on "lines of code" as a success metric is widely rejected as a flawed proxy for engineering quality, while the use of LLM-generated clichés in the messaging is seen as a form of "casual gaslighting" that obscures a lack of verifiable substance.

29 comments tagged with this topic

View on HN · Topics
They never specified what exactly the product was, without which it's impossible to judge the post. For some reason most of the uses of "agents" are to build yet other AI products, it's turtles all the way down. Maybe that says more about the field of harnesses than it does about the power of "agents".
View on HN · Topics
Feels like the active discovery going on is trying to understand what is computer vs what is AI, for every product. Agents help a ton with the discovery, but the act of building a product needs a deeper level of thought and validation to make it actually better than what came before. So IMO what you see is people still learning what needs to be understood and crafted first hand to make a product better (including economics) We’ll get there if more of us try
View on HN · Topics
Oh I definitely agree that AI can and will help create great software. It's just that creating great software isn't really the SV/VC/big tech business model or main goal.
View on HN · Topics
> What if AI lets you create new versions of those tools, but without the enshitification? I'm not sure I fully understand what you're saying here. Isn't the value of these tools almost entirely independent of their actual software? That is, we have many good open source, self-hostable forges (Forgejo, sr.ht, etc.), lots of great music player software (Jellyfin, Symphonium, etc.), and decent maps software (OsmAnd and Organic Maps). People use GitHub, Spotify, and Google Maps -- perhaps even _put up_ with their often bad/glitchy software -- because of network effects (all three) and content/licensing partnerships (Spotify/GMaps). That proprietary data isn't something AI can help you with, right?
View on HN · Topics
It really depends on the use-case. For example, my most starred github repo is a tool to convert Spotify playlists to YouTube Music (that was done pre-AI). Github depends on what issues you have with it, what your use case is, and whether you can leverage some of the network effects via API from the github source. Maps, same story.
View on HN · Topics
This is a lot tamer than what Claude Code's team claims tbf.
View on HN · Topics
I've been doing the same experiment in tsz[1] for a while now (the same past five months in fact) and I have come to very similar conclusions. Lots of harness to enforce good architecture splits. Lots of tests and CI. My point of working on tsz is to learn how to do very big projects with AI. Eventually the same workflows and attitude can be leveraged to build customer product apps with UI as well. I see that OpenAI is leveraging automated browser testing and even videos as part of their workflow. I think as models get better this direction for making software would eventually make sense. I don't think we're there yet though. But at least, unlike OpenAI vague claims I can share the output with you to see! Most of the solutions that offer a very high level of automation like Lovable are a bit too optimistic and solutions are not tightly coupled with lots of automated testing. [1] https://github.com/tsz-org/tsz
View on HN · Topics
> We tried this early on — used ChatGPT as "project manager" to set up the entire harness before writing any code. After a week it produced 140+ docs of rules, architecture, frameworks. Zero lines of code. When we finally brought in another tool to review, the verdict was: "a perfectly secure empty safe." The harness was immaculate. There was just nothing inside it. > > Harness matters, but if you're not shipping code alongside it, you're just writing fiction.
View on HN · Topics
I'm not an AI skeptic but I'm skeptical of the intent of this article. It makes great claims about agent-first engineering and tries to make a real case based on a real product, with real users, and a real team that's been growing — all without even saying what was built or showing it, just like every other AI hype article.
View on HN · Topics
At the time we wrote the article we hadn’t released the product and weren’t ready to talk about it. It was an internal prototype that looked very much like the current Codex app.
View on HN · Topics
And this thread too is filled with users that "I also have done this or that" but bar one user, nobody followed up with any link to anything.
View on HN · Topics
I wish these breathless blog posts would actually try to be more didactic. For example, actually doing a walkthrough of how to set up these allegedly super powered workflows and concrete demonstrations. I’m not an AI skeptic. Rather I’d don’t want to miss out on any actual super powers.
View on HN · Topics
A lot to these blogposts are trying to catch on the next buzzword "harness". It's almost close to the productivity porn mindset that we witnessed 10-15 years ago where creating the complicated system is more exciting than using the system for daily tasks.
View on HN · Topics
we interviewed Ryan here: https://www.latent.space/p/harness-eng and he gave a talk version of it in london: https://www.youtube.com/watch?v=am_oeAoUhew
View on HN · Topics
> It's surprising to me that people who know about reward hacking choose a simple objective like lines of code generated as a signal for quality. The simple answer is that promoting locs as a relevant metric is also reward hacking. Is it easier to promote big loc counts as a key metric, or is it easier to prove agentic engineering against harder metrics? On a more general note, software practice marketers have been pushing in that direction for quite a while. "You need cloud", "Here's how to do agile at scale", "microservice everything", etc.
View on HN · Topics
People want to do X, so the metric is how much X can be done. Everyone is over-complicating the explanation. The answer for "why are we fixating on this bad metric" is almost always the same pattern. Broad audiences need simple metrics to talk about. If the metric itself requires nuance, it's hard to communicate and hard to reason about. It's easier to push the need for nuance from understanding the metric itself down the road to where the metric is applied, which allows everyone to ignore it in immediate conversation.
View on HN · Topics
I don’t think the flex here is the amount of code alone. Their goal is to show that AI can improve productivity, the number of lines is just the proxy to that. This article is a marketing piece after all. Now someone can argue that lines of code are not a good proxy of engineering productivity, but I wouldn’t be surprised if the audience they target with this content is not the HN commenters of this thread.
View on HN · Topics
This would be much more convincing if the repos, issue trackers, etc. were accessible.
View on HN · Topics
Isn't this essentially normal AI usage and what everyone has been doing for 6 months?
View on HN · Topics
I understand that the’ve written zero lines of code for this application, but would it kill them to write a few lines of the blog post by hand? Forcing readers to wade through an unceasing string of LLM clichés demonstrates the opposite of the point you’re trying to make—that the consumers of your work are worse off because you exercised no human judgment in creating it.
View on HN · Topics
Given that we can code at 10x speed for at least half a year, one would expect to see at least some pieces of machine-created software with 5 years' worth of equivalent human engineering work. Anyone know some?
View on HN · Topics
But this is almost what we have been doing for the last 3/5 months, isn’t?
View on HN · Topics
Well to a lot of people this is still a foreign concept.
View on HN · Topics
I guess orders of magnitude ain’t what they used to be.
View on HN · Topics
The world is now agent-first already?
View on HN · Topics
Individual voices aren't strong enough to drown the marketing machine. Artists and writers are unionized, why they have a more powerful collective voice. Second, there are enough peole for which their jobs are very well paid and too cozy to dare to rock the boat. The economy and job market isn't so hot either at the moment for people to quickly be able to jump ship. Can you even be sure that you find a tech company that isn't jumping head first onto the AI hype train? Even politicians can't have enough of AI in their mouth.
View on HN · Topics
I for one am not protesting because I know that this is bullshit marketing nonsense. Look at reliability metrics of OpenAI, they’re terrible. Everyone knew a long way ahead that it’s a scam, now they’re cranking up pricing and trying to rug pull. There will be a lot of developers who will come out very well once the stock tanks. That’s my two cents
View on HN · Topics
> in an agent-first world casual gaslighting
View on HN · Topics
> Over the past five months, our team has been running an experiment: building and shipping an internal beta of a software product with 0 lines of manually-written code. This is such a common thing among software engineers nowadays that I was very surprised that OpenAI would open with that line as if it were mind blowing. But then I saw it was published in February and OP is just reposting it to farm karma.