Hacker Newsnew | past | comments | ask | show | jobs | submit | markbao's commentslogin

And to be specific, the METR study was using the Cursor harness with Claude Sonnet 3.5/3.7, along with other models of that era of the participant’s choosing.

Which is ancient at this point, and half a year older than the November 2025 inflection point when agentic coding got really good.

The original article is from August 2025, and the overall message to not trust ‘how it feels’ and rather measure outcomes seems right to me despite the outdated figures. On my team at least, we are seeing a noticeable inflection in work shipped with AI according to Weave.


And the comment you’re replying to, which is entirely a reasonable opinion, was flagged, proving their point. I’ve been here for 18 years and similarly alienated by the cynicism here.


I don’t think this is entirely wrong, in that there is a ‘class thing’ about riding the bus, but it’s more practicality than a class marker for a lot of people.

- In SF you can either walk 1-10 minutes to the bus, wait 0-15 minutes for the bus, tap on (while watching most other passengers evade the fare), get dropped off, and then walk 1-10 minutes to your destination… or spend an additional $5-10 to get Ubered door to door at a third of the time. First and last mile are real costs.

- In SF I Uber, unless Muni/BART is a straight shot. In NYC I take the subway. It’s not really a class thing. In NYC it takes longer to Uber much of the time and it costs several more times than the subway. You still have a 1-5 minute first and last mile problem, but headways on trains is decent and above ground taxis are incredibly inconsistent with traffic.

That about matches up with the experience with social groups in similar classes in these areas too. Most of my SF friends Uber. Most of my NYC friends take the subway.


This comment is legit! So many of these comments here are wildly biased to people's own personal experiences (usually the male/female Karens are the most noisy). You make many good points here.

My question(s): Why do you think Uber works so well in SF? Why don't they get trapped on Market Street with crawling speeds?

I lived in NYC (Manhattan) many years ago and I always felt that when I needed a taxi (cold/snow/rain), they were hard to get. As a result, I almost never took a classic yellow cab in NYC/Manhattan.


I’m all for standardization but you could just use this argument to keep any suboptimal status quo in place. XML is good enough and a standard. SOAP is good enough and a standard. etc.

The claim is that Conventional Commits are good enough and standardized enough that having another structure isn’t really worth it. But “worth it” is subjective. I’d say that if you are making commits and reading PRs every work day, and the conventional commits format causes a little bit of friction, that friction can add up. Having another option other than seeing conventional commits as a law of nature gives options for teams who prefer it. (Most teams aren’t generating changelogs anyway.)


The new structure needs to be "better enough" that it overcomes the built-in deficits of the older structure, and it can't introduce so many new problems that make it a net negative.

JSON was definitely a huge improvement in simplicity and readability compared to XML for many contexts. Similarly REST a much better option than SOAP (and all of these are examples of the general over-engineered, design-by-committee architectures that came out of the late 90s/early 00s - see also the original EJB spec - before a larger trend towards simplicity and ease of use won out).

But it this case, a lot of the differences just feel like potayto/potahto, i.e. minor stylistic preferences. And I have been in jobs where more than 50% of my time was doing code reviews, and while often there were e.g. some linter rules or whatever that I found suboptimal, it was a lot easier to just go with it than waste the energy to have the battle over why I think for loops are actually OK.


That seems a bit reductive. Even with humans, there’s a range of interpretations and ways that something can be built or a task completed. Engineers remember stuff so you don’t have to keep repeating yourself. Skills are a way to describe your outcome without similar repetition.


I don’t really buy that Claude Design will remove all the complexity around design. Vibe-coded apps using Claude look simpler because they are simpler. They’re not a gigantic product suite with extremely specific UI components tailored to each use case. The ‘simplicity’ is an illusion coming from conflating the complexity of a bicycle (a vibe coded app) with an airplane (an app like Figma).

Building the same design system component in code versus in Figma is going to be slightly more succinct in code; Figma’s primitives don’t have the sort of conditionals and control flow that code has. But code is much less malleable than drawing on a screen, and creative freedom is harder to achieve in code.

UI can fix the gap where code feels less malleable than Figma, but complexity comes largely from the worlds that humans create, and humans apparently want to create 8 modes for 4 products and 2 light/dark modes. If you want the same setup in Claude, it’ll be a little easier to maintain, but not much less complex.


Most of the times people just want a bike or a car. Not everybody needs an airplane. This is going to hit Figma very hard.


I admit I'm having a visceral reaction to this analogy. A bicycle is a sophisticated product whose form is almost pure function. Despite being apparenty simple, almost no regular person can draw even a reasonable facsimile of a bicycle from memory ( https://www.youtube.com/watch?v=L0_vXZ-3LFU ). Which is to say, for actually designing a functioning bicycle, the devil is in the details, and details are exactly where vibecoded apps fall down. Our lower bound for this analogy should instead be the downhill go-kart cobbled together from scrap wood you found in the dumpster.


> Most of the times people just want a bike

with a pelican on it


that runs on a local model and looks better than Opus


Milestone achieved with Qwen3.6


Figma has been in trouble for a while. All the designers at my company switched to Cursor nearly a year ago. They made live mockups that don’t even need a spec to implement, because the expected behavior is already captured in the prototype. Claude Design makes it just marginally easier.


Not to mention all the people hiring UX just because they don't want to deal with it themselves, not because they need something that requires a lot of skill.


Stitch has been around for a few months from hole and it does a better job than this. I bet designers are in the honeymoon phase of people don’t know this exists and this does my whole job phase.


The people that want just a bicycle wasn't going to buy figma


Are those people using Figma?


Feels like a lot of them are using Bootstrap, curated Tailwind elements or Wordpress templates. In which case the challenge Anthropic faces is convincing them the extra flexibility and magical "type into prompt" approach to customization of Claude Design is worth the compute cost...

Figma is answering a different question which is "are you prepared to spend time and money on full time designers to have pixel perfect layouts agreed with managers and consistent across platforms" and non-AI tooling has been orders of magnitude faster and cheaper at generating something that looks good enough as end results rather than mockups since before it existed.


> Vibe-coded apps using Claude look simpler because they are simpler.

This one seems pretty decent?

* https://github.com/2GT-Media-Group-LLC/mikrotik-manager

It was recently Vibe coded by an IT savvy non-dev (or non-traditional-dev anyway ;> ):

* https://www.youtube.com/live/qhZ5q6tlwq0?si=tAtmb04_WwhyaGkn...

Note - If you want to try it out DO NOT deploy it on a hostile network. ie the public facing internet


Isn’t Figma actually the tool used to plan and design the “airplane” level app? (Which of course does not mean it can not be airplane level itself…)


Making the complexity simpler is the whole thing. Any software that does it wins.


Goody | Remote | $150–250K + equity and benefits | US and Canada | Full-time

I'm Mark, the technical co-founder and CTO at Goody. We're building a gifting product that every business can use to recognize employees, retain customers, and accelerate sales. Despite being something everyone does, gifting is one of the areas of commerce yet to be disrupted, and we're working on building the best and most delightful product in this space.

Our product is used by Google, Stripe, Anthropic, Meta, NBCUniversal, Notion, and others, and we also offer a developer API for commerce. Tech stack is Ruby + React + TypeScript, though we're flexible on backend language if you know Python or Node.js better. All roles are full-stack.

We're coming off of a big year and planning for scale in 2026 with openings in our engineering team.

• Staff Software Engineer ($200–250K) — for those who ship at a startup pace and have a great eye for detail

• Senior Software Engineer, Customer Engineering ($150–200K) — if you like to hear a customer request in the morning and tell them it’s shipped in the afternoon

• Senior Software Engineer, Growth ($150–200K) — be the engineer who has the most direct impact on our growth

We're looking for people who have great startup energy, want to win, and bring great vibes to our tight-knit team.

https://jobs.ongoody.com/#hn


> What’s become more fun is building the infrastructure that makes the agents effective.

Solving new problems is a thing engineers get to do constantly, whereas building an agent infrastructure is mostly a one-ish time thing. Yes, it evolves, but I worry that once the fun of building an agentic engineering system is done, we’re stuck doing arguably the most tedious job in the SDLC, reviewing code. It’s like if you were a principal researcher who stopped doing research and instead only peer reviewed other people’s papers.

The silver lining is if the feeling of faster progress through these AI tools gives enough satisfaction to replace the missing satisfaction of problem-solving. Different people will derive different levels of contentment from this. For me, it has not been an obvious upgrade in satisfaction. I’m definitely spending less time in flow.


If you save 3 hours building something with agentic engineering and that PR sits in review for the same 30 hours or whatever it would have spent sitting in review if you handwrote it, you’re still saving 3 hours building that thing.

So in that extra time, you can now stack more PRs that still have a 30 hour review time and have more overall throughput (good lord, we better get used to doing more code review)

This doesn’t work if you spend 3 minutes prompting and 27 minutes cleaning up code that would have taken 30 minutes to write anyway, as the article details, but that’s a different failure case imo


> So in that extra time, you can now stack more PRs that still have a 30 hour review time and have more overall throughput

Hang on, you think that a queue that drains at a rate of $X/hour can be filled at a rate of 10x$X/hour?

No, it cannot: it doesn't matter how fast you fill a queue if the queue has a constant drain rate, sooner or later you are going to hit the bounds of the queue or the items taken off the queue are too stale to matter.

In this case, filling a queue at a rate of 20 items per hour (every 3 minutes) while it drains at a rate of 1 item every 5 hours means that after a single day, you can expect your last PR to be reviewed in ((8x20) - 1) hours.

IOW, after a single day the time-to-review is 159 hours. Your PRs after the second day is going to take +300 hours.


This is the fundamental issue currently in my situation with AI code generation.

There are some strategies that help: a lot of the AI directives need to go towards making the code actually easy to review. A lot of it it sits around clarity, granularity (code should be committed primarily in reviewable chunks - units of work that make sense for review) rather than whatever you would have done previously when code production was the bottleneck. Similarly, AI use needs to be weighted not just more towards tests, but towards tests that concretely and clearly answer questions that come up in review (what happens on this boundary condition? or if that variable is null? etc). Finally, changes need to be stratified along lines of risk rather than code modularity or other dimensions. That is, if a change is evidently risk free (in the sense of, "even if this IS broken it doesn't matter) it should be able to be rapidly approved / merged. Only things where it actually matters if it wrong should be blocked.

I have a feeling there are whole areas of software engineering where best practices are just operating on inertia and need to be reformulated now that the underlying cost dynamics have fundamentally shifted.


>Finally, changes need to be stratified along lines of risk rather than code modularity or other dimensions.

Why don't those other dimensions, and especially the code modularity, already reflect the lines of business risk?

Lemme guess, you cargo culted some "best practices" to offload risk awareness, so now your code is organized in "too big to fail" style and matches your vendor's risk profile instead of yours.


> Why don't those other dimensions, and especially the code modularity, already reflect the lines of business risk?

I guess the answer (if you're really asking seriously) is that previously when code production cost so far outweighed everything else, it made sense to structure everything to optimise efficiency in that dimension.

So if a change was implemented, the developer would deliver it as a functional unit that might cut across several lines of risk (low risk changes like updating some CSS sitting along side higher risk like a database migration, all bundled together). Because this was what made it fastest for the developer to implement the code.

Now if AI is doing it, screw how easy or fast it is to make the change. Deliver it in review chunks.

Was the original method cargo culted? I think most of what we do is cargo culted regardless. Virtually the entire software industry is built that way. So probably.


> when code production cost so far outweighed everything else, it made sense to structure everything to optimise efficiency in that dimension

Oh, for sure. Those people making electro-mechanical computers at the end of the 19th century certainly did that a lot.


You are considering a good-faith environment where GP cares about throughput of the queue.

I think GP is thinking in terms of being incentivized by their environment to demonstrate an image of high personal throughput.

In a dysfunctional organization one is forced to overpromise and underdeliver, which the AI facilitates.


If your team's bottleneck is code review by senior engineers, adding more low quality PRs to the review backlog will not improve your productivity. It'll just overwhelm and annoy everyone who's gotta read that stuff.

Generally if your job is acting as an expensive frontend for senior engineers to interact with claude code, well, speaking as a senior engineer I'd rather just use claude code directly.


Linting, compiler warnings and automated tests have helped a lot with the grunt work of code review in the past.

We can use AI these days to add another layer.


Except that when you have 10 PRs out, it takes longer for people to get to them, so you end up backlogged.


And when the PR you never even read because the AI wrote it gets bounced back you with an obscure question 13 days later ..... you're not going to be well positioned to respond to that.


Goody | Remote | $150–250K + equity and benefits | North/South America | Full-time

I'm Mark, the technical co-founder and CTO at Goody. We're building a gifting product that every business can use to recognize employees, retain customers, and accelerate sales. Despite being something everyone does, gifting is one of the areas of commerce yet to be disrupted, and we're working on building the best and most delightful product in this space.

Our product is used by Google, Stripe, Anthropic, Meta, NBCUniversal, Notion, and others, and we also offer a developer API for commerce. Tech stack is Ruby + React + TypeScript, though we're flexible on backend language if you know Python or Node.js better. All roles are full-stack.

We're coming off of a big year and planning for scale in 2026. We have a few new roles to accelerate our growth.

• Staff Software Engineer ($200–250K) — for those who ship at a startup pace and have a great eye for detail

• Senior Software Engineer, Customer Engineering ($150–200K) — if you like to hear a customer request in the morning and tell them it’s shipped in the afternoon. US and Canada only for this one

• Senior Software Engineer, Growth ($150–200K) — be the engineer who has the most direct impact on our growth

We're looking for people who have great startup energy, want to win, and bring great vibes to our tight-knit team.

https://jobs.ongoody.com/#hn

My email is open for any questions: mark@ongoody.com


Your position detail pages apparently require WebGL to prevent them from "crashing." That and "complex SPA" in the text (found from reader mode) tell me you might very well be overengineered.


Yes, this is what the world needs. Definitely better than a living wage and health care.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: