Hacker Newsnew | past | comments | ask | show | jobs | submit | manveerc's commentslogin

Totally agreed. I sometimes wonder if they are making the model "lazy" with each iteration, it keeps getting better at avoiding work.


This is why Fable was so good. It followed instructions and it was in no way lazy.


People keep making comments about fable like this? You could only use it for what like a week? How is that at all enough time to evaluate? Opus 4.6 didnt suffer from this problems for a hot minute and then when newer models were released it got worse. I think they change a ton behind the scenes and allocate compute however they want, so the model you use today may behave much differently than how it behaved yesterday


> You could only use it for what like a week? How is that at all enough time to evaluate?

By observing how in 4 workdays it achieved more than Opus in ~11 days. I am my team's backend lead and the Fable 5 model finally turned the tide on my overwhelming backlog. Back to Opus and I have to treat it like special-education kid multiple times a day.


The ~72 hours I had access to Fable were by far the most productive I've had in months. Re-wrote massive parts of my codebase and caught a ton of bugs and logic issues that had silently slipped through before. I went over my subscription limit and immediately kept paying the API price to keep going. It was that good.


Heh, it's not crazy if you're here in the Bay: I know multiple people who more-or-less disappeared for days when Fable came out because they were running their benchmarks, and only emerged blinking into the sunlight when the USG banned it. That's just how things are here now, most people are normal but there are some serious LLM dope addicts out and about.


It was a pretty stark difference. I had the opposite problem where it did too much and overshot what I wanted from it so I certainly assume that if it had stuck around it would have gotten tuned back a bit pretty quickly.


You didn't really have to use it more than a day honestly to tell what kind of shocking paradigm change it was. Man do I miss it.


For me claude-fable-5 failed to follow the instruction following test I'm making against various models https://github.com/marcindulak/claude-fails-to-follow-claude...


I've been seeing LLMs act lazy from the very beginning. They got a little better but smaller models really only want to have a single task given to them. Mythos at least does work. RIP


Congratulations on the milestone! Arcade is the best product for solving reliability and security in the action layer. It’s a no-brainer, given the team.


And time to value is another dimension. In your case finding the right person, scheduling the project, you implementing and delivering at best is a week if not more. With AI they get it in seconds. I may have fudged the numbers but the scale relative gains likely will be same.


WhatsApp had real network effects built in, and network was the moat. Don’t think Cursor has any real moat.


Well i built an equivalent of OpenClaw using Claude Code and hooking it up with WhatsApp. For mew I'm currently using it for following things

1. Morning brief + meeting preps 2. Managing client work and action items (tracking status, deliverables, etc) 3. Executing our AI workflows on my laptop. We have built several AI workflows for our agency and this setup gives the ability to seamlessly execute and control them through both mobile and desktop

Next on my to-do list is to build additional workfows for me and my wife around family logistics (travel, childcare, etc)


In my opinion sites that want agent access should expose server-side MCP, server owns the tools, no browser middleman. Already works today.

Sites that don’t want it will keep blocking. WebMCP doesn’t change that.

Your point about selenium is absolutely right. WebMCP is an unnecessary standard. Same developer effort as server-side MCP but routed through the browser, creating a copy that drifts from the actual UI. For the long tail that won’t build any agent interface, the browser should just get smarter at reading what’s already there.

Wrote about it here: https://open.substack.com/pub/manveerc/p/webmcp-false-econom...


So... an API?

Most sites don't want to expose APIs or care enough about setup and maintenance of said API.


Are you asking if Agents should use API?


Not something new. They were recording audio on Facebook app and messenger for the longest time without a people using the microphone. They were tracking people using network data. The list is pretty long.


Facial recognition would be able to detect all the strangers around you, whereas audio would surely only pick up people nearer the device, and presumably wouldn't be able to tell people apart/identify them. You're right about network data; if they're using Wifi/BT probes then they can already find and identify everyone in the vicinity.

I'm curious why you're using past tense by the way?


Oh yeah I agree it’s bad. I just meant, company has no morals. It is very data hungry and doesn’t care about people’s privacy.

And with respect to past tense, I don’t know if they still do when they were caught red handed about some of these things. Unless there is a court order I am sure they still do, but I have no proof point.


Thats a good question. I would recommend MCP for the bulk of 'chatty' soft data to keep the database clean. However, you should selectively ingest 'high value' data into ClickHouse for vector search.

For e.g. you wouldn't ingest every 'good morning' message. But once an incident is resolved, you could ETL specific threads (filtering out noise) and the resulting RCA into ClickHouse as a vectorized document. That way, the copilot can recall the solution 6 months later without depending on Slack.


The interesting part is that only one of them is software only. I get it is economists but I guess this is also telling that while AI is really talk of town in silicon valley, long term it is one, minor if I may, part of the future.


Yes! And the layoffs


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: