inline execution. agents execution takes 4x as long, uses more tokens and is full of holes. less time, less money, less bugs. who keeps pushing this trash
2026-07-28 12:23:25
0
nullwristmotion :
This seems like a completely useless experiment. They rebuilt an open source software that was almost certainly in its training data?
2026-07-22 07:36:50
328
jks :
Thanks I always needed SQLite to be rewritten in rust… oh wait no I didn’t
2026-07-22 03:39:03
178
Félix :
Big caveat here (not mentioned in the article) is that SQLite is in the training set for all of these. It’s likely smaller models wouldn’t be as successful at a rewrite task for a thing they had never seen before
2026-07-22 07:19:44
53
septemberpiano :
Copying something that’s so structured and so well documented is such a specific test case though I don’t really know how relevant the test really is in the wider conversation. It’s basically the perfect scenario
2026-07-22 10:12:52
42
hahsh68g :
100% bs - their coding models are absolute trash
2026-07-25 04:31:19
0
spaceb0iz :
The main thing I’m not understanding with this is the accuracy vs time. You’re telling me each of these implementations are exactly identical?
2026-07-22 04:12:39
22
richie :
composer is so good. $20 cursor plan can do so much!
2026-07-22 02:11:55
5
oriactica :
"look how fast AI recreates this extremely popular open source project it was probably trained on, after we prompted it with the 800 page manual/specification document of the project"
I mean I guess...
I just realized rust rewrites are also kind of a meme so they might've been trained on that as well
2026-07-22 13:04:27
9
v🇷🇴 :
"hey guys, if I can get a step by step guide to build something, I can build it" lmao
2026-07-22 12:51:51
21
objectivelycorrect_ :
I always tell Fable—do NOT fan out 100 Fable agents! You plan, make opus or sonnet work
2026-07-22 20:42:14
4
Beck :
Ah yes let’s do windows next
2026-07-22 06:26:28
0
James :
yes let's see the code it wrote tho...also using a manual for a literal planner is not a fair test, besides what even was the plan? just copy the steps out of the manual and prompt a model with them
2026-07-22 19:10:46
1
Behrad ™ :
No way the quality is the exact same
2026-07-22 15:29:44
1
Jonas :
Yes I understand that it is cheaper but is it the same quality?
2026-07-24 13:53:41
0
cryptoniun93 :
So Opus for planning and Sonnet for coding should be good as we have seen
2026-07-22 15:15:24
0
jant2534 :
what harness did they use? Aider?
2026-07-22 17:47:26
0
wndr :
how about they build something unfamiliar? what is the quality of the code looking like? how much time did it take them? what if we wanted to add a new feature that doesn't even exist yet?
those are all questions to be asked here
2026-07-22 08:09:20
1
James X :
This is exactly how I’ve been using Fable. It’s actually terrible at implementation. It does 10% of the work and then just gives up. But it’s great for design.
2026-07-22 07:01:40
6
wolph :
This is just an ad for cursor... they're just trying to show how well their system works. It might be true and it might be wrong depending on how they used it. But, I'll readily admit that Anthropic is really expensive when it comes to tokens, Opus 4.8 feels like 5 times compared to Gpt-5.5 for the same outputs. Fable and Gpt-5.6 sol behave different so that's harder to compare, but here too a massive difference in pricing.
2026-07-22 22:02:13
2
JeffyTurner :
People just rediscovered division of labor but with computers
2026-07-23 05:50:11
2
ideafactoryy :
used the promo tokens to ask fable to do some basic data analysis on decent sized chunks of data. it used about $2 per question. again, basic stuff like "create a pivot table"
2026-07-22 11:41:01
1
user9985185848733 :
But composer is built on kimi k2.5
2026-07-22 06:55:41
2
Viktor :
LLMs are good at passing tests: def test(): assert True
2026-07-22 06:28:34
2
bobnapibadung :
rust still cant beat c. whats the point?
2026-07-22 12:52:44
0
To see more videos from user @barneslucas_, please go to the Tikwm
homepage.