Discussion about this post

User's avatar
Option AI's avatar

The scale of this testing makes the findings especially useful. The ability to manage subagents, monitor experiments, and preserve coherence across long projects feels more consequential than any single benchmark. I would still compare costs at the level of reliable completed work, including human review and operational risk, rather than model runtime alone.

Oli's avatar

the SaaS replacement part is what gives me pause — building your own GitHub+Vercel is fun until the model that wrote it ships a new version and you're suddenly the maintainer of code you never really understood. $6/hr rents the engineer, but the codebase you end up owning is the real subscription.

10 more comments...

No posts

Ready for more?