Better AI models are just the start
How AI-era operating model redesign needs to evolve
Hi there,
In today’s newsletter:
Benchmarks are getting saturated
AI-era operating model redesign
Shipping use cases to production
Implementing this playbook (mega-prompt for paid subscribers)
Benchmarks are getting saturated
Every few months, we get much better AI models. A benchmark is a set of tasks of varying complexity that gives us a basis for comparing Claude Fable 5 with Kimi K3.
Benchmarks are getting saturated - in June, Artificial Analysis removed IFBench from its Intelligence Index because it no longer distinguished frontier models well enough.
Artificial Analysis recently launched Doomscroll, a randomly shuffled feed of charts showing how AI models compare across many measures.
Many well-specified tasks can already be done faster and more cheaply by AI agents.
You don’t need to spend your money running a pilot project to prove that an AI model can extract text from an invoice and turn it into something useful.



