On AI · Forecasting
The one AI metric that matters: H₅₀
Most people track AI benchmarks that don't matter for real work. Here's the one that does: the longest task an agent completes with 50% reliability.
Why this matters more than ChatGPT demos. Real work isn’t one-shot questions. It’s chains of dependent steps. If each step has 90% success, a 10-step process only succeeds 35% of the time. Improve each step to 95%? Now you’re at 60%. Reliability compounds — and so does its absence.
Where we are right now.
- Best agents handle roughly 2.5-hour tasks at 50% success.
- Cost: about $1.50/hour for complex workflows.
- Good for: code reviews, research synthesis, routine analysis.

The exponential trend that changes everything. We’ve seen roughly 6× improvement per year in H₅₀. At this pace, we hit about 7 hours by April 2026.
What 7-hour reliability unlocks.
- Multi-repo bug fixes → testing → pull requests.
- Month-end financial reconciliation across systems.
- RFP responses with compliance matrices.
- Research reports synthesized from hundreds of sources.
The economics get wild. If costs drop as fast as capabilities grow — and they’re dropping faster — that 7-hour workflow still costs about $3 total. That’s $0.40/hour for work that today requires expensive specialists.
What smart operators are doing now. Track your own H₅₀ — the longest workflows your team actually trusts AI with. Invest in checkpoints, rollbacks, and observability. When H₅₀ crosses into multi-hour territory, entire job categories shift.
The curve is steep. The implications are steeper.