← All essays

On AI · Forecasting

The one AI metric that matters: H₅₀

Most people track AI benchmarks that don't matter for real work. Here's the one that does: the longest task an agent completes with 50% reliability.

Why this matters more than ChatGPT demos. Real work isn’t one-shot questions. It’s chains of dependent steps. If each step has 90% success, a 10-step process only succeeds 35% of the time. Improve each step to 95%? Now you’re at 60%. Reliability compounds — and so does its absence.

Where we are right now.

  • Best agents handle roughly 2.5-hour tasks at 50% success.
  • Cost: about $1.50/hour for complex workflows.
  • Good for: code reviews, research synthesis, routine analysis.
Agent 50% time horizon vs release date, with 8-month projection
Fig. 1H₅₀ on a log scale: a straight exponential from o1 to a ~7-hour projection by April 2026.

The exponential trend that changes everything. We’ve seen roughly 6× improvement per year in H₅₀. At this pace, we hit about 7 hours by April 2026.

What 7-hour reliability unlocks.

  • Multi-repo bug fixes → testing → pull requests.
  • Month-end financial reconciliation across systems.
  • RFP responses with compliance matrices.
  • Research reports synthesized from hundreds of sources.

The economics get wild. If costs drop as fast as capabilities grow — and they’re dropping faster — that 7-hour workflow still costs about $3 total. That’s $0.40/hour for work that today requires expensive specialists.

What smart operators are doing now. Track your own H₅₀ — the longest workflows your team actually trusts AI with. Invest in checkpoints, rollbacks, and observability. When H₅₀ crosses into multi-hour territory, entire job categories shift.

The curve is steep. The implications are steeper.

Jack Challis builds governed AI systems for regulated industries.

Book a 20-min call →

More essays

  1. Sep 2026Through the Looking Glass→
  2. Jul 2026More or Less→
  3. Apr 2026Vibe DevOps: the boring parts that make it work→