← All essays

On living · Optimization

More or Less

A physicist's framework for changing your life — feel the local slope, take one small step downhill, and repeat.

Yesterday I went bike riding around Lake Tahoe. Sun on the water, cold air, legs working, nothing to optimize. It was wonderful, and somewhere on the second climb a very simple thought arrived: I want more of this in my life.

A few hours later I opened Slack to a thread on fire — two groups misaligned, everyone mad, everyone pulling in a different direction, nothing produced except heat. And the mirror-image thought arrived just as cleanly: I want less of this.

That’s the whole framework. More or less. For anything that happens in your life, ask one question: did that give me energy, or take it? Do I want more of it, or less? Then step in that direction. Not a five-year plan. Not a reinvention. One step.

I’ll be the first to admit this doesn’t sound like much of a framework. But I’ve come to believe its simplicity is doing a lot of work, and it’s worth spelling out why.

01Algorithms for better living

I’m a physicist by training, so forgive me, but: more-or-less is gradient descent.

Gradient descent is the workhorse optimization algorithm underneath most of modern machine learning, and it is almost embarrassingly simple. Write your life as a point \(x\) in some enormous configuration space — where you spend your hours, who you spend them with, what you say yes to. Write \(f(x)\) for how much that configuration drains you. You would like to minimize \(f\). The catch: you cannot see \(f\). There is no map. All you can measure is the slope directly under your feet — the gradient, \(\nabla f(x)\) — and the algorithm is just:

\[x_{t+1} = x_t - \eta \, \nabla f(x_t)\]

Feel the local slope. Take a small step downhill, of size \(\eta\). Measure again. Repeat.

Gradient descent
Fig. 1No map needed — you only ever read the local slope and step downhill.

That’s the entire method. No global knowledge required, no model of the whole terrain. And the bike ride and the Slack thread are exactly gradient measurements: this direction is downhill (more of this), this direction is uphill (less of this).

The competing approach — the one our culture is much more romantic about — is the wholesale transformation. Quit the job, move the family, get religion, become a new person. In optimization terms, this is trying to teleport to a global optimum you’ve imagined but cannot actually see. And that’s precisely the problem: you can’t see it. The imagined destination is a model, and the model is usually wrong in ways you only discover after you’ve paid the moving costs. Big change is high-variance, requires many pieces to land simultaneously, and offers no cheap way to correct course mid-jump.

Gradient descent asks almost nothing of you. Each step is small, cheap, and reversible. If it turns out to be wrong, the next measurement tells you, and you adjust. The algorithm is robust to a noisy world and an imperfect self — which is fortunate, because that’s the world and the self we’ve got.

02What gradient descent is good for

It works without a map. You do not need to know what your ideal life looks like. This is a bigger advantage than it sounds. Most of us are terrible at predicting what will make us happy at five years’ distance, but we are pretty reliable narrators of whether today energized or drained us. Gradient descent only ever asks the question we can actually answer.

It’s cheap. No step requires you to coordinate a thousand pieces. You don’t have to renegotiate your identity to ride a bike more often.

It compounds. Small steps, taken consistently, cover surprising distances. The cliché is numerically real:

\[1.01^{365} \approx 37.8\]

One percent better per day is not 365% better per year — it’s thirty-seven times. I don’t take the exact number seriously (life is not a clean exponential), but the shape of the claim is right: energy recovered on the margin is the input to everything else, and inputs compound.

It’s self-correcting. Because each step is followed by another measurement, errors don’t accumulate. Compare that to the wholesale change, where an error in the initial vision can go undetected for years.

03What it’s not good for

Gradient descent has one famous failure mode, and it’s worth naming honestly: local optima.

Sometimes you’re standing in a basin. Every direction you can step from here feels worse — the local gradient is zero, \(\nabla f(x^*) = 0\) — and yet you know, or suspect, that there’s much better ground across the valley. The job that’s fine but not right. The city that works but doesn’t fit. More-or-less will keep you comfortably pinned to the mediocre spot forever, because escaping requires a stretch of steps that each feel worse before anything feels better.

Local optimum and the valley crossing
Fig. 2The escape hatch: a one-time valley crossing to higher ground the local slope can't reach.

This is the one legitimate use case for getting religion. Wholesale change isn’t the rival of incremental change — it’s the escape hatch you pull when you have actual evidence you’re trapped. And the evidence looks specific: you’ve been running more-or-less honestly for a long time, and the gradient has gone flat. Nothing local moves the needle anymore. That’s when you cross the valley. Not because transformation is romantic, but because the local algorithm has told you, by its own exhaustion, that it’s out of moves.

Two other honest limitations.

The measurement problem

The algorithm is only as good as the signal, and some things read as unpleasant in the moment but compound into things you’re glad to own — what climbers call type-two fun. In utility terms: the instantaneous read is negative, \(u(t_{\text{now}}) < 0\), but the integrated, discounted value is positive:

\[\int_{0}^{\infty} e^{-\rho t}\, u(t)\, dt > 0\]

The moment lies; the integral tells the truth.

My best real-time discriminator: does the discomfort do work, or is it waste heat? The climb hurts, but the pain is building something — fitness, a summit, a memory. The Slack flame war hurts and builds nothing; it’s energy dissipated as friction, a thousand vectors summing to zero. And there’s a retrospective tell: a week later, has the sign flipped? The hard ride reads as gladness. The flame war is still just sour.

The retrospective tell
Fig. 3A week later the sign flips for work; friction just stays sour.

The sensor itself

Here’s the thing I’m most sure of in this whole essay: when I don’t sleep enough, everything reads worse. The bike ride is less sweet. The Slack thread is more enraging.

In estimation language: you never observe the true gradient, only a noisy version of it,

\[\hat{g} = \nabla f(x) + \varepsilon, \qquad \varepsilon \sim \mathcal{N}(b, \sigma^2)\]

and sleep sets both parameters. Well-rested, the noise \(\sigma\) is small and the bias \(b\) is near zero — your reads are trustworthy. Under-slept, \(\sigma\) blows up and \(b\) goes negative: everything is noisier and systematically reads worse than it is. The noise you can live with — stochastic gradient descent, the version that powers all of deep learning, converges just fine on noisy-but-unbiased estimates. The bias is the killer. A biased sensor doesn’t slow the algorithm down; it marches it confidently in the wrong direction.

Sleep sets the noise and bias on the sensor
Fig. 4Rested, the sensor is tight and unbiased; under-slept, it's noisy and biased negative.

Which is why sleep isn’t another item on the more-or-less list. It’s the gain on the sensor. A few variables in life don’t get sorted by the algorithm; they set the fidelity of the sorting. Calibrate the instrument before you trust the readings.

04Isn’t this… unambitious?

You might read all this and think: this is a framework for riding bikes more and muting Slack channels. Where’s the ambition? Where’s the leap?

Fair. And there’s some truth in it — gradient descent, by construction, never proposes the moonshot.

But here’s what I’d say back. Ambition runs on surplus. The big, risky, valley-crossing moves require slack — energy, attention, money, time — and slack is exactly what incremental optimization produces. Every drained hour you recover on the margin is an hour available for something larger. The person who has spent a year quietly accumulating more of what energizes them and less of what doesn’t is, almost as a side effect, the person who can afford the ambitious bet when it appears.

So the modest algorithm and the bold move aren’t in tension. One funds the other. Run more-or-less as your default loop, let it free resources on the margin, and save the wholesale transformation for the rare moment the flat gradient tells you it’s genuinely time.

Until then: more bike rides. Less flame wars. Feel the slope. Take the step.

Jack Challis builds governed AI systems for regulated industries.

Book a 20-min call →

More essays

  1. Sep 2026Through the Looking Glass→
  2. Apr 2026Vibe DevOps: the boring parts that make it work→
  3. Feb 2026Cheap or expensive is the wrong question→