What are you actually maximizing — and what happens when the dice only roll once?
Expected value is a promise the universe only keeps if you play forever. You play once. This page now optimizes for that fact.
Three Quantities
Joy, happiness, and regret are different quantities. One is the experienced level. One is the gap between reality and expectations. One is the gap between the policy you ran and the policy that was available.
QTY 01
Joy
$J = A(S_t) + \beta \cdot H_t$
The experienced level. Absolute value state plus happiness scaled by expectation sensitivity.
Maximize
QTY 02
Happiness
$H_t = \sum_i \lambda_i(i_t - E_t^i)$
The expectation gap. Reality minus expectations, dimension-weighted. Drives action when negative.
Feel
QTY 03
Regret
$R = \int (J^* - J)\, dt$
The counterfactual gap between the policy you ran and the best policy available.
Minimize
Before The Maximization — Choose Your Objective
The Life Metaproblem
Before the framework can maximize anything, you have to choose what you are maximizing.
That choice is called the contribution function $C$. The framework optimizes given $C$ — it does not pick $C$ for you.
Different choices of $C$ produce entirely different lives, even with identical machinery underneath.
The standard choice for joy. If you are maximizing joy weighted by survival, then $C(S_t, X_t^\pi) = \omega_t \cdot J(S_t)$.
The rest of this page assumes this choice.
Other choices are real. If you are maximizing Olympic gold, $C$ pays out a one-time bonus when $h_t$ at age 25 exceeds the threshold.
If you are maximizing legacy, $C$ depends on the states of other people, summed across them.
If you are maximizing wealth, $C$ collapses to $\lambda_w \cdot w_t$ alone. Each is a different game.
The metaproblem. Choosing $C$ is Step 1 of the Six Steps, applied to life itself.
Most people never make this choice consciously. They inherit one — from culture, family, peers — and run optimization
inside an objective they never selected. That is how someone wins at life by every conventional measure and still feels hollow.
They optimized $V^C$ beautifully. $C$ was just not theirs.
$V_t^C(S_t)$ is the value of being in state $S_t$ at time $t$, conditional on having chosen objective $C$.
The agent searches over policy $\pi = (f \in \mathcal{F},\, \theta \in \Theta)$.
The contribution function is $C = \omega_t \cdot J(S_t, E_t)$ where $J = A(S_t) + \beta H_t$.
The change from earlier versions of this page is the operator out front. The objective is no longer aggregated
by a bare expectation. It is aggregated by $\rho_{\varphi}$ — a weighting over which of your possible lives
count, and how much. The parameter vector now carries its two shape parameters:
$\theta = (\theta_{-E},\, E^*,\, a,\, b)$.
Expectations $E^*$ still enter the contribution in two places — the happiness gap $H_t$ and the action elasticity —
and are called out for that reason. The weighting is the second distinguished sub-parameter, and
Component 10 is where it earns its place.
Set $\varphi \equiv 1$ and you recover the risk-neutral objective this page used to carry.
Five real-state variables — health, wealth, relationships, purpose, survival —
plus one expectation variable per real-state dimension.
Expectations are partly inherited from starting state, partly shaped by policy.
$h_t$ health
drives survival directly through biological pathway
Joy is Absolute Value State plus Happiness scaled by Expectation Sensitivity $\beta \geq 0$.
$\beta$ controls how strongly the gap between reality and expectations amplifies or crushes
the absolute experience. High $\beta$: the gap dominates. Low $\beta$: you mostly feel what you have.
Two people, two levels of joy. A healthy, well-connected multi-millionaire
who just moved in next to Bezos. High $A$ across every dimension — but $H_t$ is deeply
negative because his new peer group rewrote his expectations overnight. Now compare a lead
foreman at the local factory who has spent twenty years rising through the ranks, consistently
exceeding what he thought he would achieve. Modest $A$, but $H_t$ is strongly positive.
With high enough $\beta$, the foreman's joy is higher. That is not a paradox. That is the formula.
Stoic practice is deliberate $\beta$ reduction. Not the elimination of happiness
or the cultivation of low expectations. The decoupling of lived experience from the gap.
A low-$\beta$ person feels what they have.
The weighted sum of your actual state variables — health, wealth, relationships, purpose —
independent of what you expected to have. This is what you have, full stop.
The $\lambda$s drift with age: kids weight stuff and immediate experience; adults weight
relationships and purpose; elders weight legacy and meaning.
Happiness is the gap between reality and expectations, dimension-weighted.
It can be positive or negative. When $H_t > 0$, reality is beating expectations.
When $H_t < 0$, expectations exceed reality — dissatisfaction.
The cancer-patient inversion. A recovering patient has positive $H_t$
because reality exceeds their lowered expectations. A healthy person who expected to feel
athletic has negative $H_t$ despite a much higher $h_t$. The level of the variable does
not determine happiness — the gap does.
Unhappiness drives action. $H_t < 0$ is dissatisfaction.
Through the elasticity $\eta_i$, that dissatisfaction converts into policy adjustments
that raise $i_t$. You do not want to eliminate unhappiness. You want it calibrated:
enough to drive the policy, not so much — scaled by $\beta$ — that it destroys the experience.
Each year contributes joy weighted by the probability of being alive to experience it.
Joy now depends on both the state $S_t$ and the expectation profile $E_t$ — both of which
are in $S_t$, but written explicitly to make the dependence clear.
The mortality rate $\mu$ depends on $h_t$ (biological pathway), $r_t$ (loneliness mortality
is robust and not fully mediated by health behavior), and $p_t$ (will-to-live, ikigai, and
post-retirement mortality effects). $w_t$ does not enter directly — it routes through $h_t$.
Expectations update each period. $g_i$ governs how: they adapt toward recent reality
($i_t$), respond to deliberate policy choices ($X_t^\pi$ — Stoic practice, reflection,
goal-setting), and absorb exogenous shocks ($W_{t+1}$ — a diagnosis, a peer group change,
a windfall that rewrites your sense of what is normal).
The fan-of-forecasts in The Lives is $g_i$ made visible over a lifetime.
Component 8
The Coupling Structure
The transition function $S^M$ encodes how each state variable feeds the others.
Strong couplings are robust and unimodal. Mild couplings are smaller. Bimodal couplings,
marked with *, can flip sign depending on the quality or type of the source variable,
not just its level.
From → To
$h_t$
$w_t$
$r_t$
$p_t$
$\omega_t$
$h_t$
—
mild
mild
mild
STRONG
$w_t$
mild
self-loop
mild
mild
—
$r_t$
mild
STRONG
—
STRONG
STRONG
$p_t$
mild
mild
mild
—
mild
* marks a bimodal coupling.
The sign depends on the quality of the source variable, not its level.
Two-level search. (1) Choose a policy class $f \in \mathcal{F}$ — the structure of the
decision function. (2) Tune its parameters $\theta$ within that class.
The expectation profile $E^*$ is part of $\theta$ — not a separate search variable, but a
distinguished sub-parameter worth naming explicitly because it enters the objective in two places:
through the happiness gap $H_t = \sum_i \lambda_i(i_t - E_t^i)$, and through the elasticity of action $\eta_i = \partial X_t^\pi / \partial E_t^i$.
The weighting $\varphi$, carried by its two shape parameters $(a, b)$, is the other named sub-parameter.
It does not change any single decision — it changes which of your possible lives the whole trajectory
is scored against.
What policies actually look like. A policy for maintaining health might be:
brush your teeth every day — the parameter is once, twice, or three times. Work out regularly —
the parameters are frequency, intensity, and type. These are simple rules.
A complex policy is identifying and pursuing a marriage partner — or deciding not to.
The parameters include what you value, what you are willing to compromise on, and how long you search.
In practice, your full life policy is a large hybrid: simple daily rules layered with long-horizon
structural bets, financial projections, and deeply embedded cultural and societal norms you may
never have consciously examined.
The explore-exploit trade-off. Searching over policies raises an immediate problem:
you cannot evaluate a policy without running it, and running it costs time.
Do you stick with what is working (exploit), or try something new that might work better (explore)?
The practical answer: you try policies, observe outcomes, and keep what works.
You trade off exploration and exploitation across long horizons.
Major unexpected life events — a job loss, a diagnosis, a move — often force a new round of
policy exploration. That is not a failure. It is the system working as designed.
New information about your state should update your policy.
Component 10 New
Which Lives You Count
Expected value is an honest objective for a casino. The casino plays a billion hands.
The law of large numbers does its work, the average shows up, the house collects.
The casino is not risk-neutral because it is brave. It is risk-neutral because it plays
every hand in the distribution. Weighting every outcome equally is the correct objective
for whoever gets to live all the lives. You get one draw.
Let $G^\pi$ be the lifetime total you accumulate under policy $\pi$ — a random variable
across the worlds you might land in:
You will never receive $\mathbb{E}[G^\pi]$. Nobody receives an average. Sort every life you
might live from worst to best and index them by $u \in [0,1]$. The life sitting at rank $u$
is the quantile $q_u(G) = F_G^{-1}(u)$. One $u$ gets drawn, once, and you spend the rest of your
life inside it.
So the question was never what your possible lives average to. The question is which of them
you built for. That is a weighting. Put a function $\varphi$ over the ranks — how much each
possible life counts when you choose a policy — and every objective anyone has ever proposed
turns out to be one shape of it:
The distribution of $G$ is the fact about the world. $\varphi$ is what you decide to do about it.
Every argument anyone has ever had about risk is an argument about the shape of this curve.
Flat. $\varphi \equiv 1$. Every possible life counts the same, which gives you
$\mathbb{E}[G]$ back. The casino's answer, and the right one for the casino, because it actually
plays every $u$.
All the mass at $u = 0$. Maximin. Wald. Build the entire life around the worst
thing that could happen to it.
A flat shelf over the worst $(1-\alpha)$. $\mathrm{CVaR}_\alpha$ — the page you
were reading a moment ago.
All the mass at $u = 1$. Maximax. Buy the ticket and assume you win it.
Mass at both ends, thin in the middle. Hurwicz's optimism-pessimism blend, 1951.
Also, empirically, most human beings: vivid about catastrophe, vivid about the jackpot, bored by
the eighty percent of the distribution they will almost certainly land in.
A hump in the middle. Plans carefully for the typical case and refuses to look at
either tail. This one has no famous name and no defenders, and it is very common.
Why CVaR is the good pessimistic shelf. $\mathrm{VaR}_\alpha$ gives you the
threshold of the bad five percent and says nothing about what lies past it — a life mildly below
the line and a life catastrophically below it look identical to VaR. $\mathrm{CVaR}_\alpha$ averages
the whole tail instead of pointing at its edge, which is what makes it the mean of your worst
$(1-\alpha)$ lives rather than a fact about one of them.
Now the question that broke the previous version of this page. Every repair on offer was pessimistic.
Why can you not weight the good tail? Hope is not obviously less rational than fear.
Because of a theorem. $\rho_\varphi$ is a coherent risk measure if and only if
$\varphi$ is non-increasing — if and only if a worse life never counts for less than a better one.
Tilt any stretch of the curve upward and you lose subadditivity, the property guaranteeing that
spreading a bet across two independent things is never worse than concentrating it. An optimistic
$\varphi$ prefers concentration. It wants everything on one number.
There is a second bill. Maximizing $\rho_\varphi$ is a concave program under the same condition,
and only under that condition. Outside it the optimizer runs to the corners and hands back the
highest-variance policy available, every time, regardless of what the distribution looks like.
Taken literally and left unconstrained, optimism does not give advice. It says buy the lottery ticket.
None of which makes optimism forbidden. It makes it expensive, and precise about the price.
You can weight the top. You are spending an axiom to do it, you should know which one, and you should
put a floor under the bad tail before you start.
The honest optimistic program
Do not blend the good tail into the objective and hope. Constrain instead:
maximize $\mathbb{E}\!\left[G \mid G \ge q_\beta(G)\right]$ subject to $\mathrm{CVaR}_\alpha(G) \ge \underline{g}$.
Swing for the upper tail, but only among policies whose bad tail stays livable. The objective is still
not concave; the constraint set is, and the floor is what stops the corner solution from being ruin.
This is Roy's safety-first criterion, and it is what people mean when they say they can afford
to take a shot.
Making it a dial. A function is not a control you can hand a reader, so take a
two-parameter family that moves continuously between the shapes above:
Two numbers. Push $a$ below one and weight piles onto the bad lives; push $b$ below one and it piles
onto the good ones. At $a = b = 1$ the curve is flat and you are the casino. Both below one gives the
bathtub. Both above one gives the hump. And the theorem stops being an axiom and becomes territory:
$\varphi_{a,b}$ is non-increasing exactly when $a \le 1$ and $b \ge 1$. That is a rectangle.
You can stand inside it or step out of it, and the chart below will tell you which you are doing.
One honesty note before you touch it: $\mathrm{CVaR}_\alpha$ is not in this family. A Beta
curve is smooth and CVaR has a hard edge at its cutoff. The family is CVaR's smooth relative, not its
container — so the chart keeps CVaR as its own preset and draws the step exactly as it is.
Which lives you are counting
Left: the weighting $\varphi$ over ranks, and the $(a,b)$ pad it comes from — the shaded rectangle is coherence. Right: the same distribution of lifetime totals, shaded where you are counting a life for more than the casino would. Drag the dot, or pick a shape.
The weighting
Your possible lives
$\varphi$ is part of the policy, not a fact about the world. It joins $\theta$ as the
pair $(a, b)$. Two people facing identical odds with different $\varphi$ run different lives, and both
can be right, because the cost of the bad tail is not the same for a 25-year-old with no dependents as
it is for a sole earner with three. The Regret section
already said the risk metric is part of the policy. This is that claim promoted from a footnote on
regret to the operator on the objective.
Read it both ways, and keep them straight. Descriptively, $\varphi$ is a fact about
a person: the shapes above are where real people actually sit, and the incoherent ones are the
crowded ones. Prescriptively, $\varphi$ is the parameter you can change on purpose. This page is
descriptive about $\varphi$ and prescriptive about $\pi$ given $\varphi$. Coherence is not a claim
that anyone behaves that way — it came from bank capital regulation, not from watching people —
and calling a weighting incoherent says what the objective stops guaranteeing, never what a person
is doing wrong.
Two thresholds, not one. Because $\rho_\varphi$ averages the order statistics, it
lands exactly $k$ standard deviations from the mean:
$$\rho_\varphi(G) \;=\; \mathbb{E}[G] \;+\; k_\varphi\,\sigma(G), \qquad
k_\varphi \;=\; \int_0^1 \varphi(u)\, q_u\!\left(\tfrac{G - \mathbb{E}[G]}{\sigma(G)}\right) du$$
That $k$ is the slope on spread, and it is what the readout reports. Coherence is one boundary —
the rectangle $a \le 1,\, b \ge 1$. The sign of $k$ is a different one, roughly $b < a$. Cross the
first and you lose a guarantee. Only after crossing the second does spread start paying, and only
then does an unconstrained search run away. The gap between them is wide, and most of the shapes
people actually hold live inside it: no longer coherent, still spread-averse.
Which leaves the instruction. Not be less risk-averse, and not be more optimistic.
Pick the shape on purpose, know what it costs, and stop inheriting it from your nerves.
A note for the careful
A spectral measure on the whole life total $G$ is static and terminal, and static risk measures are
not time-consistent: the tail you optimize at 20 is not the conditional tail you re-optimize at 50.
Nesting the measure restores consistency and gives up the clean one-life interpretation. For a life you
run exactly once, the terminal object is the honest one. Re-solve as the filtration updates, and do not
pretend the re-solve was the original plan.
Ergodicity, again
This is the same disease the log-utility move treats. The expectation is an ensemble average over a
multiverse of yous. Log utility fixes it by reshaping the contribution into a growth rate. A pessimistic
$\varphi$ fixes it by reweighting the distribution toward the tail you would have to live in. They compose.
Apply $\rho_\varphi$ to log contributions and you get both: protection against the bad draw and
respect for the fact that you compound through time, not across copies.
Where this comes from
The $\varphi$ representation is the theory of spectral risk measures (Acerbi), sitting on Kusuoka's
representation of law-invariant coherent measures; the non-increasing condition is the coherence
requirement. The two-tailed blend is Hurwicz's optimism-pessimism criterion, 1951. The observation that
real people weight both tails and ignore the middle is rank-dependent utility and cumulative prospect
theory. The convex form of $\mathrm{CVaR}$ that keeps any of this optimizable is Rockafellar and Uryasev:
$$\mathrm{CVaR}_{\alpha}(G) \;=\; \sup_{\zeta \in \mathbb{R}}\;\left\{\, \zeta \;-\; \frac{1}{1-\alpha}\,\mathbb{E}\!\left[(\zeta - G)^+\right] \right\}$$
with optimizer $\zeta^\star = \mathrm{VaR}_{\alpha}(G)$.
Ex post regret. $S_t^* \mid W$ is the state trajectory under the optimal policy
given the same starting state $S_0$ and the same realized world $W$. $S_t$ is the trajectory
under the policy actually run. Ex post regret holds the realized world fixed and asks:
given everything that actually happened, how far from the best policy did you run?
This is hindsight regret. It is real, but it slides easily into bias.
Ex ante regret. $S_t^* \mid \mathcal{F}_t$ is the state trajectory under the
best policy given only the information available at time $t$ — the filtration $\mathcal{F}_t$.
This holds $W$ fixed to what was knowable at each decision point, not to the outcome.
No one should regret not buying Bitcoin in 2010. That was not in $\mathcal{F}_{2010}$.
Ex ante regret is the fairer and more actionable quantity for a life framework.
Two properties of both. Regret is policy-level, not decision-level —
you do not regret a single bad call, you regret a policy that produced bad calls systematically.
And regret holds $W$ fixed in both versions — bad luck is not regret.
The risk metric is part of the policy. Minimizing $\mathbb{E}[R_{\text{ante}}]$ produces
one policy. Minimizing a tail-weighted regret produces a different one.
This is the same weighting and the same operator from Component 10,
pointed at regret instead of joy. Regret is a loss, so the ranking reverses: apply $\varphi(1-u)$ instead of
$\varphi(u)$ and the shelf lands on the worst 5% of regret, which is the high end. Same curve,
read backwards. The idea is unchanged — build a life robust to bad luck, even at the cost of some
expected value. Most people implicitly do this and call it being responsible. Now it has a shape.
Component 12
Expectation Elasticity
Expectations do two jobs at once. They are a reference point for happiness —
$H_t = \sum_i \lambda_i(i_t - E_t^i)$ is the gap, and happiness is what you feel.
And they are an input to the policy — high expectations drive action, low expectations keep you still.
The two roles point in opposite directions on net welfare.
$\eta_i$ is the elasticity of action with respect to expectation on dimension $i$.
High $\eta$ means raising expectations actually changes what you do.
Low $\eta$ means it just makes you miserable without changing behavior.
Per-dimension pattern.
$\eta_w$ is high — expectations shape career choices, risk appetite, negotiation.
$\eta_h$ is moderate — expectations shape lifestyle, but above a ceiling produce anxiety not action.
$\eta_r$ is high but bimodal — can drive vulnerability and repair, or trigger withdrawal, depending on attachment style.
$\eta_p$ is very high — people who expect their work to matter find ways to make it matter.
The dilemma. Optimal expectation level depends on $\eta_i$. High-$\eta$ dimensions reward
bold expectations because the action gain dominates the gap cost. Low-$\eta$ dimensions reward
calibrated expectations because there is no compensating action benefit. This kills two pieces of generic
advice at once. "Have low expectations to be happy" optimizes for gap-cost while ignoring action-loss.
"Shoot for the stars" applies high expectations uniformly without checking whether the dimension actually
has the elasticity to convert ambition into outcome.
Component 13
Expectations Are Partially Controllable
The model treats $E^*$ as a choice variable. In practice, expectations are not free parameters —
they are inherited, biological, environmental, and only partially adjustable. The agent has limited but real
authority over $E^*$, and that authority varies by dimension and by individual.
What sets your expectations regardless of your wishes.
Peer group and social environment. Childhood imprinting (especially for $r$ and $p$).
Biological temperament. Cultural narrative absorption. These forces install your starting expectations and
drag you toward your environment's mean.
What lets you adjust expectations with effort.
Deliberate practice (Stoicism, reflection, meditation, gratitude work). Environmental choice
(who you live among, what you consume). Explicit goal-setting. Cognitive reframing in the moment.
With sustained effort over time, expectations on a given dimension do shift —
but not all the way, and not instantly.
The exploit and its cost. If $E^*$ were fully controllable, there is a degenerate solution:
set $E_t^i$ slightly below $i_t$ on every dimension, harvest positive joy automatically, never try hard at anything.
This is real. People do it on purpose. Monks. Secular contented underachievers. The cost is twofold:
it kills elasticity-driven action on every dimension simultaneously, and it requires constant effort to hold
expectations against environmental drag. The exploit exists. It is not free.
Implication for the search. $E^*$ in the master objective should be read as
the expectation profile achievable given your starting point and willingness to do the work,
not any expectation profile you imagine. The feasible set of expectation profiles is bounded.
The search happens inside those bounds.
What This Tells You
01
Joy is absolute value plus happiness, scaled by how sensitive you are to the gap.
$J = A + \beta H$. A healthy, well-connected multi-millionaire who just moved in next to Bezos
has high $A$ but negative $H_t$ — his new peer group rewrote his expectations overnight.
A factory foreman who has spent twenty years exceeding his own expectations has modest $A$
but strongly positive $H_t$. With high enough $\beta$, the foreman's joy is higher.
Both absolute conditions and the expectation gap matter. $\beta$ determines which dominates.
02
Happiness is what you feel. Joy is what you have.
Like temperature — you do not feel absolute degrees, you feel hot or cold relative to
your internal baseline. $H_t$ is that baseline gap. You can have low joy and high happiness
(the cancer recovery — conditions are bad but improving fast), or high joy and low happiness
(the depressed millionaire — conditions are great but expectations keep outrunning them).
Maximize joy. Monitor happiness as feedback. A sustained negative $H_t$ is both a signal
and an engine — it drives the policy changes that raise $i_t$.
03
Survival is endogenous.
$\omega_t$ depends on $h_t$, $r_t$, and $p_t$. The lonely die younger.
The purposeless die younger. The unhealthy die younger. Neglect compounds twice:
once in joy, once in survival.
04
$r_t$ is the densest hub.
Three strong outgoing arrows — to wealth, purpose, and survival.
Standard cultural scripts treat relationships as a consumption good — something to enjoy
once other things are sorted. The system says they are a production good.
05
The policy is the lever, not the decision.
$X_t^\pi$ is well-defined — it is what the policy produces at time $t$ given the current state.
But optimizing over individual decisions is myopic: it ignores how today's decision shapes
$S_{t+1}$, $S_{t+2}$, and every state that follows. The policy search reasons over the whole
trajectory. You design the policy. The policy produces the decisions.
06
Expectations are not free. They cost you in disappointment when reality misses,
but they pay you in action when reality is shaped by belief. The optimal expectation level
is dimension-specific, governed by $\eta_i$. Set bold expectations on $w$ and $p$ where elasticity
is high. Set calibrated expectations on $h$ where elasticity has a ceiling. Set context-dependent
expectations on $r$ where the elasticity sign depends on you. The blanket prescriptions
— "manage expectations" or "dream big" — both fail because they treat a per-dimension question
as a global one.
07
The choice of $C$ is upstream of everything. The framework optimizes given $C$.
It does not pick $C$ for you. Most people inherit a $C$ from their culture and run optimization
beautifully inside an objective they never selected. That is how someone wins at life by every
conventional measure and feels hollow. Choose your $C$ deliberately. It is the metaproblem.
08
You can adjust your expectations, but not arbitrarily. The Stoic exploit — "set $E$ low,
get free joy" — is real but bounded. Expectations are inherited from environment and biology, adjustable
with sustained effort, dragged toward your peer group's mean. The exploit costs willpower to maintain
and kills elasticity-driven action across every dimension at once. Use the lever. Know its limits.
09
You get one draw, so decide which of your possible lives you are building for.
Expected value is the right objective only when you play enough times to average out, and it is what
you get by counting every possible life equally. You do not live every possible life. Sort them worst
to best and the weighting $\varphi$ says which ones count — flat is the casino, weight on the bottom is
prudence, weight on the top is hope, and hope costs you an axiom you should know you are spending.
Pick the shape deliberately, scale it to what a bad draw would actually cost you, and never let an
unrepeatable life be planned as if it were repeatable.