Decisions under Uncertainty: Utility and Information
Why a bet with positive expected monetary value can be rational to refuse, what the curvature of a utility function measures, and how to price an observation before buying it - including the common case where the honest price is zero.
Prerequisites: Bayesian Networks and Probabilistic Inference
Inference produces beliefs. Beliefs do not tell you what to do. Acting needs something beliefs cannot supply - preferences - and decision theory is the arithmetic that puts the two together.
A. Expected utility
A utility function gives one number for how desirable a state is. Since an action in an uncertain world does not fix its outcome, write for the random variable of possible outcomes. The expected utility of an action, given evidence , is
and the maximum expected utility principle is simply .
That is not "maximise the best case", which ignores risk, and not "reach a goal", since outcome quality here is continuous rather than pass-or-fail.
B. Money is not utility
You have won a game show. Take $1,000,000, or flip a fair coin for nothing against $2,500,000. The gamble's expected monetary value is $1,250,000, which beats the sure million. Most people decline.
Rate your current position at 5, the sure million at 8, and the top outcome at 9. Then
so declining is what the principle prescribes. The averaging is identical in both calculations; only the scale differs, and utility is the scale that encodes preference.
The disagreement is about the shape of a curve, not about caution. A billionaire's utility is nearly linear over a few million, giving , so the same principle says take the gamble. Two agents, identical odds, opposite rational choices.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
The argument is about the shape of a curve, so here is the curve. The three outcomes are points on it, the slider bends it, and the odds never move. The dashed line is the rating the top outcome would need for accepting to be right: 11, at the ratings above. That is worth reading twice - it says the second million and a half would have to add at least as much as the first million did, which is the exact opposite of the diminishing returns the whole argument rests on.
Interactive: the coin, and the curve that decides it
The odds never change. Only the shape of the utility does.
- Expected utility, accept
- 7.00
- Expected utility, decline
- 8.00
- Top rating that would tie
- 11.00
- What MEU prescribes
- decline
- Expected money
- $1,250,000
The gamble is worth $1,250,000 in money against a sure $1,000,000, so money says accept. In utility it is worth 7.00 against 8.00, so the principle says decline. Nothing here is a criticism of averaging - the average is computed the same way both times, and what differs is the scale being averaged. For accepting to be right the top outcome would have to be rated 11.00, which is to say the second million and a half would have to add at least as much as the first million did. At the lesson’s rating it adds 0.33 times as much, which is what diminishing returns MEANS, and why declining is not merely understandable but prescribed.
C. Curvature is risk aversion
Grayson found the utility of money close to logarithmic, an idea going back to Bernoulli in 1738. For one subject the fit was . Step through it in equal increments of $100,000 and the utility added shrinks: , then , then , then .
Shrinking gains is concavity, and for a concave curve for every lottery : facing the gamble is worth less than being handed its expected value. That is risk aversion, and it is measurable. On the same curve, a fair coin between $0 and $800,000 has expected utility against for the certain $400,000. The sure amount he values equally to the gamble - the certainty equivalent - is about $227,000, so the risk premium is about $173,000.
The curve is not concave everywhere. Someone already $10 million in debt might rationally accept a coin for a further $10 million gain against a $20 million loss, since the outcomes barely differ from where they stand. That gives an S-shape: risk-seeking when desperate, risk-averse over positive wealth.
Utilities have no absolute scale, so fix one by pinning a best and worst outcome at 1 and 0, then elicit any prize by offering it against a standard lottery and adjusting the probability until the agent is indifferent. That probability is the utility.
Where the worst outcome is death, people resist pricing it, and the trade-off happens regardless. A micromort is a one-in-a-million chance of death. Driving 230 miles costs one, so a car's 92,000-mile life costs 400; people pay about $10,000 for a car halving that risk, saving 200, which implies $50 per micromort - matching what studies report directly. The scope is small risks only; nobody accepts $50 million to die outright.
Russell and Norvig describe an agency that rejected an asbestos study for assuming a dollar value for a child's life, then declined the removal - thereby implying a lower value than the one it refused to state.
D. Pricing an observation
A decision network extends a Bayesian network with rectangular decision nodes for the agent's choices and a diamond utility node scoring the outcome. Evaluating it is a loop: set the evidence, then for each action run ordinary inference over the utility node's parents, average, and take the best. The inference engine is unchanged.
That structure lets you ask what an observation is worth before buying it. The value of perfect information is the expected utility of deciding after the observation, averaged over what it might say, minus the expected utility of deciding now:
An oil company can buy one of indistinguishable blocks; exactly one holds oil worth and each costs , so expected profit is zero either way. A definitive survey of block 3 changes that. With probability it finds oil and the company profits ; otherwise the odds among the rest improve from to , worth . Together:
The survey of one block is worth exactly the price of a block, for every .
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
E. The case where the price is zero
The same reasoning has a sharper consequence. Two routes: a highway worth expected utility 10, and a dirt track worth 4 that a satellite report could move anywhere in . The report is worthless. Whatever it says the track tops out at 6, you still take the highway, and expected utility afterwards is still 10, so .
Make the routes close instead - 10 against something in - and over readings you take the better of 10 and the reading each time:
Then mind the baseline. Deciding now is not taking the highway: it is taking whichever action has the higher expected utility, and the track is already worth . So , and all of it comes from a single reading. Without the report you drive the track; a reading of sends you to the highway instead, which is worth exactly . At the two routes tie, so switching gains nothing, and readings of and are good news and worth nothing, because you were taking that route anyway.
Where this leaves you
Information is valuable exactly when it can change what you do. A measurement that cannot alter the choice is worth nothing however precise or expensive, and value peaks when the candidate actions are close and the observation is wide enough to reorder them. That is a cheap thing to check before commissioning a study, and it is the same arithmetic that explains why a sensible person turns down a favourable bet. The training path Making Decisions under Uncertainty works each of these by hand and in code.
References & further reading
- Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.