Value at Risk vs. Stress TestingHard

One asks how bad an ordinary bad day is and answers with a probability. The other asks what a specific catastrophe would cost and refuses to say how likely it is. Neither is the other's approximation.

5 min read · 897 words

Two questions that sound like one

  • Value at risk asks a question about the distribution. Over one day, at 99% confidence, what loss will not be exceeded? The answer is a number with a probability attached, estimated from history or from a model of it.
  • A stress test asks a question about a case. If rates rose 200 basis points while spreads widened 150 and equities fell a fifth, what would we lose? The answer is a number with no probability attached, and that absence is deliberate.
  • They are not two precisions of the same measurement. One describes the middle of the distribution with statistical machinery; the other describes one specific corner of the tail with a story. Risk measures covers the first, stress testing the second.

What each one is silent about

  • Value at risk says nothing about the size of the losses beyond it. A 99% one-day figure is the threshold of the worst one day in a hundred; on the day it is exceeded it makes no claim at all about by how much. That is not a flaw to be corrected but the definition, and it is why expected shortfall — the average loss given that the threshold is breached — exists beside it.
  • A stress test says nothing about likelihood. Attaching a probability to a scenario that has never occurred would be the least defensible number in the exercise, so none is attached — which means a stress loss can never be compared with a value-at-risk figure as though they were the same kind of quantity.
  • The most common error is exactly that comparison: reading a scenario loss as "our 99.9% figure". It is not a percentile of anything. It is one path, priced.

Where the arithmetic actually differs

$$ \text{scenario} = L_a + L_b \qquad\quad \text{aggregated} = \sqrt{L_a^2 + L_b^2 + 2\rho L_a L_b} $$
What the symbols mean
  • Lleverage, or a loss given default
  • rhocorrelation between two things
  • A scenario adds. It says both things happen, so the losses sum, and there is no diversification because none was assumed.
  • A distribution-based measure asks how often they happen together, and below a correlation of one that returns a smaller number for the same two positions. The gap between the two figures is the correlation assumption, priced.
  • Which is why a scenario loss is normally the larger number, and why comparing them tells you something useful: the difference is exactly what the firm is assuming about co-movement, made visible. The calculator on the stress-testing page moves both.
  • And correlation is the parameter least stable under stress, estimated from calm periods and drifting towards one when it matters — see diversification.

How each fails

  • Value at risk fails by being calibrated on a sample that does not contain the event. A model fitted to two quiet years has no view on a shock it never saw, and it says so nowhere in its output. Backtesting — counting how often the threshold was actually breached — is the honest check, and a model breaching more often than its confidence level allows is telling you something.
  • A stress test fails by testing what somebody thought of. The search is bounded by memory and by what is defensible in a meeting, so the scenario library systematically excludes the scenario nobody imagined. Reverse stress testing — start from failure, work backwards — exists precisely to search where the library does not.
  • Both fail together on the second round. Forced selling, funding withdrawal, liquidity evaporating in the same instant: a naive version of either revalues today's positions under new inputs and stops, and the losses that turn a bad quarter into a failure are in what happens next.
  • Every case study on this site is one of these two failuresLTCM the first, Archegos and LDI the second.

The comparison

Value at riskStress test
The questionHow bad is an ordinary bad day?What would this specific event cost?
Probability attachedYes — the confidence level is the pointNone, deliberately
Where the numbers come fromHistory, or a model calibrated to itA scenario somebody constructed
AggregationCorrelation does the workLosses add
FrequencyDaily, mechanicallyPeriodic, and slower to build
Checkable against outcomesYes — backtest the exceptionsNo: the event has not happened
Blind toAnything outside the sample; the size of the tailThe scenario nobody wrote down
What it is used forLimits, capital, daily monitoringCapital guidance, planning, board discussion

Why a firm needs both, said precisely

  • They are complements rather than a primary and a sanity check. One is continuous, comparable across desks and testable against what actually happened; the other reaches places the first cannot see and cannot be tested at all.
  • A firm with only value at risk has a number every day and no view on the event that would end it.
  • A firm with only stress tests has several vivid stories and no way to set a limit, aggregate across desks, or discover that its model has been wrong for six months.
  • And both should be quoted with their conditions attached. A stress number stripped of its assumptions becomes a forecast; a value-at-risk figure stripped of its confidence level and horizon means nothing at all.

Information and education only. This compares two risk measures in general terms. It is not advice, not a recommendation of either, and nothing here takes account of your circumstances.

Information and education only. Every page, figure and calculator on this site exists to explain how financial instruments work. Nothing here is investment, tax or legal advice, a recommendation, or a valuation you can rely on. Full disclaimer