Risk Management · Capital Markets

Measuring Risk Badly: Volatility, Value at Risk, and What Comes Next

Julian Gretzinger  ·  August 2, 2026  ·  Substack

Abstract

Financial risk management rests on two dominant metrics — standard deviation as a proxy for volatility, and Value at Risk as the industry's preferred loss-threshold measure. Both are deeply embedded in regulatory frameworks, internal models, and investor communication. Neither was designed to capture what investors actually fear. This article examines the structural properties of each measure, the conditions under which they perform and those under which they fail, and why those failures tend to coincide with the moments risk measurement matters most.

Volatility, defined as the annualised standard deviation of returns, is symmetric by construction. It penalises upside deviations identically to downside ones and assumes approximately normal return distributions — an assumption the empirical record has refuted repeatedly, from 1987 through the 2008 crisis to March 2020. Taleb's distinction between Mediocristan and Extremistan provides the sharpest diagnosis: standard deviation is a Mediocristan tool applied to an Extremistan world, and the turkey problem explains precisely why calm historical windows produce dangerously low risk estimates.

Value at Risk converts distributional uncertainty into a single monetary threshold. Its appeal is administrative. Its weaknesses are structural: it says nothing about loss severity beyond the threshold, it is sensitive to estimation methodology, and — through the mechanism Taleb calls silent evidence and the ludic fallacy — it creates perverse incentives that cause portfolios to accumulate the very tail risk it cannot see. Basel III and its successors have moved toward Expected Shortfall in partial acknowledgement of these limits.

The final section surveys alternatives: Expected Shortfall, semi-variance and the Sortino ratio, drawdown-based metrics, the Omega ratio, and scenario analysis. Each corrects specific distortions introduced by its predecessors. Taleb's fourth quadrant frames the outer limit: in the domain of complex payoffs and fat-tailed distributions, no summary statistic provides reliable guidance, and the appropriate response is acknowledged ignorance rather than quantified false confidence.

The deeper problem is not which formula to choose, but whether the industry's attachment to tractable summary statistics is itself a form of risk.

#finance#markets#riskmanagement#regulation#portfoliomanagement

I — The Measurement Problem

Risk, in its colloquial sense, is the possibility of an outcome worse than the one expected. In finance, that intuition has been operationalised in a remarkably narrow way. Two metrics — return volatility and Value at Risk — account for the overwhelming majority of how risk is estimated, communicated, and regulated across asset management, banking, and insurance. Their dominance is not the result of proven superiority over alternatives. It is the result of mathematical convenience, regulatory codification, and the institutional inertia that follows from both.

The choice of metric is not neutral. The measure chosen to represent risk shapes portfolio construction, determines capital requirements, influences manager compensation, and sets the terms on which investors evaluate performance. A metric that systematically misrepresents the distribution of potential outcomes does not merely produce inaccurate estimates — it produces distorted incentives, misdirected capital, and a false sense of security that breaks down precisely when the underlying stress is most acute.

Taleb's distinction between Mediocristan and Extremistan is the appropriate starting point (Taleb, 2007). In Mediocristan — the domain of height, weight, and caloric intake — no single observation can materially distort the population average. In Extremistan — the domain of wealth, book sales, and financial returns — one observation can dominate the entire sample. A single day of trading in October 1987 produced losses exceeding the cumulative daily losses of the preceding decade. Standard deviation and VaR are Mediocristan tools applied to an Extremistan world. The distortion that follows is not incidental; it is structural.

II — Volatility as Risk Measure

Construction and appeal

Volatility is almost universally operationalised as the annualised standard deviation of periodic returns. Its theoretical foundation lies in Modern Portfolio Theory (Markowitz, 1952), where the mean-variance framework treats expected return and variance as the sufficient statistics for evaluating any portfolio. Under this framework, rational investors select portfolios that maximise return for a given level of variance — the efficient frontier. Volatility is the square root of variance, which makes it unit-compatible with return and therefore directly usable in portfolio optimisation.

The measure has genuine virtues. It is mathematically tractable, decomposes cleanly across assets given a covariance matrix, and allows the construction of the capital market line, the Sharpe ratio, and the apparatus of factor-based risk attribution. An asset with annualised volatility of 8% can be compared meaningfully to one at 18% without additional distributional assumptions.

Structural deficiencies

The core problem with standard deviation as a risk proxy is symmetry. The measure treats positive and negative deviations from the mean identically. A portfolio that earns 5% above expectation in one period and 5% below in another receives the same volatility score as one that earns 10% above and 10% below — but investor exposure to loss is identical in both cases only if utility is symmetric around expected outcomes. Most investors, and certainly most institutional mandates, are not indifferent between upside surprise and downside loss. Standard deviation cannot distinguish between them.

A second deficiency concerns distributional assumptions. Mean-variance optimisation is optimal for investors with quadratic utility, or for portfolios whose returns are normally distributed. Empirical return distributions are neither. They exhibit negative skewness and excess kurtosis — the fat tails representing a concentration of extreme events well beyond what a normal distribution would predict. The historical frequency of daily equity moves greater than three standard deviations is several times the normal distribution's implication. Volatility calibrated to normality will systematically underestimate the probability and severity of extreme losses.

Volatility is what you can measure. Risk is what can hurt you. The two overlap, but they are not the same thing.

Third, volatility is backward-looking and regime-sensitive. Estimates based on historical windows will be low precisely when markets have been calm — often just before stress events — and high after crises have already materialised. Taleb's turkey problem captures this precisely: a turkey fed daily for a thousand days accumulates overwhelming statistical evidence that its situation is safe, right up to the day before Thanksgiving (Taleb, 2007). The VIX index exhibits exactly this pattern — lowest just before volatility spikes, highest at the point of maximum fear. The absence of a black swan from the sample is not evidence of its absence from the world.

Fourth, volatility aggregation across portfolios requires a full covariance matrix — and covariances, like volatilities, are unstable. The 2008 financial crisis demonstrated that correlations across asset classes converge toward 1 under systemic stress: equities, credit, commodities, and real estate fell together precisely when diversification was most needed (Embrechts et al., 2002). The mean-variance framework, which assumes stable covariance relationships, is most misleading in exactly the conditions it was supposed to model.

The symmetry problem illustrated. Consider two portfolios over five sub-periods. Portfolio A returns +12%, −2%, +8%, −14%, +6%. Portfolio B returns +18%, −8%, +4%, −8%, −6%. Both may exhibit similar standard deviations, but Portfolio B's larger downside draws are qualitatively different if the investor has liquidity constraints or a defined liability. Standard deviation cannot see this. A downside-only measure can.

Where volatility remains appropriate

None of this renders volatility useless. For liquid, approximately symmetric return distributions — short-term government bond portfolios, diversified equity indices over long horizons — standard deviation provides a workable first-order approximation. Its decomposability makes it valuable for attribution. The Sharpe ratio remains a useful comparative tool provided one is comparing assets with similar distributional characteristics. The problem arises when the measure is extended beyond its domain of validity: to hedge funds with option-like payoffs, to credit portfolios with embedded convexity, or to any context where the tails of the distribution are the primary source of economic risk.

III — Value at Risk

Design and regulatory embedding

Value at Risk was popularised by J.P. Morgan's RiskMetrics framework in the early 1990s (J.P. Morgan, 1996) and rapidly absorbed into financial regulation. The Basel I Market Risk Amendment (1996) permitted banks to use internal VaR models to calculate market risk capital requirements — a decision that entrenched the measure in global banking practice for three decades. Under Basel II and III, VaR at the 99% confidence level over a ten-day horizon became the standard input into market risk capital calculations, supplemented from Basel 2.5 onward by stressed VaR, and progressively replaced by Expected Shortfall under the Fundamental Review of the Trading Book (BCBS, 2016).

The measure is defined as the maximum loss not exceeded at a specified confidence level over a defined holding period. A one-day 99% VaR of $10 million means there is a 1% probability of losing more than $10 million in a single day. It tells the risk manager where the tail begins, not how deep it runs.

VaR's institutional appeal is legible. It translates distributional mathematics into a single currency figure that non-technical stakeholders — boards, regulators, investors — can interpret without statistical training. It is aggregable across desks and asset classes, and it provides a common language across firms and a mechanism for regulatory capital calibration. These administrative properties, more than any statistical virtue, explain its dominance.

Structural deficiencies

VaR's most fundamental deficiency was formalised by Artzner, Delbaen, Eber, and Heath in their axiomatisation of coherent risk measures (Artzner et al., 1999). A coherent measure must satisfy monotonicity, subadditivity, positive homogeneity, and translation invariance. VaR violates subadditivity: there exist portfolios where the VaR of the combined position exceeds the sum of individual VaRs. This means VaR can penalise diversification — an outcome that contradicts the basic logic of portfolio construction.

More practically, VaR is silent about the distribution of losses beyond its threshold. A portfolio with a 99% VaR of $10 million may face expected tail losses of $12 million or $120 million — the measure does not distinguish between these. In credit portfolios, structured products, and any strategy with negative skewness and excess kurtosis, the conditional distribution beyond the threshold can be catastrophically heavy. Lehman's internal risk models reported VaR figures that were, by their own model's standard, well within manageable bounds.

VaR tells you that the storm is unlikely. It says nothing about the size of the waves once it arrives.

Second, VaR is acutely sensitive to estimation methodology. The three principal approaches — historical simulation, variance-covariance, and Monte Carlo — produce materially different results for the same portfolio. Historical simulation is bounded by whatever stress events appear in the lookback window; if the window excludes 2008, the resulting VaR will not reflect credit crisis dynamics. Parametric VaR inherits the normality assumption. Monte Carlo VaR depends entirely on the distributional assumptions embedded in the simulation engine. The regulator receives a single number; the assumptions underneath it are invisible.

Third — and arguably most consequential for systemic risk — VaR creates perverse incentives at the portfolio level. Because it is a threshold measure, it can be gamed through strategies that shift loss probability from inside the threshold to outside it. A short put strategy collects premium in normal conditions and shows a favourable VaR profile until the underlying falls sharply. Taleb identifies this as the problem of silent evidence: the strategies that blow up are not in the sample used to calibrate the model, because they have not blown up yet (Taleb, 2007). The model sees only the survivors. The graveyard is invisible.

This connects to what Taleb calls the ludic fallacy — the error of confusing the controlled randomness of a casino, where outcomes are bounded and probabilities known, with the open-ended randomness of financial markets, where the rules themselves can change (Taleb, 2007). A VaR model assumes the game being played is stable. Financial crises are events that violate this assumption. They are not draws from a known distribution; they are changes in the distribution itself.

VaR and the 2008 crisis. The Basel Committee's own post-crisis analysis found that bank VaR models failed to predict the magnitude of losses incurred during the 2008 financial crisis. Actual trading losses exceeded banks' 99% one-day VaR on significantly more days than the 1% the model implied — in some cases by a factor of three or four. The subsequent Basel 2.5 and Basel III reforms (stressed VaR, incremental risk charge, the Fundamental Review of the Trading Book) were direct regulatory acknowledgements of VaR's empirical failure under stress (BCBS, 2009).

Where VaR retains utility

As an internal management tool, VaR is not without value — provided its limitations are explicitly acknowledged. For liquid, approximately normally distributed portfolios over short horizons, it provides a reasonable daily risk estimate. As a relative measure — tracking whether a desk's risk profile is growing or shrinking, or comparing concentration across business lines — it is operationally useful even if the absolute figure is imprecise. The problem is not VaR as a component of a broader risk toolkit. The problem is VaR as the primary, often exclusive, external risk disclosure metric, embedded in regulatory capital calculations in ways that obscure the assumptions on which it rests.

IV — Alternatives and Complements

Expected Shortfall

Expected Shortfall (ES) — also called Conditional Value at Risk or Expected Tail Loss — addresses VaR's principal structural deficiency by measuring the expected loss conditional on the loss exceeding the VaR threshold. Where VaR reports the boundary of the tail, ES reports the average severity of outcomes within it. A 97.5% ES figure represents the expected loss given that one is already in the worst 2.5% of outcomes — a materially more informative quantity than the threshold alone.

ES is coherent in the Artzner et al. sense: it satisfies subadditivity, meaning the ES of a combined portfolio cannot exceed the sum of individual ESs, preserving the mathematical logic of diversification (Rockafellar & Uryasev, 2000). The Basel Committee adopted ES at the 97.5% confidence level as the primary market risk metric under the Fundamental Review of the Trading Book, replacing 99% VaR — the most significant revision to bank market risk measurement since 1996 (BCBS, 2016).

ES is not without drawbacks. It is harder to backtest than VaR because violations are, by construction, rare and the sample of tail observations is small (Acerbi & Szekely, 2014). It also inherits distributional assumptions: if the underlying model misspecifies the tail, the ES estimate will misspecify it further, since it is an average across the tail rather than a point on it.

Semi-variance and the Sortino ratio

Semi-variance — the variance of returns below a target or minimum acceptable return — directly addresses the symmetry problem in standard deviation. By restricting the calculation to the downside distribution, it produces a measure aligned with the investor's actual objective function: avoiding outcomes below a threshold, not minimising dispersion in both directions. The Sortino ratio divides excess return by downside deviation (the square root of semi-variance), producing a risk-adjusted performance measure that does not penalise upside volatility (Sortino & van der Meer, 1991).

For investors with asymmetric mandates — pension funds managing against a liability benchmark, endowments with minimum spending requirements, or any strategy where capital preservation is primary — semi-variance is structurally more appropriate than standard deviation. In practice it is underused, partly because mean-variance optimisation has a fifty-year head start and partly because semi-variance does not decompose as cleanly, making portfolio optimisation more demanding.

Drawdown-based metrics

Maximum Drawdown measures the largest peak-to-trough decline in portfolio value over a given period. Calmar and Sterling ratios relate return to maximum drawdown as an alternative to the Sharpe ratio's use of volatility. These measures have direct economic relevance: for a leveraged investor, a manager operating under a high-water mark, or any portfolio subject to forced liquidation risk, the path of returns matters as much as their distribution. A portfolio with low standard deviation but a single large drawdown may be more dangerous in practice than one with higher measured volatility and smaller peak-to-trough declines.

Drawdown metrics are, however, sample-dependent in a particularly acute way. Maximum drawdown increases mechanically with the length of the observation period and is not forward-looking in any straightforward sense: past maximum drawdown provides a lower bound on potential future drawdown, not an estimate of the expected one.

The Omega ratio

The Omega ratio is defined as the ratio of probability-weighted gains above a threshold return to probability-weighted losses below it (Shadwick & Keating, 2002). Unlike the Sharpe or Sortino ratios, it makes no distributional assumptions and incorporates all moments of the return distribution by integrating across its full range. A ratio above 1 indicates that probability-weighted gains exceed probability-weighted losses for the chosen threshold; below 1, the reverse.

The Omega ratio is theoretically appealing precisely because it is distribution-agnostic. It can be applied to portfolios with option-like payoffs, hedge fund strategies with non-normal return profiles, or any context where standard deviation is a poor summary of the risk-return trade-off. Its practical limitation is interpretive: unlike the Sharpe ratio, it does not produce a single number independent of the return threshold chosen, which complicates comparisons across strategies and periods unless the threshold is standardised.

Stress testing and scenario analysis

No summary statistic, however well-constructed, captures the full dimensionality of risk under extreme conditions. Stress testing — the simulation of portfolio losses under specified adverse scenarios — complements quantitative risk measures by asking directly: what happens if a 2008-style credit crisis, the 1987 crash, or a correlated equity-rate sell-off were to recur? Regulatory stress testing has become a primary supervisory tool precisely because model-based VaR and ES calculations failed to anticipate losses generated under genuine systemic stress.

Scenario analysis shifts the question from "what is the probability distribution of losses?" to "what would this specific set of conditions produce?" That is a different question, and in some respects a more honest one for portfolios with significant tail exposure. The limitation is that scenarios are necessarily backward-looking: they model the crises already experienced, not the ones that have not yet occurred.

V — A Framework for Honest Risk Reporting

The practical implication of this survey is not that volatility and VaR should be abandoned. Both have legitimate uses within a broader measurement framework, and abandoning them would create comparability problems without necessarily improving the quality of risk information. The implication is that they should be presented honestly — with explicit acknowledgement of the distributional assumptions embedded in each estimate, the conditions under which those assumptions are most likely to fail, and a complement of tail-sensitive measures that provide information about the distribution beyond the threshold.

A defensible minimum standard for institutional risk reporting might include: annualised return and volatility (for comparability), Sortino ratio (to separate downside from total dispersion), Expected Shortfall at 97.5% (for tail severity), maximum drawdown over the period (for path-dependence), and at least one adverse scenario analysis (for non-parametric stress exposure). This is not an exhaustive framework, but it is a materially more honest representation than a single VaR figure and a Sharpe ratio.

The question is not which formula to choose, but whether the industry's attachment to tractable summary statistics is itself a form of risk.

The deeper tension runs through the entire history of quantitative risk management: the measures that are most tractable — those that decompose cleanly, aggregate across portfolios, and slot into regulatory formulae — are often the ones furthest from the actual distribution of outcomes in stress conditions. The measures that best describe the tail are harder to estimate, harder to aggregate, and more resistant to the compact reporting that institutions and regulators prefer.

Taleb's fourth quadrant sharpens this into a practical prescription (Taleb, 2007; Taleb, 2012). He maps the space of decisions across two axes: payoff structure (simple versus complex) and distributional knowledge (thin-tailed versus fat-tailed). In the first quadrant — simple payoffs, thin tails — statistical models work well and VaR is a reasonable tool. In the fourth quadrant — complex, non-linear payoffs combined with fat-tailed, poorly characterised distributions — no model-based risk metric provides reliable guidance. This is precisely the domain of structured credit, leveraged derivatives, and concentrated single-name exposure. The appropriate response in the fourth quadrant is not a better formula. It is position limits, optionality, and an explicit acknowledgement that the loss distribution cannot be estimated with useful precision. The industry's instinct is to respond to fourth-quadrant problems with more sophisticated models. That instinct makes things worse — replacing acknowledged ignorance with quantified false confidence.

Closing Note

The history of financial risk management is, in large part, a history of elegant models applied just beyond their domain of validity. Volatility and VaR are not wrong tools. They are tools used too far from the conditions that make them right. The corrections surveyed here — Expected Shortfall, semi-variance, drawdown metrics, the Omega ratio, scenario analysis — do not solve the fundamental problem. They push it back. In the fourth quadrant, where the losses that actually matter are concentrated, the honest answer may simply be: we do not know. That answer is uncomfortable. But it is more useful than a number that implies otherwise.


The views expressed are the analytical position of the author in a personal capacity and do not constitute investment, legal, or risk-management advice.

Sources

  • 1. Markowitz, H. (1952). Portfolio Selection. Journal of Finance, 7(1), 77–91.
  • 2. Artzner, P., Delbaen, F., Eber, J.-M., & Heath, D. (1999). Coherent Measures of Risk. Mathematical Finance, 9(3), 203–228.
  • 3. J.P. Morgan / Reuters (1996). RiskMetrics Technical Document, 4th ed. J.P. Morgan.
  • 4. Sortino, F., & van der Meer, R. (1991). Downside Risk. Journal of Portfolio Management, 17(4), 27–31.
  • 5. Shadwick, W., & Keating, C. (2002). A Universal Performance Measure. Journal of Performance Measurement, 6(3), 59–84.
  • 6. Taleb, N. N. (2001). Fooled by Randomness: The Hidden Role of Chance in Life and in the Markets. Texere.
  • 7. Taleb, N. N. (2007). The Black Swan: The Impact of the Highly Improbable. Random House.
  • 8. Taleb, N. N. (2012). Antifragile: Things That Gain from Disorder. Random House.
  • 9. Acerbi, C., & Szekely, B. (2014). Backtesting Expected Shortfall. Risk, December 2014.
  • 10. Basel Committee on Banking Supervision (2016). Minimum Capital Requirements for Market Risk (FRTB). Bank for International Settlements.
  • 11. Basel Committee on Banking Supervision (2009). Revisions to the Basel II Market Risk Framework. Bank for International Settlements.
  • 12. Rockafellar, R. T., & Uryasev, S. (2000). Optimization of Conditional Value-at-Risk. Journal of Risk, 2(3), 21–41.
  • 13. Danielsson, J. (2011). Financial Risk Forecasting. Wiley.
  • 14. Embrechts, P., McNeil, A., & Straumann, D. (2002). Correlation and Dependence in Risk Management. In Risk Management: Value at Risk and Beyond. Cambridge University Press.

Julian Gretzinger

Investor and writer on monetary history, real wealth mechanics, and financial markets. substack.com/@juliangretzinger