
Expected returns matter most over long horizons. Practitioners typically estimate them using either forward-looking forecasts or historical averages. Using 119 years of US data, research shows that an equilibrium model based on Tobin's Q and dividend yields forecasts long-run returns more accurately than historical averages. The model suggests that US equity returns between 2019 and 2040 could be around 1.5 percentage points lower annually than historical averages, with important implications for UK price controls.
In two ways: forward-looking predictions and outturn averages. Most practitioners take the first route, and the largest asset managers and advisers publish capital market expectations built from current information, among them JP Morgan, Schroders, Vanguard, Willis Towers Watson, Aberdeen Standard, SEB, UBS and Aon. Regulators and valuers have traditionally taken the second, leaning on outturn averages of the returns history has delivered: the Competition Commission took that approach in its 2014 Northern Ireland Electricity determination, and Ofgem in its 2019 RIIO-2 finance methodology. The paper's contribution is to put the two approaches into a fair head-to-head test over 119 years of data.
Table 1: The two ways practitioners estimate the expected market return, and what each gives up.
Outturn averages are verifiable but assume the future resembles the whole past; forward-looking estimates use current information but vary between producers. The paper's contribution is to test the two against each other rather than argue for one.
Two conclusions emerge from the literature, and both point the same way as this paper.
The first is that the expected market return is likely lower than historical averages suggest. The realised returns of the twentieth century were flattered by outcomes that cannot be projected forward: the era's valuation ratios climbed, delivering a one-off revaluation gain on top of the returns fundamentals generated, and the markets that survived to be measured were the fortunate ones. Forward-looking estimates built from dividends, earnings and growth consistently sit below the outturn average. Taking the historical average as the expectation therefore bakes past good fortune into the future. The evidence is longstanding: Mehra and Prescott found realised equity returns difficult to reconcile with standard models of risk, Campbell and Shiller found stock prices exhibiting excess volatility relative to the present value of dividends, and Blanchard found returns at times exceeding what reasonable levels of risk aversion can justify.
The second is that ex-post returns over long horizons are less uncertain than the uncertainty of one-year returns would imply. If annual returns were independent draws, risk would compound mechanically with time; instead, long-horizon outcomes cluster more tightly than that arithmetic predicts, which is the statistical signature of mean reversion and the reason patient investors have been rewarded. It is also the opening this paper exploits: returns that partially correct themselves are returns that can, at long horizons, be forecast. Here too the evidence is longstanding: Fama and French showed the explanatory power of dividend-yield regressions growing with the investment horizon, Dimson, Marsh and Staunton documented long-term consistency in world equity returns across markets, and Siegel argued that over long periods bonds carry risks of their own relative to equities.
The case against predictability is serious and rests on four criticisms. Econometrics first: the predictor ratios are highly persistent and the regressions use overlapping horizons, a combination that biases coefficients and overstates significance in small samples, as Boudoukh, Richardson and Whitelaw argued. Statistical fragility second: apparent predictability shifts with the sample period chosen, and much of it weakens or disappears when tested out-of-sample; Kim, Nelson and Startz held that mean reversion is entirely a pre-war phenomenon. Predictive power third: even where the relationship is real, it explains little of any single year's return, and beating the simple historical average in real time has proved genuinely hard; Ang and Bekaert found that once small-sample properties are accounted for, dividend-yield predictability is not statistically significant. Interpretation last: if a valuation ratio's own mean can shift, as the dividend yield's arguably did when buybacks changed payout policy, then reversion to the mean is not well defined, and what looks like predictability may be an artefact of measuring the wrong equilibrium.
Where does the literature actually agree? The disagreement is narrower than the volume of studies suggests. Few dispute that valuations and subsequent long-horizon returns are related in the data; the dispute is whether that relationship is estimated reliably enough to act on in real time. That framing explains this paper's design, which confronts each criticism directly: 119 years of data against small-sample bias, cointegration methods built for persistent series, a payout-adjusted dividend yield against the shifting-mean objection, and out-of-sample tests as the final arbiter.
No, and the paper does not need it to. Its position is that markets are very efficient in the short run, which is why returns over days or single years are close to unpredictable, but less efficient in the long run, where valuation ratios pull prices back toward equilibrium. On this view the expected market return, the EMR the model forecasts, is not a fixed number but varies through time with valuations. Long-horizon predictability then reflects a required return that moves, rather than an everyday inefficiency waiting to be exploited, so the findings sit alongside the Efficient Market Hypothesis instead of against it.
Because among valuation measures it has an economic anchor, not just a statistical one. Q compares the market's price for companies with what their productive capital would cost to replace, and that comparison polices itself: if the market values assets well above replacement cost, building new capital beats buying it, and the resulting investment pulls the ratio back down; well below, acquisition beats construction, and the ratio is pulled up. This arbitrage logic gives Q a behavioural reason to revert toward equilibrium, where earnings-based ratios offer only the historical observation that they usually have. Q also connects the stock market directly to the real capital stock, a relationship the paper's four-variable system of market value, adjusted dividends, capital and liabilities is designed to exploit.
Measuring Q precisely is another matter, and the difficulties are real: replacement cost rests on estimated depreciation, intangible assets are captured patchily, and the markets for much corporate capital are illiquid, so any measured level of Q carries error. The paper's methodological answer is the key decision to highlight: it works with deviations of Q from its own long-run average, not with its absolute level. Errors that persistently shift the level largely cancel when only deviations are used, so the signal that matters for forecasting, how far the market currently sits from its equilibrium, survives uncertainty about where exactly that equilibrium lies. This is why the analysis is expressed as log deviations from the mean (Figure 3). It also explains why different producers, such as Smithers & Co, report different absolute levels of Q without undermining the forecasting approach: the model depends on deviations from equilibrium rather than the precise level of the ratio.
One assumption underpins the entire methodology, and it belongs up front: equity prices cannot deviate indefinitely from the economic fundamentals captured by Tobin's Q and the adjusted dividend yield. Everything else in the framework is machinery for estimating how, and how quickly, prices return toward that equilibrium.
The first piece of machinery is leverage, and it is a genuine strength of the approach. Simpler valuation models assume companies are financed entirely by equity; this model brings net liabilities in as a variable in its own right, because real companies are leveraged and debt changes both the risk and the cash flows that equity holds a claim on. Doing so lets the framework represent the full chain of equity valuation rather than a single ratio: companies raise liabilities, invest them as capital, capital generates the profits that fund distributions, and expected returns price the equity claim on what remains. Modelling that chain makes the framework a more complete representation of how equity value is created, and it is why the later finding that prices respond positively to changes in liabilities is interpretable economics rather than a statistical curiosity.
The second is the dividend yield, whose rationale as a fundamental indicator deserves stating plainly: the yield measures the cash a shareholder's pound actually buys. When prices are high relative to the cash companies deliver, the yield is low and the market is expensive in the only currency that ultimately matters to an equity holder; when the yield is high, the market is cheap in the same terms. The adjusted version counts all distributions, dividends, buybacks and other cash flows to equity, so the indicator survives changes in how companies choose to pay. Between them, the two anchors tether the market to fundamentals from both sides: Tobin's Q from the cost of the assets, the adjusted yield from the cash the assets deliver. The machinery serves a single demonstration: that this equilibrium, not the historical average, is what expected returns ultimately obey.
Because it no longer measures what it once did. Dividends were the dominant way companies returned cash to shareholders for most of the twentieth century, but since the 1980s buybacks have taken an increasing share, so the traditional yield tracks a shrinking slice of the distributions actually reaching investors. The consequence is not a rounding error: the traditional yield appears to drift permanently downward and fails statistical tests for mean reversion, which would disqualify it as an anchor. That apparent drift has generated its own literature, including demographic explanations for a supposed structural change in valuations. This paper's answer is simpler. The paper's answer is simpler: count all distributions, not just dividends. Once buybacks, new issues, cash-financed acquisitions and private equity issuance are incorporated, the apparent downward drift disappears. The adjusted yield reverts to its mean and passes the statistical tests the traditional dividend yield fails. What looked like a permanent change in valuations was largely a change in how firms returned cash to shareholders.
A century-deep core with international cross-checks. The core dataset covers the US nonfinancial corporate sector from 1900 to 2018, assembled from Wright's long-run fundamentals, the Federal Reserve's FRED database and Shiller's market data. The paper supplements this with UK data and international evidence to test whether its findings hold beyond a single market.
The historical series combines Wright's long-run US fundamentals with the Federal Reserve's Financial Accounts and Shiller's market data. To test whether the findings extend beyond the United States, the paper also draws on the Dimson, Marsh and Staunton dataset covering 21 countries, together with UK data from the ONS national balance sheet, S&P Capital IQ and historical Financial Times series.
The supplementary datasets answer a specific objection rather than adding volume. A result from one market, however long the sample, invites the survivorship criticism from the literature: the US is the market that happened to win. The paper therefore checks its story against UK fundamentals and against the Dimson, Marsh and Staunton dataset of returns across 21 countries. The three sources complement one another by design: the US series provide depth, the only place all four fundamentals exist for 119 years; the 21-country returns provide breadth over more than a century, testing whether patient investors were rewarded everywhere or only in the winner; and the UK data provides both a second test bed and the bridge to the paper's regulatory application. No single source could do all three jobs.
Yes, and by more than chance alone can explain: The annualised uncertainty of investment returns falls relative to what an independent-return process would predict as the holding period lengthens. That is the key message to hold onto before any figure or ratio, because it is the empirical heart of the paper. If annual returns were independent, the dispersion of long-horizon outcomes would grow mechanically with time, in proportion to the square root of the horizon. The evidence shows dispersion growing more slowly than that, which means good and bad years partially offset one another, indicating mean reversion: returns are not entirely independent across time.
This is what the paper's opening figure is designed to demonstrate, and it should be read as a comparison, not a level [Figure 1]. It sets the dispersion of returns actually experienced over lengthening holding periods against the dispersion that independence would predict, and the gap between the two lines is the finding: the longer the horizon, the further actual uncertainty falls below the independence benchmark. A reader need not take anything from the figure beyond that widening gap; the gap is the empirical signature of patience being rewarded: investors who hold longer historically experienced less uncertainty than short-term volatility alone would suggest.
The variance ratios formalise the same comparison across 21 countries, and their practical significance matters more than the percentages themselves. A variance ratio below one says that a market's long-horizon returns are less risky, per year invested, than its own short-horizon volatility implies, so equities appear less risky for patient investors than short-term volatility statistics suggest, and horizon belongs alongside risk appetite in any allocation decision. The ratios sit below one at long horizons in most of the 21 countries in the Dimson, Marsh and Staunton data, and that consistency is itself evidence of the strongest kind: a pattern appearing across markets with different histories, institutions and fortunes cannot be a quirk of the US sample or a survivor's accident, which is precisely the objection the international data was assembled to meet. Patient investors have been rewarded nearly everywhere, not only in the market that won the century.
The numbers behind that comparison. In the US market a 15-year holding period has a 95% probability of producing a positive real return. Across the 21 markets, 19 show variance falling as the horizon lengthens, with the decline setting in after roughly nine years; only Spain and Japan show it rising. At 20 years the US variance ratio sits 22% below its one-year level and the UK's 31% below, with Denmark the extreme case at 31% of its one-year variance. Since the US accounts for 53.3% of world market capitalisation, followed by Japan at 8.4% and the UK at 5.5%, the US result carries most of the weight for a globally diversified investor.
Figure 1: Equity returns in the US, 1900 to 2018: actual outcomes against a simulation assuming independent annual returns, 90% confidence
The actual band sits inside the simulated one and the gap widens with the horizon. If each year's return were an independent draw, uncertainty would grow in proportion to the square root of time; it grows more slowly than that. This single comparison is the empirical claim the rest of the paper explains.
Figure 2: Variance ratios for 21 stock markets over 119 years: h-year variance relative to one-year variance
Nineteen of the 21 markets show variance falling once the holding period passes roughly nine years, the US to 78% of its one-year level and the UK to 69%. Only Spain and Japan rise. Patience is rewarded almost everywhere, which answers the survivorship objection that the result is an artefact of studying the market that happened to win.
Yes, but the two ratios tell very different stories, and the difference drives the paper's design. The central question is whether valuation measures behave like temporary deviations from fundamentals or whether they permanently drift over time. The principal finding comes from the formal tests, and it deserves stating first: the adjusted dividend yield is stationary, with non-stationarity rejected at the 1% level across almost all specifications and sub-samples, while the traditional dividend yield is not [Table 1]. The charts show why: the adjusted yield crosses its mean regularly, while the traditional yield has sat below its mean for most of the past 40 years [Figure 4]. That divergence between two measures of nominally the same thing is the section's most telling result, and its cause is simple. Since the 1980s companies have shifted payout from dividends toward buybacks, so the traditional yield tracks a shrinking slice of distributions while the adjusted yield counts them all (Bunn and Shiller, 2014). Alternative explanations for the traditional yield's drift attempt to explain a non-stationarity that disappears once the full distribution of shareholder payouts is measured. The practical consequence follows directly: the adjusted yield is the more stable valuation measure and the one the model uses for forecasting.
The confirmatory test points the same way. Under KPSS, where the null is stationarity rather than a unit root, the adjusted yield scores 0.102, comfortably inside the stationary range, while the traditional yield is rejected at the 1% level at 0.91 and Tobin's Q at 5% at 0.47. All three tests agree on the adjusted yield across every period, with one exception; no two agree on either of the others.
Tobin's Q tells the slower story, and read as evidence rather than description, its pattern says two things about how the equilibrium binds. Q is highly persistent, crossing its mean infrequently, with the crashes of 1929, 2001 and 2008 each helping to pull it back toward its long-term average [Figure 3]. The persistence means markets can sit away from fundamental value for years at a stretch, which is the prolonged-deviation worldview the model is built on and the reason level-based market timing fails. The corrections mean the restoring force is nonetheless real, arriving irregularly and often abruptly, which is why forecasting can work at long horizons even though it cannot at short ones. Q's persistence also explains why its own stationarity tests are less conclusive, particularly for samples running to 2018 after its prolonged departure since 1990: a slow tether is hard to prove statistically, which is exactly why the paper rests the formal system on the adjusted yield and on cointegration between the fundamentals, rather than on Q's stationarity alone.
Table 2: Unit root tests on the three value indicators. A rejection means the indicator reverts to its mean and can serve as an anchor; a failure to reject means it drifts, and cannot.
This is the case for measuring payouts comprehensively. The adjusted dividend yield rejects non-stationarity in five of six tests, three of them at the 1% level. The traditional yield fails in half of them, and Tobin's Q weakens as the sample extends to 2018 because of its prolonged departure from its average since 1990. The reader's takeaway is not that one indicator beats another but that the ordinary dividend yield has stopped measuring what it used to: once buybacks and other cash flows to equity are counted, the apparent drift disappears. A yield that drifts cannot anchor a forecast; one that reverts can.
Figure 3: Measured Tobin's Q, log deviation from its mean, 1900 to 2018
Q crosses its mean rarely and sits away from it for years at a time, with the crashes of 1929, 2001 and 2008 each pulling it back. Both halves of that pattern matter: the persistence is why level-based market timing fails and why Q's own stationarity tests are weak, while the corrections are why forecasting works at long horizons even though it cannot at short ones.
Figure 4: Dividend yield and adjusted dividend yield, log deviations from their respective means, 1900 to 2018
The adjusted yield crosses its mean throughout the century; the traditional yield has sat below its mean for most of the past 40 years. Two measures of nominally the same thing diverge because payout shifted from dividends to buybacks. This is the case for a comprehensive payout measure in one picture, and the reason the paper does not need a demographic explanation for a drift that turns out not to exist.
Because each covers the other's weakness, and only together do they pin down the equilibrium. Tobin's Q is economically anchored but statistically slow, so persistent that tests struggle to confirm its reversion; the adjusted dividend yield is statistically sharp, stationary at the 1% level, but speaks only to the cash side of valuation. Plotted together, the two tell one consistent story about when the market has been dear or cheap [Figure 5], and combining them means the valuation signal survives periods when either single ratio is misleading on its own. The paper's specific combination pairs Q with the inverse of the adjusted yield, a deliberate departure from Robertson and Wright (2002), which puts both indicators on the same footing: for each, high means expensive. This combined reading of value is not a presentational choice but the restriction the formal model tests, which is why it belongs among the paper's key results rather than its preliminaries.
Figure 5: Tobin's Q and the inverse adjusted dividend yield, relative to their mean values
Both ratios put market value in the numerator, so above zero means expensive on either measure. For most of the past two decades they have disagreed, Q calling the market dear while the inverse adjusted yield called it cheap. Neither indicator is reliable alone, which is precisely why the model imposes both at once.
By one binding equilibrium, and the section's verdict deserves stating up front: imposing the combined restriction, Tobin's Q together with the adjusted dividend yield, is strongly supported by the data, which is direct statistical backing for the paper's theoretical framework.
The system has four variables, each with a plain meaning. Stock market value (p) is what the market says companies are worth. Cash flows to equity (d) are what companies actually deliver to shareholders: dividends, buybacks and net issues. Capital at replacement value (k) is what the companies' productive assets would cost to rebuild. Net liabilities (l) are the debt standing between the firm's assets and its equity holders.
The two value indicators are built from exactly these four, Q comparing market value with capital, the adjusted yield comparing cash delivered with price, so if the indicators revert to equilibrium, the four underlying variables cannot drift apart indefinitely, even though each may wander substantially on its own. (Figure 6).
Lag-length tests select two lags on the Akaike criterion, so the four variables are estimated as a VAR(2). The Johansen test then finds exactly one cointegrating relationship in all three sample periods, with maximal eigenvalue statistics of 31.5, 34.87 and 30.24 against a 90% critical value of 26.1, and no evidence of a second, third or fourth. The paper proceeds to impose two relations rather than one, on the grounds that the test may lack power against Q's persistence and that theory predicts two, and it says so openly rather than resting on the test alone.
Table 3: Johansen test for how many cointegrating relationships bind the four-variable system.
One relationship is found in every sample period, and no second, third or fourth. That establishes that an equilibrium binds the system; it does not establish which one, which is what the next table is for.
Table 4: Testing whether the four-variable system is bound by an equilibrium, and if so by which one. A high p-value means the restriction is not rejected, so the theory survives the data.
Neither indicator holds on its own; together they do. Imposed alone, Tobin's Q is rejected in every specification and the adjusted yield in three of six. Imposed jointly, the restriction survives in all six, with p-values from 0.12 to 0.65. That is the paper's central econometric result and the reason the framework needs both: Q measures the market against the cost of the assets, the adjusted yield against the cash those assets deliver, and only the pair pins down where equilibrium lies.The test alone does not say what the relationship is, so the paper restricts the system to its theoretical priors, the combined Q and adjusted-yield equilibrium, and the data accept the restriction. That acceptance is the load-bearing result: the equilibrium the theory predicted is the one the century of data contains, and the vector error correction model used for forecasting is built on it. The point of the statistics is not the model itself but what it establishes: market prices and fundamentals remain tied together over time. That long-run equilibrium provides the foundation for the paper's argument that long-horizon returns are predictable.
Figure 6: Log levels of stock market value, adjusted dividends, capital and net liabilities, in constant prices
All four series wander individually and none is stationary, with the 1929, 2001 and 2008 crashes visible as dips in price. Cointegration is the claim that despite wandering, they never drift far apart. This chart is what the claim is made about.
The fundamentals come first. The causality results, reported in full in the paper, show past values of the fundamentals carrying predictive content for prices, which is the direction the framework requires: departures from equilibrium are corrected through subsequent price movements, not through the fundamentals bending to meet the market. The question matters because the equilibrium story depends on it; if prices led and fundamentals followed, the model would be reading the causation backwards.
In the price equation, a market sitting one unit above its Q equilibrium is followed by a fall of 0.29 the next year, and one unit above the inverse adjusted yield equilibrium by a fall of 0.27, both significant at the 1% level, with an R-squared of 15%. Both indicators Granger-cause price changes in all three sample periods at the 5% level or better. Two short-run effects run the other way and are worth noting: a rise in dividends is partly offset, by 0.12, in the following year's return, which is what dividend irrelevance predicts, while a rise in liabilities raises the next year's return by 0.20.
The price equation explains 15% of the variation in annual price changes, the dividend equation 22% and the capital equation 24%. Both value indicators Granger-cause price changes in every sample period, with chi-squared statistics of 3.93, 4.16 and 5.57 for the adjusted yield and 11.90, 5.49 and 6.09 for Tobin's Q, all significant at the 5% level or better. The liabilities equation is not reported at all, because none of its explanatory variables reaches significance, and that absence is itself a finding: liabilities appear to influence the system without themselves responding systematically to the other variables.
The test behind the result deserves its precise reading. Granger causality examines one thing and no more: whether past values of one series improve forecasts of another. It establishes precedence in time, not causation in the everyday sense, a distinction the paper is careful to keep.(Brooks, 2008)
The impulse-response analysis reveals two key findings. In the short run, each variable responds primarily to its own history. In the long run, the equilibrium's central channel, from adjusted dividends to stock prices, emerges once the cointegrating structure is imposed.
Own history dominates. In the unrestricted model, each variable is impacted primarily by its own lagged value (Figure 7). This matters because it is what short-run efficiency looks like in the data: little cross-prediction from one variable to another year to year, consistent with the paper's position that markets are very efficient at short horizons.
Capital is the slowest variable, taking six to seven years to absorb a shock, which fits the lead times on construction and investment programmes. Over a 25-year horizon the variance of prices falls by about 30% on the 1900 to 1990 sample and 34% on the full sample to 2018, and adjusted dividends account for almost all of that reduction, 25 and 33 percentage points respectively. That is the same predictability the variance ratios show, arriving by a different route.
The paper therefore re-runs the analysis under the alternative ordering, and the results are broadly similar, so the findings are robust to the assumption rather than an artefact of it. The numbers themselves: over 1900 to 1990, the unexplained variance of price falls by roughly 30% as the horizon extends to 25 years, with roughly 25% explained by adjusted dividends; over the full 1900 to 2018 sample the reduction is larger still, roughly 34%, with roughly 33% explained by adjusted dividends (Figures 9 and 10). Long-horizon returns are demonstrably less uncertain than short-horizon returns, in numerical form, and this feeds directly into the forecasting tests that follow.The contrast between the short run and the long run is the section's sharpest conclusion, and it deserves emphasis. Year to year, each variable answers mostly to its own past and payout decisions look like noise against price movements.
At long horizons, the equilibrium asserts itself and the response of price to adjusted dividends turns clearly positive: companies delivering more cash to shareholders are associated with higher subsequent value. That finding sits uncomfortably with the Dividend Irrelevancy Theorem, under which payout policy should not affect value at all; on this evidence, over the horizons where the equilibrium binds, it does. The paper supports the conclusion with Shiller's argument on the centrality of dividends to long-run value.
Figure 7: Accumulated response to one standard deviation shocks, unrestricted model, two standard errors, 1900 to 2018
In the unrestricted system each variable answers mainly to its own past, which is what short-run efficiency looks like in data and is consistent with returns being close to unpredictable year to year. The exception is price responding positively and significantly to liabilities, which is why the model treats leverage as a variable rather than assuming it away.
Figure 8: Accumulated response to one standard deviation shocks, equilibrium model, 1900 to 1990
Once the cointegrating structure is imposed, the response of price to adjusted dividends becomes visible and positive, as the theory predicts. The apparent fall in liabilities after a price shock is statistically insignificant; companies do not restructure their debt because the share price moved.
Figure 9: Variance decomposition of the equilibrium model, 1900 to 2018
As the horizon extends to 25 years, the variance of prices falls by about 34%, and adjusted dividends explain almost all of that reduction. Long-horizon returns are less uncertain than short-horizon ones, and the fundamentals are what does the explaining. This is the same finding as the variance ratios, reached by a different route, which is why the two together are more persuasive than either alone.
The answer is yes. The VECM outperforms the historical-average benchmark in all three in-sample test periods and in two of the three out-of-sample exercises, broadly consistent with Campbell and Thompson (2008), who likewise find that predictive models can outperform the historical average.
The exercise has one objective, tested two ways: whether the VECM forecasts returns more accurately than the historical average, judged first in-sample, using the full dataset, and then out-of-sample, rebuilding the model with only the information available at each forecast date, so that the two approaches compete on equal information. The test periods are chosen to make the race fair before they are chosen for anything else: each is long enough for the equilibrium forces to operate, they deliberately include a low-return period and a high-return period so that neither approach is flattered by the era, and enough data sits outside them to contrast the in-sample and out-of-sample results. The three periods are 1967 to 1986 (low return), 1987 to 2006 (high return) and 2007 to 2018 (most recent at the publication time).
Table 5: Forecast accuracy in-sample: the model against the historical average over three twenty-year windows.
The model is more accurate in all three windows, and by most in 1986 to 2006, where the historical average was five percentage points adrift. Mean absolute error falls from 3.16 percentage points to 1.55.
Table 6: Forecast accuracy out-of-sample: the model rebuilt using only data available at each forecast date.
This is the fair test, because the model and the historical average now have the same information. It wins in two of three windows, with mean absolute error of 1.38 against 3.16.Table: Out-of-sample VECM forecasts, rebuilt using only information available at each forecast date.The overall out-of-sample finding is the one to hold onto: even with the information advantage removed, the VECM beats the historical average in two of the three periods, and in the earliest it cuts a near three-point error to a third of a point. The exception, 2006 to 2018, has identifiable causes rather than mysterious ones: at 12 years the period is arguably too short to provide a fair test of forces that operate over decades, and the forecast made in 2005 is inflated by the boom of the 1990s dominating the data available at that date. Taken together, the forecasting exercises show that the equilibrium model retains predictive value even when deprived of future information. That is a much stronger test than fitting the historical sample, and it suggests the model captures relationships that extend beyond the data used to estimate it.
The practical implication is straightforward: expected future returns are lower than the historical average suggests. The forecasting results show that the VECM consistently outperforms the historical-average benchmark, supporting the idea that long-horizon expected returns should be estimated from current valuations rather than assumed to equal the past.
Applied forward, the model predicts a downward dip in stock prices and dividends, implying US equity returns from 2019 to 2040 around 1.5 percentage points below the historical average each year (Figure 13). For anyone using an expected return, an investor setting allocations, a valuer discounting cash flows, a regulator setting allowed returns, the practical content is the same: at 2019 valuations, the past overstates the future.
The source literature reinforces the point: long-run averages of realised returns are an unreliable guide to expected returns, and estimates should be informed by current valuations.
Figure 10: Forecast to 2040: the model implies an expected market return around 1.5 percentage points below the historical average
Applied forward from 2019 valuations, the model predicts a dip in prices and dividends. For an investor this is a forecast; for a regulator it is a warning that an allowed return anchored on the historical average would be set too high for the period it covers. This is the exhibit that connects 119 years of econometrics to a decision about customer bills.
Because the expected market return determines the allowed cost of equity. At each price control, regulators must estimate the return equity investors require, and the conventional anchor is the long-run historical average of market returns. This research supports a different weighting: if returns contain a predictable component tied to valuations, then at times of high valuations, as the model identified in 2019, the forward-looking expected return sits below the historical average, and an allowance anchored purely on history overstates the cost of equity that customers fund. The size of the effect is material at the scale of regulated revenues, since allowed returns apply to asset bases in the tens of billions of pounds.
The numbers a UK regulator would use. UK real total returns over 1899 to 2018 were 5.4% geometric on the DMS inflation series and 5.1% on the Bank of England's, or 5.0% in log terms. The uplift from log to arithmetic returns falls from 1.9% at a one-year horizon to 0.7% at twenty years once predictability is allowed for, against 1.3% on outturn variance alone. On that basis the preferred estimate of the expected market return is 5.7%, a more conservative reading gives 5.9%, and ignoring predictability altogether gives 6.3%. Recent UK regulatory assumptions have been 6.5%.
The paper's practical contribution to that exercise is two concrete methods for estimating the expected market return. The first reads the market's position against equilibrium: Tobin's Q and the inverse adjusted dividend yield show whether the market currently stands above or below its long-run equilibrium, and the expectation should be set below the historical average when the market stands above it, and vice versa. The second corrects the averaging of returns for horizon: predictability models estimate the difference between arithmetic and geometric average returns across holding periods, a gap that mean reversion narrows as good and bad years offset, so the appropriate expected return depends on how long the money is invested, which for price controls is the length of the regulatory period and beyond.
The conclusions for regulators (present as a highlighted panel)
Return averaging matters: the paper sets out the appropriate use of log and geometric returns for long-horizon expectations, where mean reversion narrows their gap with arithmetic averages.
The preferred estimates of the expected market return for UK price controls sit below recent UK regulatory assumptions.
Table 7: What the forecasting result implies for a UK price control, in the numbers a regulator actually sets.
The findings, not the machinery, are the conclusion. Stock returns contain a predictable component at long horizons. Over 119 years, the US market and its fundamentals are bound by a single equilibrium; long-horizon returns are markedly less uncertain than short-horizon volatility implies; a forward-looking model built on that equilibrium forecasts more accurately than the historical average, in-sample and out; and at 2019 valuations it implies US equity returns around 1.5 percentage points below the historical average annually through 2040.The implication runs straight into practice: the expected market return moves with valuations, so any decision that rests on one, an asset allocation, a discounted valuation, an allowed return, should look at where the market stands, not only at where it has been.
For UK price controls the stakes are concrete and cumulative. Allowed returns apply to regulated asset bases in the tens of billions of pounds over multi-year periods, so the gap between a purely historical anchor and a valuation-informed estimate compounds into material customer bills, and the credibility of long-term regulatory decision-making depends on the estimate being right on average across successive controls, which is why methods that incorporate valuations provide a stronger basis for repeated long-term regulatory decisions.
The message for the next price control is short: at current valuations the expected market return is measurably lower than the past suggests, and the tools for measuring it are on the table.It is worth being clear about what this work is and is not for. It is not a case for trying to beat the market, and nothing here suggests an investor can time entries and exits on a valuation ratio. Its purpose is regulatory: to give the bodies that set allowed returns a more robust method for estimating the expected market return that underpins the regulated cost of equity, one that reads the market's position against its long-run equilibrium rather than resting on an average of the past alone. That is why MCC pursues this work. The same number that decides whether an investor's portfolio meets its objective decides what millions of households pay for water and energy, and it deserves to be estimated with the best evidence available.