In 2020, the most successful quant fund in history lost more than 20 percent and told its investors it would not be changing the models. The investors who left anyway converted a temporary drawdown into a permanent loss. Every systematic strategy has years like this. What decides your outcome is the yardstick you use to judge them.

In the spring of 2020, the most successful quantitative fund in history was having a terrible year. Renaissance Technologies' public funds lost more than 20 percent while the market recovered from the COVID crash. Meanwhile Medallion, the firm's private fund, gained 76 percent over the same stretch. Investors in the public funds were furious, and many redeemed. Renaissance's response was remarkable for what it did not do. The firm told investors that the models were built on decades of data, that the environment was an outlier, and that it would not be changing them. It held the line. The public funds recovered in the years that followed, and the investors who left near the bottom converted a temporary drawdown into a permanent loss.
I think about that episode a lot, because it captures the hardest problem in systematic investing. Building a strategy that works is the easy half. The hard half is knowing what to do when a strategy that works stops working for a while. Every rules-based strategy will have stretches, sometimes long ones, where it underperforms. Whether those stretches arrive was never in question. What decides your outcome is the yardstick you use to judge them.
The intuitive answer is to judge the returns. A strategy that is losing money is failing, and a strategy that is making money is succeeding. This feels obviously true, and it is statistically almost useless over the horizons investors actually evaluate.
Andrew Lo worked out the math in a 2002 paper on the statistics of Sharpe ratios.1 The standard error on a Sharpe ratio estimated from a year or two of returns is enormous. A genuinely skilled strategy and a worthless one produce overlapping return distributions over short windows, and you cannot reliably tell them apart by looking at recent performance. You need many years, sometimes decades, before the signal separates from the noise. Nobody waits that long. Research on institutional behavior shows that even professional allocators evaluate managers on roughly three-year windows, which the math says is far too short.
Wesley Gray made the point vivid in a 2016 study he titled Even God Would Get Fired as an Active Investor.2 He constructed a portfolio with perfect foresight, one that held the exact stocks that would perform best over the following five years. That portfolio, built on literal omniscience, still suffered drawdowns of more than 20 percent and long stretches of underperformance against the index. A manager running it would have been fired on the strength of the track record alone. If perfect foresight cannot survive evaluation by short-term returns, no honest strategy can.
The damage from this evaluation error shows up in real money. Goyal and Wahal studied thousands of hiring and firing decisions by pension plans and found a consistent pattern.3 Plans fired managers after periods of underperformance and hired replacements with strong recent records. The fired managers then outperformed the managers hired to replace them. The plans were systematically selling low and buying high, at the strategy level, with billions of dollars. Cornell, Hsu, and Nanigian later showed that a naive strategy of hiring the recently worst-performing managers beat hiring the recent winners.4 Chasing performance destroys value even when everyone involved is a professional.
There is a specific version of this mistake that kills quantitative strategies, and it deserves its own name. I call it redesigning at the bottom.
In August of 2007, quantitative equity funds suffered losses so fast and so synchronized that the episode is still called the quant quake. Khandani and Lo reconstructed what happened.5 A large multi-strategy fund began unwinding positions, the unwind moved prices against every fund holding similar positions, and the losses cascaded. Here is the part that matters. The strategies themselves rebounded within days. Funds that held their positions recovered most of the damage almost immediately. Funds that cut exposure at the trough locked in the losses permanently. Same models, same week, opposite outcomes, and the difference was entirely the decision made at the moment of maximum pain.
The pattern repeated during what practitioners now call the quant winter of 2018 through 2020, when value and other factor strategies underperformed for three consecutive years. Cliff Asness of AQR spent that period publishing research arguing that factor timing is deceptively difficult and that the discipline is to hold, with at most modest tilts.6 It was an unpopular position while the losses compounded. Then value snapped back violently in 2021 and 2022, and the investors who had capitulated missed the entire recovery.
The mirror image is just as instructive. Winton, one of the pioneers of trend following, reduced its allocation to trend strategies after years of mediocre results. The reduction landed in the years before 2022, which turned out to be one of the best years for trend following in decades. Cutting a strategy after a drawdown is emotionally understandable, and it is also, empirically, the worst-timed decision in the industry's history. Strategy returns are not persistent in the short run, and a drawdown tells you almost nothing about the next period.
So if returns are a noisy witness, what should you interrogate instead? The answer the best practitioners converge on is behavior. Not did the strategy make money, but did the strategy do what it was designed to do, in the environment it actually faced.
This reframing only works if you do something uncomfortable. You have to write down what the strategy is supposed to do before it does it. Clinical researchers learned this lesson decades ago. A drug trial that defines success after seeing the data can prove anything, which is why trials pre-register their endpoints. The same logic applies to investing. Bailey and Lopez de Prado showed how easily backtested performance can be manufactured by testing many variations and reporting the best one, and they built statistical corrections for exactly this problem.7 Harvey, Liu, and Zhu went through hundreds of published factor discoveries and concluded that most fail once you account for how many things were tried.8 The common thread: criteria chosen after the fact are worthless. Evaluation has to be committed in advance.
In practice, judging behavior means holding a strategy to three pre-committed standards. First, an expected range. From the strategy's design and history, you can state how deep its drawdowns tend to run, how long it typically trails its benchmark, and how often losing quarters arrive. A drawdown inside that envelope is the strategy operating as documented, not evidence of failure. Second, a mechanism check that is separate from the outcome check. A trend strategy is supposed to be long in uptrends. A defensive signal is supposed to fire near stress. You can verify whether the machinery is functioning without reference to whether last quarter was green, and this distinction is the entire diagnosis. Underperformance with intact machinery is a regime cost. Underperformance with broken machinery is decay. They look identical in a return chart and they demand opposite responses. Third, retirement criteria defined before deployment. There are legitimate reasons to kill a strategy. Markets change structurally, and edges do decay. But the evidence that justifies retirement has to be specified in advance, in mechanism terms, or the decision will always be made at the bottom, by pain.
This is the discipline that separates a process from a hope, and it is why I believe a strategy should publish its expected behavior across market regimes as prominently as it publishes its returns. When the expected behavior is on the page, an investor can ask the only question that matters during a drawdown. Is this what the strategy said it would do?
There is a deeper reason to insist on this frame. A strategy's bad environments are the price of its good environments, and the two are usually inseparable. Nobody gets to engineer the bill away.
Momentum is the cleanest example. Daniel and Moskowitz documented what they called momentum crashes, brief violent periods, usually in sharp reversals after panics, where momentum strategies suffer their worst losses.9 Those crashes come structurally attached to the same behavior that generates momentum's long-run premium. Trend following pays for its crisis protection with whipsaw losses in choppy, directionless markets. Defensive strategies pay for their drawdown protection with lag in fast recoveries. Every durable strategy carries a bill for some regime, and a strategy that claims otherwise has simply not disclosed which regime sends the bill.
Once you see strategies this way, the portfolio conclusion follows on its own. If every strategy has a characteristic failure mode, robustness cannot live at the strategy level. It lives at the portfolio level, in the deliberate pairing of strategies whose failure modes do not overlap. This year has been a live demonstration. Strategies driven by macro signals spent much of 2026 positioned cautiously while markets climbed through war headlines and rate repricing, and strategies driven purely by price rode the same trend without hesitation. Neither approach is broken. They read the world through different instruments, and they pay their bills in different years. Held together, each covers the other's characteristic weakness, which is something no amount of tuning can achieve within a single strategy.
Diversification across tickers is familiar. Diversification across failure modes is the version that actually protects a systematic portfolio.
All of this compresses into a short interrogation. Does the strategy state, in advance and in plain language, how it is expected to behave across different market environments? Does it name its bad environments specifically, rather than gesturing at generic risk? Does its manager define, before deployment, what evidence would justify retiring it? And when the drawdown arrives, as it will, is the behavior inside the documented range, and is the machinery still doing its job?
If the answers are yes, the drawdown is tuition, already priced into the design. If the strategy cannot answer these questions at all, the returns were never the problem.
Renaissance could hold the line in 2020 because the firm knew, quantitatively and in advance, what its models' bad stretches looked like. The investors who redeemed had only the returns to go on, and the returns, as always, were lying about the horizon that mattered.
1. Lo, A. W. (2002). The statistics of Sharpe ratios. Financial Analysts Journal, 58(4).
2. Gray, W. (2016). Even God would get fired as an active investor. Alpha Architect.
3. Goyal, A., and Wahal, S. (2008). The selection and termination of investment management firms by plan sponsors. The Journal of Finance, 63(4).
4. Cornell, B., Hsu, J., and Nanigian, D. (2017). Does past performance matter in investment manager selection? The Journal of Portfolio Management, 43(4).
5. Khandani, A. E., and Lo, A. W. (2011). What happened to the quants in August 2007? Evidence from factors and transactions data. Journal of Financial Markets, 14(1).
6. Asness, C., Chandra, S., Ilmanen, A., and Israel, R. (2017). Contrarian factor timing is deceptively difficult. The Journal of Portfolio Management, 43(5).
7. Bailey, D. H., and Lopez de Prado, M. (2014). The deflated Sharpe ratio: correcting for selection bias, backtest overfitting, and non-normality. The Journal of Portfolio Management, 40(5).
8. Harvey, C. R., Liu, Y., and Zhu, H. (2016). ...and the cross-section of expected returns. The Review of Financial Studies, 29(1).
9. Daniel, K., and Moskowitz, T. J. (2016). Momentum crashes. Journal of Financial Economics, 122(2).
Build, test, and deploy your own strategies in Portfolio Lab. Free to start.
Educational & Research Disclosure: The content provided is for informational and educational purposes only and is not intended to constitute investment advice, a recommendation, solicitation, or offer to buy or sell any security. Any discussion of market trends, historical performance, academic research, models, examples, or illustrations is presented solely to explain general financial concepts and does not represent a prediction, guarantee, or assurance of future results. Past performance is not indicative of future results. All investing involves risk, including the possible loss of principal.

In 2020, the most successful quant fund in history lost more than 20 percent and told its investors it would not be changing the models. The investors who left anyway converted a temporary drawdown into a permanent loss. Every systematic strategy has years like this. What decides your outcome is the yardstick you use to judge them.