Skip to content

Latest commit

 

History

History
138 lines (102 loc) · 6.11 KB

File metadata and controls

138 lines (102 loc) · 6.11 KB

Backtest-period false-positive and tail-risk timeline

Question

Did the false-positive and severe-loss problem occur throughout the corrected backtest from January 10, 2020 through June 30, 2023, or was it created by a few abnormal observations?

This analysis uses the corrected sector-capped top 40 for Ridge and XGBoost. Every half-year contains approximately 25 to 27 weekly decisions and 1,000 to 1,080 selected-stock observations.

Ridge timeline

Period Positive precision Severe-loss rate Mean selected return Severe share of negative drag Losing weeks Worst week
2020 H1 51.60% 12.00% 0.90% 61.08% 52.00% -21.75%
2020 H2 55.93% 0.74% 0.96% 7.72% 40.74% -4.53%
2021 H1 58.00% 0.60% 0.75% 6.49% 32.00% -3.13%
2021 H2 54.63% 2.13% 0.58% 18.56% 40.74% -5.30%
2022 H1 43.60% 6.10% -0.53% 32.58% 68.00% -10.23%
2022 H2 58.89% 2.59% 0.79% 20.47% 37.04% -6.72%
2023 H1 52.31% 1.06% 0.28% 9.33% 53.85% -4.35%

Ordinary negative holdings appear in every period. Even the best half-year, 2022 H2, has 41.11% negative selected-stock observations. The sign false-positive problem is therefore persistent.

The severe-loss problem is much more concentrated:

Period Severe observations Share of all Ridge severe observations Share of Ridge severe-loss drag
2020 H1 120 46.69% 53.93%
2020 H2 8 3.11% 2.81%
2021 H1 6 2.33% 1.85%
2021 H2 23 8.95% 7.18%
2022 H1 61 23.74% 21.86%
2022 H2 28 10.89% 9.26%
2023 H1 11 4.28% 3.11%

2020 H1 and 2022 H1 together contain 70.43% of Ridge's severe observations and 75.79% of its severe-loss drag.

Outside those two shock halves, Ridge still has:

  • 55.95% positive precision;
  • 44.05% sign false positives;
  • 1.44% severe-loss rate;
  • 0.674% mean selected return.

This means the ordinary false-positive problem remains, but most economically damaging tail risk is associated with stressed periods.

XGBoost comparison

Period Positive precision Severe-loss rate Mean selected return Severe share of negative drag Losing weeks
2020 H1 49.50% 12.80% -0.18% 57.53% 52.00%
2020 H2 53.52% 0.65% 0.90% 5.76% 37.04%
2021 H1 58.00% 1.80% 1.05% 19.54% 32.00%
2021 H2 51.85% 3.24% 0.27% 24.70% 48.15%
2022 H1 41.30% 5.60% -1.01% 27.27% 72.00%
2022 H2 54.63% 2.31% 0.25% 18.78% 37.04%
2023 H1 51.92% 1.63% 0.21% 15.12% 34.62%

XGBoost shows the same temporal pattern but performs worse overall. Its 2020 H1 and 2022 H1 periods contain 64.34% of severe observations and 67.66% of severe-loss drag.

Market-state attribution

The realized market state is an attribution label using future SPY and universe returns. It is not available when the portfolio is selected.

Ridge

Realized market state Dates Losing weeks Mean portfolio return Severe observations per week
Broad negative 65 89.23% -2.58% 3.18
Mixed 34 50.00% -0.11% 0.68
Broad positive 83 10.84% 3.25% 0.33

The broad-negative dates explain much of the sign false-positive behavior. When most stocks fall, relative winner selection cannot ensure positive absolute returns.

The causal decision-time SPY regime has weaker discrimination:

Decision-time regime Dates Losing weeks Mean portfolio return Severe observations per week
Strong 126 45.24% 0.53% 0.79
Weak 56 48.21% 0.57% 2.80

The weak regime identifies a higher tail-event frequency, but it does not identify losing weeks or negative average returns reliably enough to justify moving the complete portfolio to cash.

Worst Ridge episodes

Decision date Portfolio return Negative holdings Severe holdings Worst holding
2020-03-06 -21.75% 40 37 RCL, -46.83%
2020-03-13 -11.26% 35 23 BA, -33.91%
2022-06-03 -10.23% 39 22 EXPE, -17.40%
2020-02-21 -9.93% 40 21 DRI, -17.92%
2022-09-16 -6.72% 40 5 MTCH, -12.74%

The worst individual Ridge observations include RCL in March 2020, EPAM in February 2022, and Netflix in April 2022. These are genuine corrected market events rather than split-scale errors.

Concentration

Across 182 Ridge decisions:

  • 84 weeks lose money;
  • 65 weeks contain at least one severe-loss holding;
  • 13 weeks contain at least five severe-loss holdings;
  • removing the single worst stock reverses only 7 of the 84 losing weeks.

The problem is therefore not simply one bad company in every losing week. The largest drawdowns usually combine a common market decline with several selected-stock losses.

Conclusion

Yes, the same issue appears in the corrected 2020 through 2023 backtest, but it has two different forms:

  1. Persistent sign uncertainty. Approximately 40% to 56% of selected stocks can be negative in any half-year. This is inherent to noisy stock returns and cross-sectional ranking.
  2. Regime-concentrated severe losses. Most damaging losses occur during 2020 H1 and 2022 H1, especially on broad-negative market dates.

This supports a narrow tail-risk guardrail rather than replacing the winner ranker with a positive return classifier. It also explains why the first E4 treatment reduced severe-loss drag but could not establish a general CVaR or return advantage: tail events are clustered, rare outside shocks, and difficult to identify from the available causal regime signal.

Evidence