Did the false-positive and severe-loss problem occur throughout the corrected backtest from January 10, 2020 through June 30, 2023, or was it created by a few abnormal observations?
This analysis uses the corrected sector-capped top 40 for Ridge and XGBoost. Every half-year contains approximately 25 to 27 weekly decisions and 1,000 to 1,080 selected-stock observations.
| Period | Positive precision | Severe-loss rate | Mean selected return | Severe share of negative drag | Losing weeks | Worst week |
|---|---|---|---|---|---|---|
| 2020 H1 | 51.60% | 12.00% | 0.90% | 61.08% | 52.00% | -21.75% |
| 2020 H2 | 55.93% | 0.74% | 0.96% | 7.72% | 40.74% | -4.53% |
| 2021 H1 | 58.00% | 0.60% | 0.75% | 6.49% | 32.00% | -3.13% |
| 2021 H2 | 54.63% | 2.13% | 0.58% | 18.56% | 40.74% | -5.30% |
| 2022 H1 | 43.60% | 6.10% | -0.53% | 32.58% | 68.00% | -10.23% |
| 2022 H2 | 58.89% | 2.59% | 0.79% | 20.47% | 37.04% | -6.72% |
| 2023 H1 | 52.31% | 1.06% | 0.28% | 9.33% | 53.85% | -4.35% |
Ordinary negative holdings appear in every period. Even the best half-year, 2022 H2, has 41.11% negative selected-stock observations. The sign false-positive problem is therefore persistent.
The severe-loss problem is much more concentrated:
| Period | Severe observations | Share of all Ridge severe observations | Share of Ridge severe-loss drag |
|---|---|---|---|
| 2020 H1 | 120 | 46.69% | 53.93% |
| 2020 H2 | 8 | 3.11% | 2.81% |
| 2021 H1 | 6 | 2.33% | 1.85% |
| 2021 H2 | 23 | 8.95% | 7.18% |
| 2022 H1 | 61 | 23.74% | 21.86% |
| 2022 H2 | 28 | 10.89% | 9.26% |
| 2023 H1 | 11 | 4.28% | 3.11% |
2020 H1 and 2022 H1 together contain 70.43% of Ridge's severe observations and 75.79% of its severe-loss drag.
Outside those two shock halves, Ridge still has:
- 55.95% positive precision;
- 44.05% sign false positives;
- 1.44% severe-loss rate;
- 0.674% mean selected return.
This means the ordinary false-positive problem remains, but most economically damaging tail risk is associated with stressed periods.
| Period | Positive precision | Severe-loss rate | Mean selected return | Severe share of negative drag | Losing weeks |
|---|---|---|---|---|---|
| 2020 H1 | 49.50% | 12.80% | -0.18% | 57.53% | 52.00% |
| 2020 H2 | 53.52% | 0.65% | 0.90% | 5.76% | 37.04% |
| 2021 H1 | 58.00% | 1.80% | 1.05% | 19.54% | 32.00% |
| 2021 H2 | 51.85% | 3.24% | 0.27% | 24.70% | 48.15% |
| 2022 H1 | 41.30% | 5.60% | -1.01% | 27.27% | 72.00% |
| 2022 H2 | 54.63% | 2.31% | 0.25% | 18.78% | 37.04% |
| 2023 H1 | 51.92% | 1.63% | 0.21% | 15.12% | 34.62% |
XGBoost shows the same temporal pattern but performs worse overall. Its 2020 H1 and 2022 H1 periods contain 64.34% of severe observations and 67.66% of severe-loss drag.
The realized market state is an attribution label using future SPY and universe returns. It is not available when the portfolio is selected.
| Realized market state | Dates | Losing weeks | Mean portfolio return | Severe observations per week |
|---|---|---|---|---|
| Broad negative | 65 | 89.23% | -2.58% | 3.18 |
| Mixed | 34 | 50.00% | -0.11% | 0.68 |
| Broad positive | 83 | 10.84% | 3.25% | 0.33 |
The broad-negative dates explain much of the sign false-positive behavior. When most stocks fall, relative winner selection cannot ensure positive absolute returns.
The causal decision-time SPY regime has weaker discrimination:
| Decision-time regime | Dates | Losing weeks | Mean portfolio return | Severe observations per week |
|---|---|---|---|---|
| Strong | 126 | 45.24% | 0.53% | 0.79 |
| Weak | 56 | 48.21% | 0.57% | 2.80 |
The weak regime identifies a higher tail-event frequency, but it does not identify losing weeks or negative average returns reliably enough to justify moving the complete portfolio to cash.
| Decision date | Portfolio return | Negative holdings | Severe holdings | Worst holding |
|---|---|---|---|---|
| 2020-03-06 | -21.75% | 40 | 37 | RCL, -46.83% |
| 2020-03-13 | -11.26% | 35 | 23 | BA, -33.91% |
| 2022-06-03 | -10.23% | 39 | 22 | EXPE, -17.40% |
| 2020-02-21 | -9.93% | 40 | 21 | DRI, -17.92% |
| 2022-09-16 | -6.72% | 40 | 5 | MTCH, -12.74% |
The worst individual Ridge observations include RCL in March 2020, EPAM in February 2022, and Netflix in April 2022. These are genuine corrected market events rather than split-scale errors.
Across 182 Ridge decisions:
- 84 weeks lose money;
- 65 weeks contain at least one severe-loss holding;
- 13 weeks contain at least five severe-loss holdings;
- removing the single worst stock reverses only 7 of the 84 losing weeks.
The problem is therefore not simply one bad company in every losing week. The largest drawdowns usually combine a common market decline with several selected-stock losses.
Yes, the same issue appears in the corrected 2020 through 2023 backtest, but it has two different forms:
- Persistent sign uncertainty. Approximately 40% to 56% of selected stocks can be negative in any half-year. This is inherent to noisy stock returns and cross-sectional ranking.
- Regime-concentrated severe losses. Most damaging losses occur during 2020 H1 and 2022 H1, especially on broad-negative market dates.
This supports a narrow tail-risk guardrail rather than replacing the winner ranker with a positive return classifier. It also explains why the first E4 treatment reduced severe-loss drag but could not establish a general CVaR or return advantage: tail events are clustered, rare outside shocks, and difficult to identify from the available causal regime signal.