2008–2026 · disjoint 3-month returns
Does the forecast survive a full cycle?
A predeclared cross-asset experiment, 74 quarterly tests and a decision based on consistency—not the best-looking result.
Stable predictive advantage not established.
The primary model was a fixed blend of market signals, VIX and the Treasury yield curve, with expanding and recent training windows. It did not pass the predeclared historical consistency screen. The incumbent V2 remains a research forecast; neither version has demonstrated a stable live edge.
Specification and evaluation: 13 September 2026 · retrospective reconstruction · zero completed prospective observations95% interval -2.71% to +3.25%
At least 5 required before promotion
Historical reconstructions are excluded
Where did the improvement hold?
Positive bars mean lower squared forecast error than the same historical-mean benchmark. Select a period to inspect its results.
Only the fixed ensemble was eligible for promotion. Changing this view does not select a replacement model. Periods are grouped by forecast-origin year.
| Origin period | Quarterly windows | Primary error skill | V2 error skill |
|---|---|---|---|
| 2008–2010 | 12 | -2.98% | -3.90% |
| 2011–2013 | 12 | -3.71% | -2.07% |
| 2014–2016 | 12 | -4.73% | -3.41% |
| 2017–2019 | 12 | +0.07% | -0.33% |
| 2020–2022 | 12 | +1.54% | +2.83% |
| 2023–2026 | 14 | +3.71% | +3.12% |
Every gate must pass.
Lower error than historical mean
0.46 % skill
No worse than incumbent V2
0.10 % skill
Both block intervals entirely above zero
-2.71 % minimum lower bound
Positive skill in at least five of six periods
3 of 6 periods
Worst period no lower than −5% skill
-4.73 % skill
Positive skill in 2023–2026
3.71 % skill
Positive skill at both alternative quarter starts
0.38 % minimum skill
Passing these historical gates would be a research milestone, not proof of future reliability. Repeated research on existing history can bias even a careful walk-forward study.
Keep the whole comparison visible.
| Model | Error skill vs mean | Typical error · RMSE | Correct direction | Role |
|---|---|---|---|---|
| Historical mean | 0.00% | 13.27 pp | 74.3% | Benchmark |
| V2 · market signals | +0.37% | 13.25 pp | 73.0% | Incumbent research model |
| Cross-asset ridge | -0.50% | 13.30 pp | 73.0% | Diagnostic component |
| Multi-window average | +1.25% | 13.19 pp | 74.3% | Diagnostic component |
| Fixed ensemble · primary | +0.46% | 13.24 pp | 74.3% | Predeclared primary |
Always forecasting a positive return was correct in 74.3% of these windows. A high hit rate alone is not an edge.
Alternative quarter starts and market environments
Each quarter-start group is internally disjoint, but groups overlap with one another. They are robustness checks and cannot be pooled as independent observations.
| Origin months | Windows | Primary error skill |
|---|---|---|
| 1 / 4 / 7 / 10 | 74 | +0.46% |
| 2 / 5 / 8 / 11 | 74 | +0.38% |
| 3 / 6 / 9 / 12 | 73 | +1.57% |
Market environments are defined using the trend known at the forecast origin, not the future return.
| Origin environment | Windows | Primary error skill |
|---|---|---|
| Below trend | 21 | -0.57% |
| At or above trend | 53 | +1.28% |
Evidence that arrives before the outcome.
Use archived earnings estimates, original releases and dated industry histories. Today’s revised data cannot stand in for an information set available years ago.
Log the information cutoff, issue time, model version and forecast before the future return path. The next planned monthly origin is 30 September; no forecast has been issued for it and collection is not automatic.
The protocol requires at least 12 disjoint prospective quarterly outcomes before assessing a live-edge claim. This minimum is only a floor; uncertainty and stability must still pass. It takes approximately three years to observe 12 consecutive non-overlapping quarters.
Method, limitations and reproducible files
290 common monthly input rows; 0 dropped rows; 228 expanding forecast origins. All targets mature before entering training. VIX and curve observations have a seven-calendar-day delay and an additional stale-source rejection check.
V2 is the existing four-feature ridge. Cross-asset ridge adds VIX level/change and the 10-year minus 2-year yield spread level/change. The multi-window component averages 16 univariate forecasts using expanding and 120-label histories. The primary ensemble assigns one-third weight to each component. All penalties, scaling and weights were fixed before scoring.
Paired 95% skill intervals use 5,000 circular block draws: two-quarter blocks -2.59% to +3.24%; four-quarter blocks -2.71% to +3.25%. These intervals do not adjust for every earlier research attempt.
- Retrospective walk-forward reconstruction. Specification was made after V1/V2 research and all data already existed. No pristine holdout; repeated research can bias results.
- Each offset contains disjoint three-month windows internally. Different offsets overlap and cannot be pooled or counted as independent tests.
- Use latest observation at least seven calendar days before each actual monthly market close. Previous three-month feature level uses the same rule at the origin three calendar months earlier. Drop missing rows; no backward filling. Latest retrieved histories are not certified point-in-time vintages. Earlier origins use actual SOXX last trading dates. Reject sources more than seven days stale at the delayed cutoff.
- Passing every gate permits a historically robust research label only. It does not establish future reliability. Failure retains V2 as the incumbent research forecast; never pick another component because it looks better in the final test.
- No stable live edge claim without at least 12 disjoint completed prospective quarterly outcomes logged immediately after the information cutoff and before the future return path is observed, with immutable issue timestamps and no historical backdating, compared with frozen benchmarks and recalculated uncertainty. Minimum sample count is a floor, not proof. No completed prospective forecasts exist today.
- No trading strategy or after-cost alpha is established by squared-return forecast error. Forecast comparisons do not imply executable fills at their information-date close.
- Pointwise bootstrap intervals do not adjust for all previous research attempts. Passing a historical screen is not an independent discovery.