This multi-horizon forecasting study asks whether a compact Transformer actually earns its complexity over persistence, seasonal-naive, ridge, and gradient-boosted baselines on a long scientific time series. It evaluates several forecast horizons, seed stability, context-length sensitivity, early-versus-late test behavior, and performance during high solar activity.
The Transformer has lower mean error than the simpler baselines at the 1- and 6-month horizons under the recorded protocol. The 12-month result is less stable: its three-seed mean MAE is 24.206, versus 23.973 for histogram gradient boosting, despite seed 42 favoring the Transformer. The seed-42 intervals against that baseline also cross zero at 6 and 12 months. This qualifies any broad claim that the Transformer is consistently superior.
- Calculations, evidence and verification scope
- Figure sources and exact numerical paths
- Working paper
- Data and provenance
Review scope: The existing suite requires unavailable dependencies; no full-suite pass is claimed. The complete data/model experiment was not rerun in this review. Stored empirical results were inspected, not independently reproduced from raw data.
Research Bundle · AI Engineering · multi-horizon time-series forecasting
This repository asks a deliberately falsifiable question: does a compact Transformer earn its complexity over strong simple and tabular baselines on a long real scientific time series? The source is the official WDC-SILSO Version 2.0 monthly mean total sunspot number. The Transformer is a hypothesis, not the presumed winner.
- Does a compact Transformer improve 1-, 6-, or 12-month forecasting over persistence, seasonal-naive, ridge autoregression and histogram gradient boosting?
- Are Transformer results stable across random seeds?
- Does the conclusion change between early and late parts of the same final test era?
- Does performance deteriorate disproportionately during high solar activity?
- Does the result change when context length changes while trainable model size and target dates remain controlled?
The runner downloads the official SN_m_tot_V2.0.csv from WDC-SILSO. SILSO marks CSV rows with 1 for definitive values and 0 for provisional values.
This Research Bundle is frozen to definitive observations from 1749-01 through 2026-03:
- study rows: 3,327 months;
- study-input SHA-256:
02adc08ef41aca6e5a02a21417d38bd3ed1dc14ab4bb3de5df5c93488c017a9b; - study fingerprint: deterministic
YYYY-MM;sunspot-valuesequence; - later rows are excluded even after they become definitive.
The full downloaded file hash is still recorded as provenance, but it is not the research identity because SILSO is a living data service. If any month/value pair inside the frozen study era changes, the runner fails.
WDC-SILSO identifies the series as Version 2.0, DOI 10.24414/qnza-ac80, under CC BY-NC 4.0.
The comparison uses one common pair of target-date boundaries for every horizon and context-length analysis.
- primary context: 132 months;
- horizons: 1, 6 and 12 months;
- final 20% of the frozen raw monthly timeline: test targets;
- final 15% of the pre-test era: validation targets;
- the same validation/test target months are reused across all horizons;
- the same target months are reused for 60-, 132- and 264-month context sensitivity;
- normalization is fitted only through the last training target;
- evaluation is rolling-origin direct forecasting with observed history: later test targets may use earlier realized test-era observations in their context, but never the target value or any future observation.
This avoids a subtle confound where changing context length also changes the evaluation era. It also makes clear that the final 20% is not forecast recursively from one fixed historical origin.
Baselines:
- persistence;
- 12-month seasonal naive;
- ridge autoregression;
- histogram gradient boosting.
Ridge and histogram gradient boosting are selected from small declared grids using the same chronological validation era and validation MSE principle used for Transformer checkpoint selection. HGB's internal random early stopping is disabled, so baseline selection does not introduce a hidden random validation split.
The Transformer uses:
- input projection dimension 32;
- four attention heads;
- feed-forward width 64;
- one encoder layer;
- fixed sinusoidal positional encoding;
- early stopping on the chronological validation era;
- seeds 13, 42 and 73.
Fixed sinusoidal positions are important here: changing context length no longer changes the number of trainable positional parameters.
- MAE and RMSE for 1-, 6-, and 12-month horizons;
- repeated Transformer seeds;
- 12-month moving-block bootstrap summaries for Transformer-vs-baseline MAE differences for the primary seed;
- the same paired comparison for every declared Transformer seed, plus across-seed delta mean/SD and direction counts;
- retrospective high-activity versus other-period error stratification using a training-era threshold; the realized test value defines the diagnostic group and is never a model input;
- early-versus-late test-era robustness;
- context sensitivity at 60, 132 and 264 months on identical horizon-1 target dates and the same 60-epoch cap and early-stopping policy;
- 12-month moving-block Transformer-vs-baseline MAE intervals for all four primary baselines: persistence, seasonal naive, ridge and HGB.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
PYTHONPATH=. pytest -q
PYTHONPATH=. python src/run_experiment.pyGenerated evidence includes results/metrics.json, results/summary.md, three horizon figures, paper/results.md, and paper/results.tex.
A win on one horizon does not establish general Transformer superiority. Sunspot dynamics are periodic, nonstationary and physically structured. Results remain conditional on the frozen study input, forecast horizon, evaluation era, baseline set and optimization protocol. A simpler model winning is a valid and useful outcome.
README.md → DATA.md → src/run_experiment.py → results/summary.md → results/metrics.json → RESEARCH_BUNDLE.md → REPRODUCIBILITY.md → ETHICS.md → paper/paper.md.
Repository source code is released under the MIT License. The WDC-SILSO source data are not relicensed by this repository; SILSO identifies them under CC BY-NC 4.0 with attribution requirements. Generated metrics and figures are provided as research evidence, and users should preserve SILSO attribution and assess the source-data license when reusing data-derived artifacts. See DATA_LICENSE.md.
The strict-QA empirical rerun uses the frozen SILSO study input and the final all-seed comparison protocol.
- At 1 month, the Transformer beats HGB on MAE for all 3/3 seeds; mean Transformer-minus-HGB MAE delta = -0.992.
- At 6 months, the Transformer beats HGB on MAE for all 3/3 seeds; mean delta = -0.846.
- At 12 months, the conclusion is seed-sensitive: 2/3 seeds beat HGB, while the across-seed mean delta is +0.233 MAE, slightly favoring HGB on average.
- The 12-month seed-42 comparison alone is therefore not treated as evidence of general Transformer superiority.
CI binds the committed results to the frozen study fingerprint, common target dates, all four baseline comparisons, all three Transformer seeds, rolling-origin evaluation semantics, and the declared context-sensitivity controls.