The Problem of Lucky Guesses

Any prediction system - even a random one - will occasionally issue a correct prediction. If an analyst makes 100 earthquake predictions and 5 come true, that is not evidence of skill: depending on the tolerances and the earthquake rate in the region, a random system might achieve the same hit rate by chance. This is not a criticism unique to Seismikon; it is a fundamental challenge for all probabilistic forecasting systems.

The Statistical Threshold

To claim skill above chance, a prediction system needs to demonstrate that its performance is statistically significantly better than the null model - not just slightly better in a small sample. The required sample size depends on the base rate of events, the tolerances, and the desired confidence level.

Seismikon uses an empirical Monte Carlo null model: thousands of random forecasts are generated with the same geographic and temporal distribution as Seismikon's actual predictions. The skill score measures how far Seismikon's performance falls from the null distribution - the further the better.

Where Seismikon Currently Stands

The live prediction service has been issuing prospective forward predictions for a limited period. As prospective data accumulates, the statistical confidence in the skill claim increases. Current results are published on the Track Record page; the sample size needed for publication-grade statistical significance has not yet been reached for all metrics.

This is an honest limitation, not a dismissal of the system. The retrospective Bikos (2024) paper provides strong evidence of the methodology's validity; the prospective record is the ongoing real-world test that will determine how well it generalises.

Why Seismikon Publishes Every Result

Selective reporting - publishing successful predictions and downplaying failures - is one of the most damaging forms of bias in prediction research. Seismikon's commitment to publishing every forward-issued prediction, including FPs and FNs, is the structural defence against this. Any observer can download the full record and compute the skill score themselves. See Data Downloads.

A single striking prediction - even if it precisely matches a M7+ earthquake - does not establish skill. It is the population of predictions, evaluated over time, that produces a defensible scientific claim.