What Is Retrospective Testing?

Retrospective (or back-test) analysis evaluates a model's performance against historical data that was available when the model was built. In earthquake science, this means training a model on seismic catalogues from past decades and measuring how well it "predicts" earthquakes that are already known to have occurred.

Retrospective results are necessary and useful for model development - you cannot deploy a model with no performance history. But they are inherently optimistic: the model has, in some sense, been shaped by the data it is tested on, either through direct training or through design choices made by researchers who knew the data.

What Is Prospective Testing?

Prospective testing means issuing a prediction before the outcome is known, then recording whether the prediction was correct when the outcome is revealed. The model had no access to the future data at prediction time.

Prospective performance is the scientific gold standard because it eliminates all forms of look-ahead bias. A model that achieves 98% precision in a retrospective back-test but 30% in prospective testing has overfit its training data; its back-test claim is scientifically misleading.

How Seismikon Handles This Distinction

Seismikon is explicit about which figures are retrospective and which are prospective:

  • The 98% precision figure from Bikos (2024) is a retrospective back-test result. It is clearly labelled as such on the Research Methodology page and should not be cited as live performance.
  • The Track Record page shows only prospective results - every forward-issued prediction evaluated against real outcomes after the window closed.
  • The Evaluation Protocol was published before the first prospective prediction was issued, ensuring no protocol adjustments can be made in response to results.

Why This Matters for Users

Any earthquake prediction system - or any forecasting system in any domain - should be evaluated primarily on its prospective performance. If a vendor only cites back-test figures, ask how long the prospective record is and whether the evaluation protocol was fixed before predictions were issued. Seismikon publishes both, clearly labelled, so users can make informed assessments.

The Bikos (2024) retrospective figures establish scientific validity of the approach. The live Track Record is the ongoing real-world test. Both are needed; neither alone is sufficient.