Elexion election intelligence← FORECAST DESK
MODEL GOVERNANCE / VERSION 0.6.1

Forecast methodology

Evidence gates choose the public model. No challenger is promoted because it looks more sophisticated.

Current publication state

The public forecast uses a widened baseline ensemble. Version 0.6 scores that public baseline alongside Gaussian Monte Carlo, Markov momentum, polls-only, fundamentals-only, and previous-election benchmarks. Every production model-fold uses 1,000,000 predictive draws for winner probabilities and 90% intervals. Training origins must fall within 10% of the held-out horizon, bounded to a two-to-thirty-day tolerance. Production horizons must be near an evaluated fold—not merely between the shortest and longest tests. No challenger is promoted because 2–14-day U.S. evidence cannot validate the current early-cycle horizon. The 12 U.S. folds use a pinned retrospective poll compilation, not forecast-origin archived vintages, so they also fail the vintage-proof gate. Forecasts without source-vintage feature snapshots are forced to grade D and list zero model-input sources.

Türkiye now has three archive-verified forecast-origin folds, each using 1,000,000 draws after training on 2014 and 2018. All three hold out the same 2023 election, cover only 2–14-day horizons, and span nine years, so they remain diagnostic and cannot validate or promote a model. The current headline is a separate mixture of two May 2026 matchups—Erdoğan against Yavaş and Özel. İmamoğlu receives zero active-scenario weight while his degree remains annulled because higher education is a presidential eligibility requirement; a stay or reversal by the Council of State would trigger reassessment. Equal weights across the two remaining matchups are structural placeholders, not nomination probabilities.

Australia now has 14 archive-verified folds across five held-out elections and 21 years of history. Markov momentum has the best election-clustered Brier score in the 7–28-day tests (0.166 versus 0.223 for Gaussian Monte Carlo and 0.279 for the fitted baseline). The historical reliability gates now pass, but the next election remains far outside the tested horizon, so the challenger is not used for the current long-range forecast.

Research doctrine: context without guesswork

Version 0.6 disables every hand-written economy, security, conflict, crime, and incumbency coefficient. Context rows remain reporting signals, but contribute exactly zero to published probabilities. A driver activates only after its value was observable at each historical cutoff, its direction is fitted from training elections, and the complete country-specific model beats simpler alternatives on unseen elections.

The polls-versus-fundamentals weight is no longer fixed at 72/28. Every historical fold fits a constrained zero-to-one weight using prior elections only, with each election receiving equal weight regardless of archive density. The held-out election cannot influence that coefficient. A current forecast may use the fitted blend only when its country report passes every reliability and horizon gate; otherwise an available poll aggregate remains polls-only.

There is no universal “war moves voters right” rule. Research finds conditional, time-varying, and sometimes opposite security effects. The intended model first gates on current issue salience, then uses party ownership, incumbent responsibility, shock timing, geography and decay. A security coefficient stays off when current salience is low or country-specific walk-forward validation is insufficient; it is never activated merely because a conflict exists.

Time and candidate uncertainty

Model volatility is calibrated at a 90-day reference horizon. Version 0.6 applies a transparent time multiplier: (days to election ÷ 90)0.18, bounded from 0.75× to 1.60×. Long-range forecasts therefore widen automatically instead of behaving like twelve-week calls. For unsettled ballots, one million draws sample a candidate scenario first and electoral uncertainty second; the public headline is the resulting mixture, while every conditional distribution remains visible.

One million runs reduce numerical simulation noise. They do not erase polling error, candidate uncertainty, model misspecification, or missing historical validation; those remain visible through intervals, scenario splits, quality grades, fold counts, and vintage-proof status.

Forecast availability policy

Official nominations, final electoral mechanics, and machine-reuse permission affect certainty. A one-million-run forecast publishes only when a defensible electoral probability target exists. When a genuine ballot is unsettled, the model uses explicitly labeled candidate, party, alliance, or governing-versus-opposition scenarios. Where there is no national popular election—or evidence cannot support a probability—the country remains a sourced calendar-only record. Reference-only sources are linked but never ingested.

These proxy forecasts are grade D. They are not presented as validated candidate polls, and they cannot promote a challenger model. Names and mechanics replace proxies as reproducible source-vintage evidence arrives.

CHALLENGER A

Gaussian Monte Carlo

Samples exchangeable zero-sum multi-contestant shocks, turnout uncertainty, house effects, and election-system translation. Unvalidated contextual drivers have no directional effect. Contestant order never assigns a favorable or unfavorable sign.

CHALLENGER B

Markov momentum

Uses a three-state campaign-movement chain with persistent down, neutral, and up states. Transition probabilities and movement size are estimated only from training elections; remaining campaign steps follow the forecast horizon.

Promotion gates

  1. At least eight strict forecast-origin folds across three held-out elections and twenty years of history
  2. Immutable revision hashes for fundamentals, every poll snapshot, and results
  3. All inputs must have release timestamps at or before each forecast cutoff
  4. Election-clustered paired-bootstrap Brier superiority at 90% confidence
  5. Vote-share RMSE no more than 5% worse than best baseline
  6. Empirical 90% interval coverage of at least 80%

Challengers are compared with a training-fitted poll/fundamentals ensemble, polls-only, fundamentals-only, and previous-election baselines. Failure of any gate retains the simpler public fallback.

System engines and limits

Engines cover presidential runoff transfers, FPTP seat elasticity, thresholded proportional and mixed-member allocation, electoral-college translation, institutional regional paths, and unresolved national-control scenarios. District or state maps remain suppressed until validated boundary-level inputs exist. Early structural forecasts carry low quality grades and wide intervals.

Driver sensitivity

Sensitivity matrices are suppressed unless coefficients come from a promoted, source-vintage model. Qualitative context can explain what analysts are monitoring; it cannot silently alter a probability.

Data integrity

Raw responses are content-addressed and immutable. Every observation preserves observed, released, available, and retrieved timestamps. Model cutoffs enforce all four clocks. Adapters reject unapproved reuse terms before network access and retain last-known-good canonical records when parser confidence falls.

G20 coverage states

Public scope is limited to the 19 sovereign G20 countries. The European Union and African Union are excluded. All 19 now have sourced national election-status records. Sixteen carry forecasts; China and Saudi Arabia remain calendar-only because no national popular executive or legislative-control ballot exists, while Russia remains calendar-only until an official timetable, candidate field, and approved source-vintage evidence support a probability.