Goldman Sachs and a single operator working with AI coding agents independently built Monte-Carlo forecasters for the champion of the 2026 World Cup. We place the two systems side by side on the one question both answer, the identity of the winner, and report what each side did and what it cost to do it. The two converge. Both rate Spain the modal champion, at 24.0 percent for the homegrown Kickoff engine and 25.7 percent for Goldman Sachs, and both lean the same way against the betting market, overweight Spain and Argentina and underweight England and Portugal, with Spain and Argentina together holding 40.8 percent of the Kickoff mass and 40.0 percent of the Goldman Sachs mass. Their largest disagreement is France, which the Goldman Sachs scoring-talent covariate prices at 18.9 percent against the Kickoff figure of 12.0 percent. A seven-dimension assessment scores the two even as champion forecasters with orthogonal strengths: Goldman Sachs leads on estimation sample and covariate breadth, Kickoff on calibration evidence, leakage control and reproducibility. The Kickoff forecast was built in nine days on free public data and beats the devigged betting market in three of four backtest World Cups. Accuracy cannot settle the comparison, because the 48-team format has no precedent and the two forecasts were frozen on different dates.
The 2026 World Cup is the first to field 48 teams: twelve groups of four, the top two plus the eight best third-placed teams advancing to a round of 32, then a knockout to the final, with finalists playing eight matches. Two teams built a probabilistic forecaster for its champion on the same skeleton, an Elo-anchored Poisson model feeding a Monte-Carlo simulation of the actual bracket. One is Goldman Sachs Global Investment Research, whose 29 May 2026 note extends a recurring World Cup research product (Hatzius, Pierdomenico, and Stehn, 2026). The other is Kickoff, an engine that a single operator built in nine days with AI coding agents on free public data, and that the conjunction paper MKT-004 (ATOL Research, 2026) used to select a contest ticket. Both forecasters accept the group draw of 5 December 2025 as a fixed input, so any divergence between them is a property of the models rather than of the draw.
The two systems are commensurable on one axis, the champion marginal, and the comparison holds them to it and asks what each side did and what it cost to do it. The accounting is deliberately even. Goldman Sachs is the better-resourced and better-estimated model, and this paper says so first and plainly: it fits an interpretable regression on roughly twenty thousand matches and carries a richer covariate set than Kickoff chooses to. Kickoff supplies the calibration evidence, the leakage control, and the reproducibility protocol that the public Goldman Sachs note does not. On a seven-dimension assessment the two finish even, with the edges falling on different dimensions rather than accumulating on one side.
The contribution is the comparison itself. A forecasting system assembled by one person with coding agents on zero-cost data holds an even side-by-side against a model from one of the most resourced research divisions in finance, on a task that division has published on for more than a decade. Section 3 places both models in the football-forecasting and prior-ATOL literature. Section 4 describes the two systems, the validation metrics, and the reproducibility protocol. Section 5 reports the head-to-head probabilities, the shared anti-market tilt, the validation evidence, the build ledger, and the assessment. Section 6 interprets the France divergence and the breadth-against-discipline contrast. Sections 7 and 8 state the limitations and what remains for the in-tournament sequel to settle.
Probabilistic football forecasting rests on two primitives that both models reuse. The first is the Poisson goals model with a low-score dependence correction and time-decayed match weighting (Dixon and Coles, 1997), which turns team-strength estimates into a scoreline distribution. The second is the Elo rating system (Elo, 1978), a recursive strength estimate that both models adopt as the dominant driver of expected goals. Forecast quality is measured with proper scoring rules, the Brier score (Brier, 1950) and the logarithmic score, which reward calibrated probabilities and penalize confident errors.
Goldman Sachs has published a World Cup forecast under the title “The World Cup and Economics” across several tournaments, with the chief economist as the through-line and the model extended at each edition (Bloomberg, 2026). The 2026 note states that it is similar to the models for prior World Cups but extended along several dimensions, and it backtests its own 2018 and 2022 forecasts (Hatzius, Pierdomenico, and Stehn, 2026). The note is a research product of a large institution with proprietary data and a named author team.
Within the ATOL corpus, the Kickoff engine was introduced in MKT-004 (ATOL Research, 2026), which used its joint simulation to pick a four-outcome contest ticket and reported the champion backtest reused here. The construction context, a single operator directing a fleet of AI coding agents across production repositories, is documented in MKT-005 (ATOL Research, 2026). Prior ATOL papers compare a model against a market baseline; none places an ATOL system beside an external institutional forecaster. That side-by-side, and its reframing of an accuracy question as a question of what is now buildable, is the gap this paper fills.
Goldman Sachs fits a Poisson goals regression. The number of goals a team scores is modelled as Poisson with a mean driven by the difference in Elo ratings as the main term, a scoring-talent count of the team’s players among the top fifty scorers in the major European leagues capped at four, momentum from goals scored over the last ten matches and conceded over the last five, mentality effects including a reigning-champion slump and a first-World-Cup boost, and geography through a host effect and adjustments for distance, altitude, and temperature (Hatzius, Pierdomenico, and Stehn, 2026). The fitted scoreline distribution feeds a 50,000-draw Monte-Carlo simulation of the 48-team bracket, and the model is re-run after each match day during the tournament.
Kickoff runs an Elo-anchored Monte-Carlo engine over the same bracket structure. Two match engines sit behind one interface: a World-Football-Elo win, draw, and loss model, and a Dixon-Coles bivariate Poisson with the low-score correction and time-decayed weighting. For the champion forecast the Elo engine is selected by calibration, because on held-out champions it scores a log-loss of 1.889 against 2.504 for Dixon-Coles1, which over-rated CONMEBOL sides whose attack ratings are inflated by high-scoring intra-confederation qualifiers. Knockout ties are resolved by an explicit extra-time and penalty-shootout model rather than by a draw-discarding heuristic. The features that did not earn their place were rejected on calibration evidence, including an opponent-defence adjustment and a designated-penalty-taker term, each of which degraded the binding goal-count log-loss in backtest. The two systems are summarized in Table 2 and Table 3.
Table 2. Data and inputs for the two forecasters.
| Input | Goldman Sachs | Kickoff |
|---|---|---|
| Match corpus | About 20,000 mandatory internationals since 1978 | International results 1872 to the present, free |
| Strength prior | Elo plus a proprietary regression | World Football Elo (eloratings.net) |
| Player inputs | Top-50 European-league scorer counts | Not used for the champion marginal |
| Market benchmark | Polymarket and Bet365, retrieved 29 May | Polymarket, devigged and frozen |
| Group draw | Known, 5 December 2025 | Known, 5 December 2025 |
| Leakage stance | Re-run daily, not frozen | Frozen pre-kickoff snapshot, seed-reproducible |
Table 3. Methodology contrast for the two forecasters.
| Dimension | Goldman Sachs | Kickoff |
|---|---|---|
| Covariate set | Rich: Elo, talent, momentum, mentality, geography | Parsimonious: Elo, with Dixon-Coles rejected by calibration |
| Estimation | One fitted Poisson regression, interpretable effects | Elo recursion plus a calibration-selected engine |
| Feature discipline | Several small-sample terms, no reported intervals | Features tested and rejected when they hurt calibration |
| Knockout ties | Discard equal-goal draws toward the higher mean | Explicit extra-time and penalty model |
| Simulations | 50,000 draws | 20,000 to 30,000 draws |
| Updating | Daily re-run | Frozen, no in-tournament update |
Both systems can be scored by the logarithmic loss assigned to the realized outcome,
where $p_{i, y_i}$ is the probability the forecaster placed on the outcome $y_i$ that event $i$ resolved to. For the champion backtest, an event is a tournament and the realized outcome is the actual winner. Goldman Sachs reports a different quantity: the correlation between predicted and actual goal difference, 0.49 across all World Cup games since 1978 (Hatzius, Pierdomenico, and Stehn, 2026). That correlation is informative about scorelines but is not a calibration test, and the note publishes no calibration metric, confidence interval, or acceptance threshold. Kickoff measures calibration directly through the Brier score (Brier, 1950), log-loss, and expected calibration error, benchmarks the champion probability against the devigged closing line, and gates the result behind a blocking calibration check that refuses to lock a forecast when the gate is not green.
The Kickoff forecast is pinned to a content-addressed snapshot of its inputs and a fixed random seed. Re-executing the frozen package returns byte-identical output on repeated invocation,
python -m vkickoff report --outright --no-live --no-callog \
--frozen-board --json --seed 20260611 --snapshot c7f3f5afa805
which on 12 June 2026 produced a board with SHA-256 digest beginning d9023739 on two consecutive runs, with Spain the champion at 24.0 percent and the same anti-market ordering reported below. The champion board of record in Table 1 is the engine’s 30,000-simulation pre-kickoff run, recorded in the snapshot’s dry-run report; the shippable package reproduces the modal champion and the ranking and agrees with the board of record to within one percentage point on the lower-probability teams, the residual reflecting the simulation count and a later build. The Goldman Sachs forecast carries no such protocol: it is re-run daily inside the bank, and no party outside Goldman Sachs can regenerate its board for any date.
The Goldman Sachs champion probabilities for the top eight teams were verified against the 29 May 2026 note and against reputable secondary coverage (Bloomberg, 2026). The figure for Colombia, which sits outside the note’s published top eight, is taken from the internal head-to-head comparison and is flagged in the limitations rather than asserted as independently confirmed. Goldman Sachs material is used at fair-use scale, the headline champion probabilities, the named covariates, and the stated sample, with full attribution; the note’s exhibits are not reproduced.
The analysis notebook (notebook.ipynb) consumes only the committed CSVs under
data/ and regenerates the deviation table and every Vega-Lite figure with
Python 3.12 and pandas, running top to bottom with no network access. The model
runs occur in the source repositories, the world-cup engine and the poly
package; the notebook consumes the exported artifacts, matching the MKT-004
pattern.
Both models rate Spain the clear favourite for the 2026 title, at 24.0 percent for Kickoff and 25.7 percent for Goldman Sachs, against a devigged-market price of 15.9 percent (Fig. 1, Tab. 1)2. Argentina is the Kickoff runner-up at 16.8 percent and the Goldman Sachs third pick at 14.3 percent, against a market price of 8.5 percent. Because the two models were estimated on different data with different covariates, their agreement on the favourite reflects two implementations of a common recipe arriving at the same answer rather than a shared input.
Table 1. Champion probabilities for the 2026 World Cup: the Kickoff engine, the Goldman Sachs model, and the devigged market, for the nine teams that anchor the comparison.
| Team | Kickoff | Goldman Sachs | Market |
|---|---|---|---|
| Spain | 24.0% | 25.7% | 15.9% |
| Argentina | 16.8% | 14.3% | 8.5% |
| France | 12.0% | 18.9% | 16.6% |
| England | 5.4% | 5.0% | 10.8% |
| Brazil | 4.5% | 7.6% | 8.1% |
| Colombia | 4.2% | 2.0% | 1.7% |
| Portugal | 3.4% | 4.8% | 9.3% |
| Netherlands | 2.9% | 5.0% | 3.8% |
| Germany | 2.7% | 4.5% | 5.5% |
data/champion_marginal.csv.The two models disagree with the betting market in the same direction. Measured as model probability minus the devigged-market probability, both are overweight Spain, at plus 8.1 points for Kickoff and plus 9.8 for Goldman Sachs, and overweight Argentina, at plus 8.3 and plus 5.8, while both are underweight England, at minus 5.4 and minus 5.8, and Portugal, at minus 5.9 and minus 4.5 (Fig. 2)3. Two independently built models leaning against the bookmakers in the same direction is mutual corroboration, and Goldman Sachs reaches the same characterization in its own market comparison, describing itself as overweight Spain and Argentina and underweight England and Portugal (Hatzius, Pierdomenico, and Stehn, 2026). The concentration is near-identical: Spain and Argentina together hold 40.8 percent of the Kickoff mass and 40.0 percent of the Goldman Sachs mass (Tab. 1)4.
data/market_deviation.csv.The largest disagreement is France. Goldman Sachs prices France at 18.9 percent, its second favourite, while Kickoff prices it at 12.0 percent, a gap of 6.9 points (Fig. 1, Tab. 1)4. The two also fall on opposite sides of the market on France: Goldman Sachs sits above the market at plus 2.3 points, Kickoff below it at minus 4.6 (Fig. 2)3. After France, the next-largest disagreements are Brazil, where Goldman Sachs is higher by 3.1 points, and Argentina, where Kickoff is higher by 2.5 points (Tab. 1)4. Colombia is the team Kickoff most favours relative to Goldman Sachs, at 4.2 percent against 2.0 percent.
The Kickoff champion forecast has been backtested against the betting market on four tournaments. Across the 2010, 2014, 2018, and 2022 World Cups, the engine’s pre-tournament champion probability records a mean log-loss of 1.897 against the devigged closing line at 2.199, beating the market in three of the four years and losing only in 2018 (Fig. 3)5. At the per-match level, Dixon-Coles is the better-calibrated primitive, with a held-out log-loss of 0.867 and an expected calibration error of 0.037 on 1,767 matches, against 0.894 and 0.056 for Elo (Tab. 4)6; the champion forecast nonetheless uses Elo, which wins the champion marginal that Dixon-Coles distorts. Goldman Sachs reports a single validation number, a predicted-against-actual goal-difference correlation of 0.49, and publishes no calibration metric or interval (Hatzius, Pierdomenico, and Stehn, 2026). On the specific claim that these champion probabilities are well-calibrated, Kickoff supplies evidence and the public Goldman Sachs note does not.
Table 4. Validation evidence published by each system. The Kickoff per-match and backtest figures are computed against held-out data and the devigged closing line; the Goldman Sachs figure is the goal-difference correlation reported in the note.
| Measure | Goldman Sachs | Kickoff |
|---|---|---|
| Primary metric | Goal-difference correlation 0.49 | Brier, log-loss, and ECE versus history and the closing line |
| Per-match log-loss | Not published | 0.867 (Dixon-Coles), 0.894 (Elo), n = 1,767 |
| Champion backtest | Qualitative 2022 retro-cast | Log-loss 1.897 versus market 2.199, beats market 3 of 4 |
| Uncertainty | Point probabilities, no interval | Interval on the headline, Monte-Carlo standard-error gate |
| Acceptance rule | None stated | Blocking calibration gate |
data/champion_backtest.csv.The construction ledger is where the asymmetry shows (Fig. 4, Tab. 5)7. The Kickoff forecast was built in nine days, from the first commit on 2 June 2026 to the frozen pre-kickoff snapshot on 11 June 2026, across about 16,000 lines of engine code, a 5,300-line frozen forecasting package, and 6,800 lines of tests, over 58 commits. Its data cost was zero: the martj42 international-results file and the eloratings.net ratings are free (Jürisoo, 2024). It was built by one operator working with AI coding agents. The Goldman Sachs note states only what is public: three named Global Investment Research authors, a proprietary-data footing, and a research product with a lineage to the 2014 and 2018 editions (Hatzius, Pierdomenico, and Stehn, 2026; Bloomberg, 2026). The contrast does not require any estimate of Goldman Sachs headcount, hours, or budget to be stark.
Table 5. Construction ledger for the two forecasters. The Kickoff column is derived from the world-cup and poly git histories and the frozen snapshot date; the Goldman Sachs column states only public facts.
| Dimension | Goldman Sachs | Kickoff |
|---|---|---|
| Builders | Three Global Investment Research authors | One operator with AI coding agents |
| Build span | Recurring product, lineage to 2014 and 2018 | Nine days, 2 June to 11 June 2026 |
| Codebase | Proprietary, not public | 16,000 lines engine, 5,300 package, 6,800 tests; 58 commits |
| Data cost | Proprietary data footing | Zero, free public sources |
| Reproducibility | Daily re-run, not externally reproducible | Deterministic from snapshot and seed |
| Validation | Goal-difference correlation, no interval | Brier, log-loss, ECE behind a blocking gate |
Scored only as champion-marginal forecasters, the two systems finish even, with the edges falling on different dimensions (Tab. 6)8. Goldman Sachs holds the edge on estimation sample and statistical power, on model inference and interpretability, and on covariate richness. Kickoff holds the edge on calibration evidence, on market-benchmark rigor, on leakage control and reproducibility, and on honest statement of limits. Goldman Sachs leads three dimensions and Kickoff leads four, but the dimensions are orthogonal goods, estimation breadth and calibration discipline are different things, and they do not sum into a single score. The honest reading is a tie: Goldman Sachs is the better-estimated model and Kickoff is the better-validated one.
Table 6. Seven-dimension assessment of the two systems as champion-marginal forecasters, scored one to five per dimension with the edge attributed.
| Dimension | Goldman Sachs | Kickoff | Edge |
|---|---|---|---|
| Estimation sample and statistical power | 5 | 3 | Goldman Sachs |
| Model inference and interpretability | 4 | 3 | Goldman Sachs |
| Covariate richness | 5 | 3 | Goldman Sachs |
| Calibration evidence | 2 | 5 | Kickoff |
| Market-benchmark rigor | 3 | 5 | Kickoff |
| Leakage control and reproducibility | 2 | 5 | Kickoff |
| Honest statement of limits | 4 | 5 | Kickoff |
France is the clean referendum on what a richer covariate set buys. The Goldman Sachs scoring-talent term rewards France’s unusually deep roster of elite top-five-league forwards, and it lifts France to 18.9 percent and above the market (Fig. 2). The Kickoff champion engine carries no such term and prices France near its rating-implied level, below the market. Whether the covariate captures real attacking quality that Elo misses or whether it overfits a plausible story is open, and only the tournament can adjudicate it. The France gap isolates the methodological difference between the two systems more cleanly than any other team.
That difference is breadth against discipline. Goldman Sachs invests in breadth of explanation: more covariates, a larger sample, interpretable marginal effects, and a daily re-run that incorporates new information. Kickoff invests in discipline of evaluation: a minimal covariate set, a calibration gate that can veto a feature or a forecast, a frozen snapshot, and a seed that makes the result reproducible. Each choice has a cost. The Goldman Sachs covariates rest in places on small samples without published intervals, and the daily re-run is not a frozen, leakage-controlled protocol. The Kickoff parsimony may leave real signal on the table, as the France case suggests, and its champion forecast uses the less inferential of its two engines.
What the comparison establishes sits above either model. A forecasting system that one operator assembled in nine days with AI coding agents, on data that cost nothing, holds an even side-by-side with a model from a top research division that has published on this problem since 2014. The two agree on the champion, agree on the direction of their disagreement with the market, and divide a seven-dimension assessment three edges to four. For a practitioner choosing between them, the choice is between explanatory richness and evaluative rigor: the Goldman Sachs model is the one to read for an account of why a team is strong, and the Kickoff model is the one to trust for a probability that has been calibration-tested and can be reproduced.
The deepest limitation is shared and structural. The 48-team format has no prior instance, so neither model’s bracket-path distribution, the third-place combination table and the eight-match finalist path, can be validated on history; both are at best market-anchored and internally consistent. The comparison is therefore confined to the champion marginal, the one axis both models answer.
The Goldman Sachs side carries the limitations of an unvalidated small-sample design. Several covariates rest on few observations, the reigning-champion slump on roughly a dozen prior champions, without reported regularization, standard errors, or out-of-sample stability tests, and the note publishes no calibration metric or confidence interval. These are characterizations of the public note, not of any internal Goldman Sachs work that may exist but is not published.
The Kickoff side has a thinner evidence base than Goldman Sachs’s proprietary inputs, a deliberately narrow feature set, and a champion forecast that uses the Elo engine rather than the more inferential Dixon-Coles. Its published board of record is a 30,000-simulation run; the shippable package reproduces the ranking and Spain at 24.0 percent but differs by under one percentage point on lower-probability teams, a build and simulation-count residual disclosed in Section 4.3.
Two caveats apply to the head-to-head numbers themselves. The two forecasts were frozen on slightly different dates and market snapshots, which injects a small non-model wedge into any gap between them. And the Goldman Sachs figure for Colombia, outside the note’s published top eight, comes from the internal head-to-head transcription rather than from independent verification.
This comparison is pre-kickoff by construction: both published boards predate the opening match, and neither side’s probabilities are updated here with realized results. Scoring the two forecasts against what actually happens is the job of the in-tournament scorecard, the planned MKT-004 sequel, and is out of scope here.
On the one question both models answer, who wins the 2026 World Cup, Goldman Sachs and Kickoff are peers. They share an Elo-Poisson-Monte-Carlo backbone, name the same modal champion in Spain, lean against the bookmakers in the same direction, and divide a seven-dimension assessment three edges to four. Goldman Sachs is the better-estimated model and Kickoff is the better-validated one, and the choice between them is a choice between explanatory richness and evaluative rigor rather than a choice between right and wrong.
What remains open is the part neither model can settle from the armchair. France, where the two disagree most, will be priced correctly by one of them and not the other only once the matches are played, and the broader accuracy question waits on the tournament. The result that does not wait is the one this paper set out to record: the price of building a forecaster that holds its own against Goldman Sachs has fallen to one operator, a fleet of coding agents, nine days, and free data.
All datasets are released under CC BY 4.0 and inventoried in
data/manifest.yaml. The analysis is reproduced by
notebook.ipynb, which loads the committed CSVs and
regenerates the deviation table and every Vega-Lite figure with no network
access. The datasets are: champion-marginal (the head-to-head champion
probabilities of Kickoff, Goldman Sachs, and the devigged market);
market-deviation (each model’s deviation from the market baseline, rebuilt by
the notebook); champion-backtest (the Kickoff engine against the closing line
for 2010 to 2022); match-calibration (held-out per-match Brier, log-loss, and
ECE for the two engines); assessment-rubric (the seven-dimension assessment);
and build-ledger (the construction contrast). The Kickoff and devigged-market
figures derive from the frozen world-cup engine snapshot c7f3f5afa805
(reports/2026_dry_run.md, engine_choice.md, market_comparison.md,
backtests.md); the reproduction command is in Section 4.3. The Goldman Sachs
figures are from the 29 May 2026 Global Investment Research note, which is the
property of Goldman Sachs and is not redistributed here; readers who wish to
consult it should acquire it from Goldman Sachs or its reported coverage. The
free input data are the martj42 international-results file (Jürisoo, 2024) and the
eloratings.net ratings.
See references.yaml. Inline citations resolve there: Dixon
and Coles (1997) and Elo (1978) on the forecasting primitives, Brier (1950) on
proper scoring, Hatzius, Pierdomenico, and Stehn (2026) and Bloomberg (2026) on
the Goldman Sachs model, Jürisoo (2024) on the free results data, eloratings.net
on the Elo inputs, and ATOL Research (2026) for the Kickoff engine in MKT-004 and
its construction in MKT-005.
The held-out champion log-loss of the two engines is in data/engine_selection.csv, loaded in the notebook load cell and transcribed from the world-cup engine_choice report. ↩
Champion probabilities are in data/champion_marginal.csv, loaded in the notebook load cell and plotted by the fig-champion-marginal cell. ↩
Deviations from the market baseline are computed in the notebook tilt cell into data/market_deviation.csv and plotted by the fig-market-deviation cell. ↩ ↩2
The notebook summary cell computes the Spain plus Argentina concentration, the France gap, and the rubric edge split from the loaded tables. ↩ ↩2 ↩3
The champion backtest is in data/champion_backtest.csv; the summary cell computes the mean log-loss and the three-of-four win count, and the fig-champion-backtest cell plots it. ↩
Per-match calibration figures are in data/match_calibration.csv, loaded in the notebook load cell. ↩
The construction ledger is in data/build_ledger.csv, derived from the world-cup and poly git histories and the frozen snapshot date. ↩
The seven-dimension assessment is in data/assessment_rubric.csv. ↩
research/markets/tournament-forecasting/MKT-006-gs-comparison/.
Manifest:
data/manifest.yaml.
Notebook:
notebook.ipynb.