Tennis fans argue about who will win a slam. This project does something slightly different: it gives each player a percentage chance, the way a weather forecast gives a chance of rain. “Alcaraz 27%” does not mean he will win. It means that if this tournament could be played many times, he would win about one in four of them.
The work used the real results of the 2026 men’s season — thousands of matches, from the biggest events down to the smaller Challenger tournaments. From those results, every player got a strength score. That score was used to say “if these two play today, how likely is each to win?” Then those match chances were used to play out the whole US Open draw on a computer, many times, and count how often each man lifted the trophy.
To check whether those percentages were honest, they were compared with two things people already trust: bookmakers’ odds, and what actually happened at the 45 ATP tournaments already finished in 2026. For that second check, the computer was only allowed to use matches from before each event began — no peeking at the winner.
First, give every player a score. Think of it like a ladder. Beat a player above you and you climb a lot. Beat someone below you and you barely move. Lose to someone below you and you drop. That score is called Elo. Chess has used it for decades; here it is applied to tennis. A player who has not played much yet is not left at a blank: they start from their official ranking points, which already say something about how good they are. Importantly, a player’s score before a match uses only matches that have already finished. The computer never “knows” the result of the match it is trying to forecast.
Second, turn two scores into a match chance. Subtract one player’s score from the other’s. A gap of 100 points means the stronger player wins about 63 times out of 100 — a favourite, but far from certain. Other facts were tried too: how someone had been playing lately, how many days of rest they had, their record against that opponent, serve and return numbers, how many matches they had played in the last two weeks. None of those extras did better than the simple score gap, so a match is decided from that one comparison.
Third, play the tournament — not just one match. Winning the US Open takes seven wins in a row. And the draw is a tree: who you meet in the quarter-finals depends on who won the rounds before. So you cannot just look at Alcaraz’s first-round opponent and stop there. The computer plays the entire bracket 40,000 times. Each time, every match is decided by a weighted coin flip using the chances from step two. In some runs Alcaraz wins. In some he loses in the fourth round. After 40,000 runs, if he has the trophy in 26.8% of them, that is his title chance.
The sections below show those chances, then the extra facts that were tested, then whether the percentages lined up with real results.
Carlos Alcaraz is the model favourite at 26.8%. In a 128-player knockout, even the strongest player in the field is more likely to lose the tournament than to win it.
The numbers below come from a match-win model fitted on the 2026 ATP season, then run through the US Open draw 40,000 times. Jannik Sinner, who won six titles including Wimbledon, is not in the field — his last match in the data is 12 July — so this is a forecast of the draw as entered, not of who is best in the world.
Left: share of 40,000 simulated tournaments won by each player. Right: the same runs, read as the chance of reaching each round. Alcaraz is a near-lock through the third round and still only 27% to win the title.
| Player | P(title) | Fair odds |
| Carlos Alcaraz | 26.8% | 3.7 |
| Alexander Zverev | 13.3% | 7.5 |
| Ben Shelton | 7.5% | 13.3 |
| Frances Tiafoe | 6.0% | 16.8 |
| Taylor Fritz | 4.9% | 20.4 |
| Tommy Paul | 4.6% | 21.6 |
| Jiri Lehecka | 3.6% | 27.6 |
| Daniil Medvedev | 3.3% | 30.1 |
The top 20 players hold 93% of the probability mass. Across 40,000 simulated tournaments, 45 different men still won at least once.
11,789 decided ATP matches from 2 January to 5 September 2026: Tour and Challenger, hard / clay / grass. Features for match i are built only from matches that finished before it. Post-match box-score columns (sets won, aces in this match) never enter the prediction for that match.
Most players have almost no history — median six matches — which is why new names start from ATP ranking points rather than from a blank rating. Challenger matches outnumber Tour matches roughly 2.5 to 1; the shipped model is scored on Tour only.
Every input is a difference: player A minus player B. That keeps the model fair if the two names are swapped.
Fifteen variables were built and tested:
| Group | Variable | What it measures |
| Rating | Elo difference | Strength from prior results. New players start from ATP ranking points. |
| Matches in the rating | How much evidence sits behind that Elo. | |
| Activity | Career matches this season | Volume. |
| Days of rest | Time since last match (capped at 60). | |
| Recent form | Win rate over the last 10 matches. | |
| Head-to-head | Prior meetings | A’s win rate vs B, shrunk toward 50% when the sample is thin. |
| Serve / return | Service-points won % | Rolling average of A’s own serve, last 20 matches. |
| Return-points won % | Same for return. | |
| Break points saved % | Serve under pressure. | |
| Aces, double faults | Per-match averages. | |
| Fatigue | Sets, games, matches in 14 days | Workload going into the match. |
Elo is the only variable with real standalone signal. Serve and return stats come next, then sit on top of Elo — they describe the same thing twice.
Left: Elo difference dominates mutual information with the match result. Right: fatigue counts move together; rest is the inverse of recent workload. Adding correlated copies of the same story is how a 15-feature model gets worse.
A 15-feature logistic model and a gradient-boosted model both lost to Elo alone:
Each group is added on top of the last. Serve/return dips the error a hair, still inside fold-to-fold noise. The full set is worse than Elo by itself.
What actually prices a match is therefore one number: the Elo gap. Converted with a fitted coefficient (β = 0.00528):
P(A beats B) = sigmoid(0.00528 × (Elo_A − Elo_B))
A 100-point gap is a 62.9% favourite. Surface-specific ratings were tried; they did not improve the forecast enough to ship.
The title probability is not that formula applied seven times to a fixed path. The draw is a tree: who you meet in the quarters depends on who won the rounds before. The model plays the whole bracket 40,000 times (4,000 per event in the season backtest), sampling each match from the formula above, and counts trophies. Grand Slam matches are best-of-five; that longer format gives the better player more room to win.
The same match probabilities can be written onto a completed draw. Green is a result the model expected before the event started; red is an upset. Wimbledon 2026, Sinner’s title path:
Fery over Cobolli at 0.29 is the shock in this slice. Sinner’s last three matches were all priced above 0.65 — a title the ratings saw coming.
Benchmark: betting-market odds with the bookmaker margin removed. Test set: 254 ATP Tour matches, 1–22 August, locked until the end.
| Log loss (lower better) | Accuracy | |
| This model (Elo only) | 0.645 | 62.2% |
| Betting market | 0.627 | 63.8% |
The model does not beat the market. Gradient boosting on all 15 features was the worst of the ladder, not the best. The ranking of models is stable across months; the gap to the market does not close.
Red is the ceiling. Everything else is a results-only model. Distilling the market into a ridge on the same features still sits ~0.04 nats behind — the missing information is not in this file.
The gap is widest on close matches — the ones decided by fitness, motivation, and other things that never appear in a results file.
Left: predicted probabilities track observed win rates. Middle: Elo is closest to the market when the rank gap is huge, and loses most of the ground inside the top 25. Right: the pattern holds on clay, grass, and hard.
On all 45 completed ATP events of 2026, ratings frozen at each event’s start:
Left: most events sit below the uniform line — the bracket simulation beats “everyone is equal.” Middle: predicted favourite rate 21.6%, observed 20.0%. Right: the Australian Open is the left tail: Alcaraz, the champion, was given 0.05% because January ratings had almost no 2026 evidence.
An early match model hit 93% accuracy by reading who won more sets. That is leakage, not skill. It was thrown out.
Github Source:
https://github.com/grstanziola/USOpen2k26Prediction/blob/main/US_Open_2026_Prediction.ipynb