What each model knows

Each league has a model of its own. They share a method, and differ in what they know about the game.

What every model reads

  • A rating of each team, built from its results: it moves after every game, more after a wide margin, counts the home advantage, and is pulled part of the way back to the average between seasons.
  • Rest: the days since each team's last game.
  • Form: each team's last ten games, by wins and by average margin.
  • The season so far: each team's share of games won.
  • Whether the game is played on a neutral site.

A model reads a team as it stood before the game. Nothing from the game itself, or from later, can reach its own prediction.

What each league adds, and what it does not see yet

  • MLB: the starting pitcher once he is announced — the runs he allows, a measure of his own pitching that leaves the fielders out, and how deep he goes into games. Not yet: bullpen usage, lineups, the park and the weather.
  • NFL: who starts at quarterback, and what he has been worth per start, adjusted for the defenses he faced. Not yet: other injuries, and the weather.
  • NCAA football: the talent on the roster and the production that returns from last season, passing included. They weigh most in a season's first games. Not yet: injuries, changes of depth chart during the season, and the weather.
  • NBA: who is out, from the injury report. Not yet: lineup combinations, travel, and rest beyond back-to-backs.
  • WNBA: what every model reads, and nothing more yet. Not yet: who is available, and roster changes.
  • NHL: shot volume, back-to-backs, how many games in the days before, and the distance travelled. Not yet: the starting goalie and special teams.

The same “not yet” is shown on Predictions when one league is picked.

How a model is trained and judged

Time only moves forward. A model is fitted on past seasons, its probabilities are calibrated on the season after those, and it is judged on seasons after that — games it had never seen. The latest season is looked at once, after every choice has been made.

In each league three kinds of model compete: the team rating alone, a logistic regression and gradient boosting. The simplest stays unless another does measurably better on unseen games: complexity has to pay for itself.

How a model keeps learning

Every week that a league has enough new games, a challenger is trained on everything up to the first half of them. The model in use and the challenger then predict the second half, which neither has seen, and the challenger takes over only if it is not worse. Off season, a league simply waits.

The margin and the total

A second model per league estimates the margin and the total from the same inputs plus each team's recent scoring, for and against. The ± beside them is measured on a season the model had not seen.

What the models do not do

Measured on past seasons, the models are well calibrated and do not beat outside estimates of the same games: where the two disagree, the model is more often the one that is wrong. iPick Sport says so instead of hiding it — that is what reliability is for, and the Results page shows every day how the predictions did.