After reviewing how well the model fit with the actual 2006 season and estimating the luck factor, I noticed that teams with strong run defenses were not well represented in the playoffs. This result is contrary to most conventional wisdom that emphasizes "stopping the run" as a key to winning.
The conventional wisdom makes sense. In a purely logical sense if a team were infinitely bad at stopping the run, its opponents would score a touchdown on every run attempt. So every incremental improvement over an "infinitely bad" run defense would have to improve a team's chance of winning.
But look at a 2006 ranking of run defense. The playoff teams are highlighted in yellow. The horizontal line is at the league average.
Notice that only 5 of the 12 playoff teams were above average while the other 7 are below average. Only 1 playoff team was in the top 7 in run defense. Additionally, 4 of the worst 6 run defenses made the playoffs including the absolute worst 2. Incredibly, the 2006 Super Bowl winner was the very worst--and not just by a little but by .39 yds/run, almost an entire standard deviation worse than the next best team.
To illustrate this another way, I've plotted wins against defensive run efficiency and added a regression line. Notice that as run defense gets worse, wins increase. This is backwards and obviously indicates a big problem in the data. In fact, the simple linear regression illustrated above indicates that defensive run efficiency is not significant at all (p = 0.839). This is partially because of the small population of NFL teams, n=32.
This raises a larger problem with the model and its baseline data. The model's coefficients were drawn from the 2005 season only. In 2005, winning and run defense were correlated much better than 2006.
There were 256 games in 2005, which means there were 512 "game efforts" by all the teams. The model's significance calculations are based on the sample size, and n=512 seems like a healthy sample at first glance. But although there were 512 "game efforts" there only 32 different teams which creats a hybrid dataset in which n=512 and n=32.
The model can definitely be improved by adding more seasons of data to calculate the coefficients.
- Home All posts
Importance of Run Defense
Predicting Playoff Races
I'll use Baltimore as an example again, because I follow the Ravens most closely. Following week 12, the Ravens were 9-2 and it was all but certain they would win the AFC North. But they were also in the hunt for a first-round bye in the playoffs. Indianapolis held the second seed and a bye with a 10-1 record. So to earn a bye, BAL had to at least beat out IND.
Since the model had calculated, based on a game-by-game assessment, the probabilities of each team finishing the season with each possible total number of wins, they could be compared side-by-side. Here is how both teams' outlooks appeared following week 12.
Prob(BAL 12, IND 12) = Prob(BAL 12) * Prob(IND 12) = .339 * .260 = .088
So there is about a 9% chance of that particular outcome. Every possible combination of outcomes is listed in the table below. BAL win totals are listed vertically on the left and IND win totals are listed horizontally on top. (Click to enlarge.)
P(BAL wins) = .279
P(Tie) = .278
P(IND wins) = .443
(Total = 1.000)
The table above is easier to understand if pictured graphically. Below is a plot of the probabilities of each outcome. BAL's wins are along the vertical axis and IND's are along the horizontal axis. The diagonal line plots the tie outcomes, where both teams finish with an equal number of wins. Plot "density" above the line indicates IND would have more wins, and "density" below the line indicates BAL would have more wins. Here, we see the "center of gravity" of the probability plot shows the mostly likely outcome is a tie at 13 wins for both teams. (And except for IND laying an egg at home vs. HOU, this would have been the real outcome.)
You can download the spreadsheet with several selected match-ups between playoff contenders here.
Assessing the Model's Accuracy
Probability models are difficult to assess by their nature. Linear models offer an R-squared that give a definitive assessment of the explanatory/predicitive power of a model. But probility models, such as the logit model I use, offer numerous indirect assessments but none is more straightforward the the % correct score. It tells us how well our model predicts actual outcomes.
If the model predicts outcomes well, then we know two things. First, we know how to predict games, which is fun. Second, we understand what is really important in winning NFL games and we have a deeper understanding of the inner-workings of the sport as it is played.
But % correct doesn't tell the whole story. It simply draws a line at .50 and if a team that is predicted to have a .51 win probability actually wins, we consider the model correct. If a team predicted to have a .49 win probability wins, the model is considered incorrect. But that's unfair. We expect to be wrong in 49% of all such cases. That's just the reality of equally matched teams. Further, the model is expected to be wrong in 20% of the games where it predicted the favorite team to have a .80 win probability. In fact, if the model were exactly 20% wrong in such games, it would mean the model is better than if it were 100% correct.
So to assess the model's accuracy, not just in terms of how often its predicted favorite actually won, but in terms of how accurate the predicted probabilities were, I produced the table below. I divided all the games into 5 categories based on the "lopsidedness" of the probability. Where the visiting team was forecast to have a win probability of between .00 and .20 (and the home team's win probability was between 1.00 and .80), I scored the models accuracy. I did the same for win probabilities between .21 and .40, between .41 and .60, between .61 and .80, and finally between .81 and 1.00.
If the model fits relatively well, its % correct scores should reflect the predicted probabilities of each category. So, for the 1st category, between .00 and .20, we'd expect a % correct score of approximately 90%. Here is how all 5 categories scored:
The model seems to fit well accross the spectrum of games. Locks, solid favorites, and toss-ups were accurately predicted by the model.
Total Probabilities
In the final stretch of the 2006 regular season I wanted to know which teams would make the playoffs. I thought that as long as my model was somewhat valid, it would give insight on how various teams would finish the season.
As opposed to my original linear model, which predicted the number of total season wins for each team based on to-date efficiency, my new game-by-game model takes matchups and homefield advantage into account. As the number of remaining games dwindle, these factors become more important.
Baltimore is my favorite team, so it was the first one I analyzed. With 5 games remaining in the season, the probability models (both the original (model 1) and alternate (model 2)) showed the following win probabilities for the Ravens.
The probabilities for each possible season outcome, i.e. total number of wins can be computed. For example, the probability that the Ravens would end the season by winning all 5 remaining games would be the product of all the individual win probabilities. At this point in the season it was:
Prob(5 wins) = .50 * .48 * .95 * .70 * .91 = .14
Every possible combination of game outcomes was calculated. There were 2^5 possible combinations of outcomes with 5 games remaining. Each combination that leads to 4 wins is summed to give the total probability of winning 4 of the 5 remaining games. The same is done for 3, 2, 1, and 0 wins. With a record of 9-2, the Ravens' resulting probabilities were:
The cumulative probabilities of winning "at least" X number of games is easily calculated by adding the probabilities of all the possible outcomes of X and greater than X. Baltimore's probabilities of winning at least X number of games was:
As luck would have it, Baltimore indeed finished the season with 13 wins, the number predicted most likely by the model.
These calculations were performed for all 32 teams. Once computed, various teams of interest can be compared. For example, who would win the AFC North? Simply compare BAL's and CIN's total win probabilities. Who would likely win homefield advantage in the playoffs? Compare BAL's and IND's total win predictions.
An Alternate Model
Midway through the 2006 regular season, I created an alternate game-by-game winner prediction model. It used the same stats, such as yards per rush, yards per pass attempt, turnovers, and home field advantage, but it used them in a different way.
I still used logit regression to compute the outcome probability of each game, but I experimented with different forms of each independent variable to see if I could improve the model's fit. I tried exponential and logarithmic versions of each variable, but predictive power was not improved.
I finally stumbled on an idea that produced a version of the model at least as predicitive as my original. I wondered if I could model the effect of a very strong running offense against a very weak run defense (and for passing, and vice versa). I theorized that if I could mathematically represent such an interaction, it might fit reality better than considering each team's efficiency stat alone. After all, this how it works in actual games--one team's offense interacts with its opponents defense. They're not performing independently in front of judges or for time.
Instead of using each team's yards per pass attemt/run, etc., I created variables that captured the interaction of an offense vs. a defense which I called PASSFACTOR AND RUNFACTOR. When Team A plays Team B, this is represented mathematically by:
APASSFACTOR = AOPASS * BDPASS (Team A's off pass eff x Team B's def pass eff)
ARUNFACTOR = AORUN * BDRUN (Team A's off run eff x Team B's def run eff)
BPASSFACTOR AND BRUNFACTOR are computed in the same way.
In this way, if a team with a great pass offense plays a team with a poor pass defense, the mismatch will be captured because PASSFACTOR will be very high for the great passing team.
The net turnovers and home field advantage variables remain the same. However, with a better database, it would be interesting to use the same technique with team A's giveaways vs. team B's takeaways and vice versa, or even going deeper by discerning fumbles and interceptions as separate variables. The random components of turnover stats may become too strong if we divide them up like that.
The new "matchup" model was nearly as predictive as the original, correctly predicting the outcome of only one fewer the 2005 games as well as the original (74.6% correct). All variables were significant. This was no breakthrough, obviously, but it is another tool to understand the game. And by adding more seasons of data, perhaps the model will improve.
Week by week, both models produced very similar probabilities, often only differing within .03 or so.
Contact / FAQ
You can contact Brian by email at brian@mail.advancednflstats.com. Questions about the content of the posts are best as comments. All comments get read, and those with questions often get a quick response.
I'm always amazed at the overwhelming response this site has received. I appreciate all the feedback. Please keep it coming--the good and the "constructive." While I don't want to discourage anyone from writing me, I'd like to encourage everyone to use the site comments feature for any items you think might be of interest to other readers. You may get a answer to your question quicker from another reader than from me. Also, my spam filter can trap some comments and emails, so I apologize if you haven't received a response.
FAQ
Where did you get your data?
Most of my team data comes from open online sources such as espn.com, nfl.com, myway.com, and yahoo.com. It's easy for anyone to grab whatever they're interested in from those sites.
My play-by-play data comes from a source that's not publicly available, and at this time I regret that I cannot share it. However, I have been able to create a database of nearly all NFL plays since 2002 free for anyone to use for research.
Why don't you do predictions Against The Spread (ATS)?
While I don't object to gambling, it doesn't really interest me. I realize a lot of football fans are interested in it, and I'm happy if anything here helps inform fans of all kinds. I do however use Vegas odds and spreads as consensus yardsticks for my own models, and sometimes it's fun to see if some math can beat "the system." I just lack that gene for risk, and cheering for my favorite team is excitement enough for me.
I heard you're not a math professor or professional statistician, but actually a former Navy fighter pilot. Is that true?
Yes. I flew F/A-18 Hornets off carriers. The Navy sent me to graduate school where I picked up a knack for analytics. Only now do I realize that much of the combat tactics we used as pilots were derived using the same analytical techniques I use for my football research.
Why aren't you working for a team?
I'm very reluctant to relocate my family. I often do freelance consulting for NFL teams and media outlets.
What's it like flying an F/A-18?
The best way I ever heard someone describe flying a fighter plane was this: It's like simultaneously racing in the Indianapolis 500 while playing a video game and holding two telephone conversations...all while people are shooting at you. Then the hard part begins--landing on a pitching carrier deck. In reality flying a plane is a true team sport. Thousands of brave sailors and everyone else who supports them. I was very fortunate to be part of such a great organization as the United States Navy.
About the Founder
Brian Burke
Reston, Virginia
Brian is a football fan and closet math enthusiast. He has a BS in aerospace engineering and an MS in management and leadership. After spending 15 years in the Navy, most of them as an F/A-18 carrier pilot, Brian has taken up the less dangerous hobby of advanced NFL statistics.
Originally from Baltimore, Brian spent his youth either on the sidelines of Maryland Terrapin football games or in 50-yard line seats at Colts games, both thanks to his dad. His football career topped out playing tight end for the Towson High Generals. His only college football experience was playing in the Naval Academy's annual "2.7" game, an intramural game that featured the "geeks" vs "rocks" divided by the Academy's average GPA. You can guess which team Brian played for.
After graduating from Annapolis in the Class of '93, Brian attended flight school and went on to fly the F/A-18C Hornet until leaving the Navy. He attended the Naval Postgraduate School and returned to Annapolis as an instructor. Brian flew 21 combat missions and was awarded the Air Medal, the Navy Commendation Medal, and numerous other personal and unit commendations.
He now resides with his family in northern Virginia.
Is There 'Momentum' During a Season?
The 2006 season continued and I made incremental improvements to the model. Beginning at week 8, I emphasized the efficiency stats from each team's most recent 4 weeks. I hoped this would help account for a team's improvement over the course of the season and for injuries to key players. I set this to be a changable variable which I called an "emphasis factor" that I could set to 0 to see how much the predicted probabilities changed from the baseline.
What I found was that the probabilities changed, but not by much, and rarely enough to swing one team from underdog to favorite. It made intuitive sense to me, and still does, so I kept the "emphasis factor" in the model.
But this is actually a question ripe for research. Teams appear streaky. Commentators, analysts, and fans accept that teams can be on a roll or in a rut. But over the course of a season it would be reasonable that wins and losses can naturally come in bunches on occasion.
Just as if you flipped a coin 16 times, you wouldn't expect the outcome to perfectly alternate heads and tails. If you did this several times, you would find that some of the 16-coin-flip "seasons" have several instances of consecutive heads "streaks." If we were betting on heads, our human intuition would tell us that the coin is "hot" or on a hot streak. In some rare, but totally natural, cases we might even expect 8 straight heads followed by 8 straight tails, or vice-versa.
The same principal holds for sports. The 16-game NFL season is particularly short and is susceptable to natural "streaks" without any actual change in performance. In theory, an average team destined to go 8-8 could start the season 8-0 followed by an 0-8 "collapse." Admittedly this would be extremely rare. In this case the luck factor (calculated below) for this team would be +4 during the 1st half of the season, and -4 for the last half of the season.
This phenomenon may explain why the emphasis on recent performance I applied in the model didn't change the probabilities very much. Teams appear streaky in terms of wins and losses, but win/loss records are naturally more erratic than the efficiency stats, so we should expect teams to appear that they're on "a roll" or in "a rut" when winning or losing streaks are really just part of a natural distribution of outcomes.
Below is a table of Week 11 comparing the predicted probabilities using both the standard stats and the stats with the recent 4-weeks performance weighted.