Who Gets Credit for YAC?

In the last post, I discussed Yards After Catch (YAC) and introduced its complement, Air Yards. In this post, I'll take a closer look at YAC and how much credit a QB should get for YAC compared to his receivers. Ultimately, these results will contribute to an improved passer rating system, an effort I began here.

When we break down pass yards and how credit should be shared by the passer and receiver, we could start by saying that the receiver should get credit for the Yards After Catch and the QB should get credit for the Air Yards.

But a plausible alternative is that the QB's accuracy contributes to YAC. His accuracy, and his ability to read defenses and find open receivers, would logically allow him to contribute significantly to YAC. An accurate and aware QB would be able to lead his receivers, hitting them in stride and steering them to where they can best evade a tackle.

If a QB is responsible for some of his team's YAC, his accuracy and ability to read open receivers should correlate well with YAC. To test the theory that a QB's abilities contribute to YAC, I conducted a simple regression. The correlation of accuracy and YAC should be positive, strong, and significant. Additionally, the r-squared of the model should be relatively high.

As measures of a QB's accuracy and ability to read open receivers, I used pass completion percentage and interception rate. The dependent variable was YAC/completion. The data comes from every team between the 2002 and 2006 regular seasons (n=160). The results of the regression are listed below.













































































VARIABLECOEFFICIENTSTDERRORT STATP-VALUE
Comp %0.010.010.660.51
Int Rate-4.696.02-0.780.44
r-squared0.01




Neither completion percentage (p=0.51) nor interception rate (p=0.44) was statistically significant. The r-squared for the model was nearly zero. These results yield evidence that, for passes as a whole, the QB makes little contribution to YAC. Yards after Catch, therefore, belong mostly to the receiver.

This result is important because it means a valid QB passer rating should exclude YAC. In the next post, I'll look at QB performances from 2006 and who were the downfield threats and who were the "dink and dunk" throwers. Ultimately, I'll apply what I learned about Air Yards and YAC to an improved passer rating formula.

Leading Indicators 3

In the last two posts I compared the results of two regression models. The first model estimated current year wins based on current year stats. The second model predicted next year’s wins based on last year’s stats. The comparison of the regression results revealed how well various team stats persist from year to year as leading indicators of future team wins.

In this post, I'll apply the leading indicator model to each team's 2006 stats. This model accounts for only 20% of the variance in season win totals (compared to 75% for the standard current year model), but it does tell us which teams may have a "head start" due to their previous performance.

The strongest leading indicators were found to be offensive running efficiency (36%), team penalties (52%), and defensive forced fumble rate (45%). Offensive and defensive interception rates tend to be "anti-predictors" in that they are strong leading indicators, but good performance one year predicts fewer wins the next year.

Here is the list of each team's "leading indicator wins." These are win estimates above or below average calculated from the results of the leading indicator model.



































TEAMPersisting Wins
DAL2.4
MIA2.1
DEN2.0
CHI2.0
PIT1.7
KC1.6
ATL1.6
IND1.6
SD1.4
PHI1.2
NO1.1
SF0.8
SEA0.7
NYG0.5
GB0.1
STL0.0
JAX-0.2
CIN-0.3
DET-0.3
NE-0.5
CAR-0.5
ARI-0.5
NYJ-0.5
TB-0.6
TEN-0.6
BUF-0.8
MIN-0.9
WAS-1.0
CLE-1.1
BAL-1.2
HOU-1.4
OAK-2.4


The teams that fare well here are those that ran the ball well, committed few penalties, and forced a lot of fumbles. They also had poor interception rates on offense, defense, or both.

The only anomoly are the 13-win Ravens, who are buried among the league's worst teams. This is due to their phenomenal interception rates on both offense and defense in 2006. It will be very hard to repeat their 13 wins in 2007.

Leading Indicators 2

"Anti-Predictors"

In the last post I compared the results of two regression models. The first model estimated current year wins based on current year stats. The second model predicted next year’s wins based on last year’s stats. The comparison of the regression results revealed how well various team stats persist from year to year as predictors of team wins.

I found that the stats that persisted from season to season as a predictor of wins were offensive running efficiency (36%), team penalties (52%), and defensive forced fumble rates (45%). I also found two stats that could be considered anti-predictors. Offensive and defensive interception rates both reverse their direction of prediction across seasons. In other words, a low offensive pass interception rate for a team in one year foretells fewer wins the following year, all other things being equal.

This is a confusing result, to say the least. One possible explanation is random coincidence--the stats may show a connection only by chance. However, the significance of the offensive interception coefficient was 0.07 and the defensive interception coefficient was 0.09. As I wrote earlier, the FDA might not approve heart medication based on trials with marginal significance levels, but it is still highly unlikely that both coefficients suffer from statistical Type I errors. After all, we’re not splitting the atom here. We’re just talking about football.

In the last post I wrote, “My theory is that we are witnessing regression to the mean. For many teams, interception rates have a lot of variance due to luck. So a team that is unlucky with interceptions one year is not likely to be as unlucky the next year, [and vice versa]. That could partially explain the reversed signs. Another possibility is that teams systematically swing from high to low interception rates from one season to the next, something I strongly doubt.”

A friend at work suggested I actually look at teams, their players, and what happened that might cause such a result. (What? There’s more to football than statistics?) I couldn’t bring myself to qualitatively analyze what might be going on, but he did inspire me to dig a little deeper.

Below is a list teams that demonstrated the trend of impressive interception rates one year, followed by a severe drop-off in wins the next. The average for both offensive and defensive interception rates is 0.32.

































































































































































































YearTeamWinsNext WinsO Int RateD Int Rate
2002TB1270.0180.061
2002OAK1140.0160.037
2003TEN1250.0180.038
2003KC1370.0220.044
2004PHI1360.0200.031
2005TB1140.0290.036
2004NYJ1040.0250.038
2005JAX1280.0120.039
2003SF720.0290.045
2005CIN1180.0260.060
2004HOU720.0300.042
2002ATL950.0250.047
2003MIA1040.0420.042
2005DEN1390.0150.033
2005WAS1050.0230.030
2004BUF950.0370.049
2004SD1290.0180.038
2003STL1280.0380.047
2004NE14100.0290.037
2004BAL960.0240.042
2005SEA1390.0210.028
2004NO830.0300.024
2004GB1040.0320.015


These teams all exhibited a notably better than average interception rate on either offense or defense, only to suffer a dramatically worse record the next year.

Below is the list of teams that exhibit the opposite trend, no matter how slight. They exhibited better than average interception rates, then improved their record the following year.









































YearTeamWinsNext WinsO Int RateD Int Rate
2005TEN480.0240.019
2005NYJ4100.0320.045
2004NYG6110.0270.030
2002KC8130.0270.029


Whatever the reason for the phenomenon, it appears real. It might be summed up by "Live by the interception, die by the interception." If a team wins a lot of games based on superior offensive (low) or defensive (high) interception rates, it tends to be extremely difficult to repeat, and that team will very probably not win as many games the following year.

In the next post, I'll apply the leading indicators to the team stats from 2006. I'll list how each team can be expected to benefit (or suffer) in 2007 from the leading indicators of NFL wins.

Leading Indicators 1

Predicting team win totals before the season begins is a very inexact science. Although I’ve predicted the win total of all the teams for 2007 based on last year’s performance stats, the estimates are fairly vague. The reason for the lack of confidence is that team performance in one area does not necessarily remain consistent across seasons. But we can measure which team performance predictors do tend to persist from year to year. The stats that endure as predictors of following year wins can be considered leading indicators.

I ran two regressions. The first was my usual model using efficiency stats to estimate team wins for the year in question. The second model used the same efficiency stats to estimate team wins for the next year. In other words, it used 2002 stats to predict 2003 wins. The data set included the ’02-’06 seasons. By comparing the results, we can see which stats tend to be consistent predictors from year to year.

The efficiency stat predictors were converted into standardized variables. This way, they can be directly compared to each other in terms of their relative importance in estimating wins. The % Persist column calculates the proportion of predictive power retained from one year to the next. It was calculated by dividing each coefficient of the next year model by the current year model, then adjusting for the r-squared of each regression.






























































































































Same Yr WinsNext Yr Wins
VARIABLECOEFSIG.VARIABLECOEFSIG.% Persist
O Pass1.220.00O Pass0.420.119.4
D Pass-1.110.00D Pass-0.010.970.2
O Run0.420.00O Run0.560.0136.0 *
D Run-0.160.23D Run0.000.99-0.5
Penalties-0.230.07Penalties-0.440.0952.6 *
O Fum-0.420.01O Fum-0.050.833.5
D FFum0.470.00D FFum0.780.0045.3 *
O Int-0.320.03O Int0.400.07-34.2 ?
D Int0.600.00D Int-0.340.09-15.5 ?
r-squared0.75r-squared0.20

The results of the first regression produced expected results. It estimated present year wins very well (r-squared=0.75), with all variables significant. The second regression, which predicted following year wins, was expectedly much weaker (r-squared=0.20), but it revealed which stats endure from year to year as predictors of team wins.

It shows that offensive run efficiency, team penalties, defensive forced fumbles, and interceptions thrown are relatively persistent predictors of following year wins. Defensive pass and run efficiencies are not consistent predictors.

I adjusted the coefficients in each model by their respective model’s r-squared values. Then I divided the second (next year) model’s coefficients by the first. This tells us the percent of predictive value of each stat that survives from one season to the next. In other words, I calculated how much of each stat’s predictive power survives the off-season to help predict next year’s wins. For example, only 9% of the predictive power of offensive pass efficiency endures.

We see that the stronger persisting stats are offensive running efficiency (36%), team penalties (52%), and defensive forced fumble rate (45%).

Notice that the interception rate stats also show persistence (45%, 34%), but that the signs of the coefficients are reversed between models. This means that these stats could be considered ANTI-predictors. In other words, a low offensive pass interception rate in one year signifies fewer wins the following year. This is unexpected and could be just due to their marginal significance. But although p-values of 0.07 and 0.09 may not good enough for the FDA to approve heart medication, it’s still extremely likely that the results signify something is at work.

My theory is that we are witnessing regression to the mean. For many teams, interception rates have a lot of variance due to luck. So a team that is unlucky with interceptions one year is not likely to be as unlucky the next year. That could partially explain the reversed signs. Another possibility is that teams systematically swing from high to low interception rates from one season to the next, something I strongly doubt. Otherwise, I’m at a loss to explain this result.

Examining the results as a whole, including the lack of persistence in defensive stats and the anti-prediction of interception stats, indicate that defensive performance, and secondary performance in particular, is not persistent from year-to-year as an indicator of team win totals. It is not reliable as an indicator of wins from season to season compared to other facets of the game.

Continue reading part 2 of this article.

Median Rushing Yards

What's the difference between these two situations?

1. On 1st and 10 from the opponent's 30, a RB gets a handoff and breaks free for a 30 yd TD.

2. On 1st and 10 from his own 30, a RB gets a handoff and breaks free for a 70 yd TD.

In both plays, the RB read the blocks and made the moves necessary to break into the open field. In both plays the RB's speed and agility beat the safeties. But the difference of 50 yds is basically statistical trash because in situation #1, the RB likely could have kept running for another 50 yds.

In rating running ability, I've previously suggested the use of median statistics rather than average statistics. When we want to know how good a team's running game is, or how good a RB is, we want to know the central tendency of the team or player. The statistical mean is only one way of looking at central tendency. Median can often be more useful. Averages are often distorted by a very few outlier inputs.

Consider this fictitious example examining the central tendency of college dropouts who live in Redmond, WA. Let's say there are 5,000 college dropouts in Redmond, Washington, and each make $30,000/yr except this one guy named Gates. He makes $20 billion/yr. The average salary of a college dropout in Redmond is over $4 million/yr. So if I'm a student in Redmond I should skip class tomorrow, right? $4 million/yr might be the average, but it's not the central tendency and it's virtually useless information.

Which player would you rather have on your team? A RB who gets at least 4 yds on every carry, or a RB who gets 29 1-yd runs and one 91-yd run? Both players average 4 yds/carry. The first player's median rush is 4 yds and the second player's is 1 yd. It's an extreme example, but it illustrates my point. Consistency has its value.

Here are a list of the top runners of 2006 sorted in order of their percentage of runs >4 yds. It's interesting to compare to their average yds/rush, their total yards, and other stats. (Ties are broken by % of carries > 3 yds.)












































































































































































































































































































































































































































































































RBTEAM4YD PCTATTYDSAVGTDFUMLSTTD/ATT%
NorwoodATL57996336.42002.0
AddaiIND5422610814.87223.1
WestbrookPHI5024012175.17112.9
WashingtonNYJ501516504.34112.6
BettsWAS4824511544.74421.6
BarberNYG4732716625.15311.5
BarberDAL471356544.8140010.4
TomlinsonSD4634818155.228218.0
GoreSF4631216955.48552.6
JonesCHI4629612104.16112.0
McAllisterNO4624410574.310214.1
VickATL4612310398.42421.6
Jones-DrewJAX461669415.713117.8
DillonNE461998124.113226.5
JacksonSTL4534615284.413213.8
DayneHOU451516124.15103.3
BrownMIA4424110084.25422.1
McGaheeBUF442599903.86422.3
FargasOAK441786593.71100.6
BensonCHI441576474.16003.8
HenryTEN4327012114.57312.6
WilliamsTAM432257983.51220.4
JamesARI4233711593.46331.8
FosterCAR422278974.03321.3
MaroneyNE421757454.36113.4
RhodesIND421876413.45222.7
TaylorMIN4130312164.06432.0
TaylorJAX4123111465.05212.2
Bell T.DEN4123310254.42330.9
Bell M.DEN411576774.38105.1
JohnsonKAN4041617894.317224.1
JohnsonCIN4034113093.812623.5
LewisBAL4031411323.69422.9
JonesDAL4026710844.14111.5
DroughnsCLE402207583.44541.8
ParkerPIT3933714944.413643.9
DunnATL3928611404.04101.4
AlexanderSEA392528963.67532.8
MorrisSEA391616043.80110.0
GreenGB3726610594.05221.9
JonesDET361816893.86443.3

I'm not suggesting average rushing is worthless, just that it is only part of the story.

Rushing TDs and Passing

Rushing touchdowns are the holy grail of fantasy football. They're the most scarce scores. Everyone starts the draft by picking 2 RBs (except the guy who always takes Peyton Manning first).

Rushing TDs are a product of a great running game, right? That's only half true. Rushing TDs have just as much to do with a good passing game as a running game. That's not shocking news to most serious football fans. No team can run the ball all the way down the field without at least a few pass completions. But the extent to which rushing TDs are dependent on the passing game may surprise some.

Rushing TDs correlation with:
---------------------------------
Team Yds/Rush 0.50
Team Yds/Pass Att 0.46

Going a little deeper, we can run a quick regression using rushing and passing efficiency (including sack yards) to estimate rushing TDs. Interceptions are a big part of the passing game--teams that throw a lot of INTs would be expected to limit their opportunities for rushing TDs. So I'll include interception efficiency in the model. I'll also use standardized variables so we can directly compare the relative importance of each variable. Based on data from the '02-'06 seasons, the regression produces the following results:















































VARIABLECOEFFICIENTSTDERRORT STATP-VALUE
const13.510.3538.880.00
Z RUN AVG*2.560.357.270.00
Z PASS EFF*2.160.395.540.00
Z INT RATE-0.340.39-0.880.38
R-squared0.42

The coefficient of RUN AVG (yds/rush) is 2.56 while the coefficent of PASS EFF (yds/att) is 2.16, which tells us that a good passing game is almost as important as a good running game in producing rushing TDs.

But if we consider the importance of INT RATE (INTs/att), we see that the importance of the passing game nearly equals that of the running game in producing rushing TDs (2.56 vs. 2.50).

So if other teams in your league have already picked up Stephen Jackson and Larry Johnson, or if you're looking for a #3 RB to fill the gap when LT has a bye in week 7, then look for an overlooked RB on a team with a decent passing game.

Luck: Epilogue

Coincidentally, as I was posting the results of my look at the amount of luck in NFL games, Phil Birnbaum posted this at his site. He was sharing a paper he did a while back about how "truly" good an MLB team is that wins 100 games. If some games are won on merit, and some by luck, then his calculations say the average 100 game winner won by luck 7 games more than they merited based on the their talent level. In other words, 100 game winners are probably both good and lucky.

But even more interesting was another tidbit Phil linked to. If you follow the references, you land here on Tom Tango's site. He's another accomplished sabermetrician. He approached the question of luck in sports outcomes far more elegantly than I did.

There, he works through his math calculating how many games are required for a sports league to produce the "true" best team on top of the standings. For MLB he says it's 69 games, and for the NFL it's 12 games.

Along the way, Tango articulates his theory. Regarding the distribution of win-loss records, the observed variance is: variance(observed) = varance(true) + variance (luck). Since we know the variance of the observed distribution (SD^2), and we know the variance of luck from the binomial distribution (p=0.5, n=16), we can solve for variance (true), which is the variance in team records based on merit.

I'm not sure what to think about his method. His theory assumes the "true" distribution (what I call pure-skill) is narrower than the observed distribution. Then luck acts on the true distribution to widen it.

But my simulations show that the distribution of a pure merit league is much wider than either the observed or the luck distributions. So I'm not sure how to interpret his theory.

On another note, I reran my simulation against the 96-01 NFL seasons. The scheduling system was a little different then, but the effect should be minimal. The simulation maximized its goodness-of-fit at 51% luck, which is pretty much what I found for the 02-06 seasons.

Upcoming posts include a look at YAC stats and how they affect QB stats, and a look at what really produces rushing TDs--something for the fantasy football fans out there.

Luck and NFL Outcomes 3

This is the third and final post of an article discussing the amount of luck in determining outcomes in the NFL. In the first post, I compared the actual distribution of team win-loss records over the past five seasons with an idealized pure luck distribution. I found that only 78 out of 160 actual season records (48%) differed from what we’d expect if the NFL were determined completely by luck.

In the second post, I compared the actual distribution with an idealized distribution of records in a theoretical league governed by “pure skill.”

In this post, I will unify the three distributions--actual, luck, and skill--into one algorithm. The resulting equation reveals the proportion of NFL games in which the deciding factor is luck and not the camparitive strength of each opponent.

LUCK, SKILL, AND OBSERVED

The chart below is a histogram of the pure binomial distribution, the simulated pure-skill (zero luck) distribution, and the actual distribution of NFL records since the '02 expansion. (Pure luck is blue, actual is yellow, and pure skill is red.)


When I first examined the three distributions together, I was struck by how the actual distribution appeared to split the difference between the luck and skill distribution. The actual records appear to be some sort of combination of the luck and skill distributions. To me, it looked as if the pure-luck binomial distribution was pressed into a flatter and wider distribution by skill.

LUCK/SKILL SYNTHESIS

It dawned on me to create another simulation, one that synthesized the pure-luck and pure-skill distributions together in varying degrees. (10% luck/90% skill, 20% luck/80% skill, etc.) Basically, the luck% variable determined a percentage of games (chosen at random) to be decided by pure luck, essentially a coin flip. The remainder of the games were credited to the superior team. The simulation algorithm looked like this:

If rand() < %luck, then game outcome = pure luck, else game outcome = win by the better team

I varied the %luck value between 0 and 1, re-running the simulation. Here are some representative win distributions (legend is in %luck):


Next I overlayed the actual distribution.


PROPORTION OF LUCK IN NFL OUTCOMES

Then I varied the %luck value until it maximized the goodness of fit between the actual distribution and the synthesized distribution. At 52.5% luck, the theoretical distribution is statistically indistinguishable from the actual distribution (chi-square goodness of fit p=0.94). This means it is 94% probable that the discrepancies between the synthesized simulation and the actual observations are merely due to sample error.


THEORETICAL MAXIMUM PREDICTION RATE

I will be very careful in stating what conclusion I draw from this exercise. The actual observed distribution of win-loss records in the NFL is indistinguishable from a league in which 52.5% of the games are decided at random and not by the comparative strength of each opponent.

I admit 52.5% seems very high. But keep in mind, that half of the time, the better team wins by luck. In other words, half the time our prediction models are correct by chance, just like a monkey picking winners would be. If the 52.5% figure is correct, the best any prediction model could do is:

0.50 + 0.525/2 = 0.76

So 76% correct would be the theoretical ceiling for NFL game prediction models. This is consistent with the various computer models as well as odds makers. It is also consistent with our intuitive experience--upsets seem happen about a quarter of the time. Sometimes a model (or a person) can predict at better than a 76% correct rate, but anything above that would be...by luck.

I also posted a follow up to this series of articles here.