Showing posts with label other sports. Show all posts
Showing posts with label other sports. Show all posts

Lacrosse Analytics

I'm a Baltimore guy, and aside from an affinity for steamed crabs and a regrettable taste for National Bohemian beer, the mid-Atlantic has given me an appreciation for the sport of lacrosse. To most North American sports fans, lacrosse must seem like some strange niche sport, like "jousting" or "baseball." But it's very entertaining and fun to watch. It's growing fast, particularly in the super-zips around DC where the ANS headquarters is.

For those not familiar with lacrosse, imagine hockey played on a football field but, you know, with cleats instead of skates. And instead of a flat puck and flat sticks, there's a round ball and the sticks have small netted pocket to carry said ball. And instead of 3 periods, which must be some sort of weird French-Canadian socialist metric system thing, there's an even 4 quarters of play in lacrosse, just like God intended. But pretty much everything else is the same as hockey--face offs, goaltending, penalties & power plays. Lacrosse players tend to have more teeth though.

Because players carry the ball in their sticks rather than push it around on ice, possession tends to be more permanent than hockey. Lacrosse belongs to a class of sports I think of as "flow" sports. Soccer, hockey, lacrosse, field hockey, and to some degree basketball qualify. They are characterized by unbroken and continuous play, a ball loosely possessed by one team, and netted goals at either end of the field (or court). There are many variants of the basic team/ball/goal sport--for those of us old enough to remember the Goodwill Games of the 1980s, we have the dystopic sport of motoball burned into our brains. And for those of us (un)fortunate enough to attend the US Naval Academy (or the NY State penitentiary system) there's field ball. The interesting thing about these sports is that they can all be modeled the same way.

So with lacrosse season underway, I thought I'd take a detour from football work and make my contribution to lacrosse analytics. I built a parametric win probability model for lacrosse based on score, time, and possession. Here's how often a team can expect to win based on neutral possession--when there's a loose ball or immediately upon a faceoff following a previous score:

Optimizing a Swim Meet: Traveling Salesmen and Asexual Mutants

You wouldn't think there is much to swimming analytics. Compared to sports like baseball or football, swimming is extremely deterministic. Swimmers tend to have a certain speed in each stroke, and they vary only slightly around that tendency from meet to meet. There are no interactions with teammates, collisions with opponents, or bouncing balls to worry about. But it turns out that there's more to aquametrics than meets the eye.

My kids are on a summer swim team in the local league. It's a great activity--exercise, a bit of healthy competition, and all four of my kids and step-kids are on the same team for one season out of the year. It's a lot of fun for everyone.

It's complicated, though. There are five age groups for both boys and girls for a total of ten competition groups. There are 4 strokes (fly, back, breast, and freestyle). Each swimmer is assigned to one of three classes for each stroke. The A class has the faster kids, the B class the next faster kids, and the C class has the rest. First, second, and third place in each stroke-class (for each age/gender group) earn points for the team. The A, B, and C classes all count equally, in the spirit of the league. It's 5 points for a 1st, 3 for a 2nd, and 1 for a 3rd, regardless of class. This way, nearly everyone's performance can affect the outcome of the meet.

There's only one strategic variable in the meet. Each swimmer can only swim in 3 of the 4 stroke events in the meet. In other words, each swimmer has to skip one of the four strokes for the day. The manager of each team seeds the meet a couple days beforehand. You'd think that the best strategy is to have each swimmer participate in their 3 best strokes. But that's not the case.

Live NCAA Basketball Win Probability

Live win probability for the NCAA basketball championship games is available now at wp.advancednflstats.com/bball. (Final games are here. Games from the previous day are here.)

This is something I put together a couple years ago, and I've dusted it off for the tournament. The model's approach is very similar to my football model. Basketball is a much simpler sport, though. There's no field position, down, or distance in basketball. Aside from possession, score and time remaining are really the only significant statistical factors.

World Cup Penalty Kicks, Wimbledon Serves, and Intuitive Algebra

Maybe June should be "other sport month" here at Advanced NFL Stats. Except for Albert Haynesworth's principled, valiant stand against being forced to play defensive tackle a foot and a half to the right from where he is accustomed, there's not much going on in the NFL. Fortunately, there's plenty of other sports going on, including the world's biggest events in soccer and tennis.

Soccer and tennis offer two of the best examples of simple two-strategy zero-sum game theory. Soccer offers us the penalty kick, when a player kicks to the left or right extreme side of the goal so hard that the goalkeeper must simultaneously guess a direction to lunge. Tennis gives us the serve, where the server aims for either the extreme forehand or backhand side of his opponent's service box. Both examples give us the opportunity to examine how well experts are able to approximate the optimum strategy mix.

Actual vs. Theoretical WP at the World Cup

In response to a few questions on my last post regarding my World Cup win probability (WP) model, here are some actual numbers to chew on. An anonymous commenter pointed us to actual win rates at WhoWins.com (a fun site by the way). I've graphed the actual rates below.

The actual win rates are for 708 games stretching all the way back to 1930. The theoretical WPs based on a Poisson distribution are the solid lines, and the actual rates are the little triangles and squares. Keep in mind these are the WPs for the trailing team.

Open Wide for Some Soccer!

"Fast kickin', low scorin',...and ties--you bet!"

I just returned home to the good ol' USA after a week overseas. Evidently, there is some sort of sporting tournament underway in some other country somewhere. It's confusing because sometimes they call this sport "football," except that it's the same sport my daughter played when she was five. Too funny! Still, it is considered a real sport by some, and so it's my job to suck the fun out of it by creating a Win Probability (WP) model, making you realize there's no chance your favorite team can come back to win.

NASCAR Game Theory

When we talk about game theory in sports we almost always talk about zero-sum games. Whatever one team gains, the other loses, whether it’s yards, wins, or even win probability. It’s all very tidy.

One sport that features some non-zero-sum games is auto racing. I’ll admit that until recently I was a sports snob. Spending four hours watching cars make left turns and have their tires changed was never my idea of a good use of time. But now I have to say I am intrigued by NASCAR.

Drivers accumulate points throughout the season based on their finishing positions, and the top 10 drivers qualify for a playoff-type system in the final few races. Every race is worth the same number of points. The winner gets 185 points, the second place finisher gets 170 points, and the third place finisher gets 165 points. You can see the full table here. We can agree that points are not the only consideration in the value of winning a race, but for the sake of discussion I’ll limit the value just to the points.

Before I get into the meat of NASCAR’s non-zero sum game, consider the classic Ultimatum game in which two players divide a sum among themselves. The first player decides how to divide the sum and makes a single offer to the second player who can accept or reject the offer. If the second player accepts the offer, they spilt the sum accordingly, but if he rejects the offer, neither player receives a payout.

Live NCAA Basketball Win Probability

Live win probability for the NCAA basketball championship games is available now at wp.advancednflstats.com/bball. (Final games are here. Games from the previous day are here.)

The model's approach is very similar to my football model. Basketball is a much simpler sport, though. There's no field position, down, or distance in basketball. Score and time remaining are really the only significant factors.

Verducci Follow-Up

The recent post about the Verducci Effect and Let's Make A Deal didn't elicit the response I was looking for. Reactions ranged from denial to sadness to even anger. I think the game show story hurt more than helped my case of why I think the Verducci Effect is an illusion. And as I said in the original article, I'm not completely certain. But now, after developing my thoughts a little better, I'm more certain than before.

Game shows aside, I'll explain my thought process with an notional example, a mental exercise actually. I'm most interested in the injury aspect of the Verducci Effect, so I'll concentrate on that in this post. The injury rates I'm going to use are created only for clarity, and they are not intended to match the true rates. I'm also going to make some simplifying assumptions to illustrate the broader point. Please keep in mind this example is only intended to demonstrate a concept.

The Example

Assume every MLB pitcher's career lasts exactly 5 years. Also assume that the league-wide injury rate is 1 out of 5 years, defined as however you like--say being on the DL. Also assume that one year out of each pitcher's career can be identified, after the fact, as a "career year," which by definition assures us of two things: no significant injury and an upswing of innings.

Take 200 pitchers and randomly assign a number from 1 through 5 to each of their 5 respective years completely independently. When a 1 comes up, call that an injury year. So far, we've got a 1 in 5 (20%) injury rate across the league. Some pitchers will have multiple injury years, some won't have any, and they are completely independent.

In my head I'm thinking of a table of cards, 200 x 5. For now, the cards are turned up so we can see the numbers 1 through 5. 20% of the cards are 1s--injuries.

Now, take all the "career years" for the pitchers off the table by removing one card from each row, which by definition cannot be an injury year. Let's choose the highest card and remove it. What percentage of cards will now be 1s (injuries)? Before, there were 200 out of 1000 (20%), and now there are the same number of injuries (200) but fewer cards remaining (800). 25% of the remaining years are 1s (injuries). If we now selected a card at random, we'd have a 1 in 4 chance at finding an injury.

Shrink the Sample

Let's do the same exercise but with a sub-sample of 50 pitchers instead of 200. There are still 20% injury years, and if we take each pitcher's known "career year" off the table, 25% of the remaining years will be injury years. Turn over the remaining cards, so you can't see the numbers 1-5. Turn one card face up at random--what are the chances of finding a 1-card (an injury)? It has to be 25%. We started with 250 cards, but there are now 200 cards remaining and 50 of them are injury cards.

Shrink It Again

Repeat the exercise with 10 pitchers. Does this change anything? No. Originally 20% of the cards were injuries, and after removing the career year cards, 25% of those remaining are injuries.

Down to One Player

Now consider a sub-sample of just a single pitcher. Again, turn over all the cards so we can't see the numbers 1-5. There are originally 5 cards on the table, with a probability of 1 in 5 being an injury card. Take away one card, which we know after the fact is not an injury card, leaving 4 cards. Turn one card over--the card dealt immediately following the "career" year. What is the chance it's an injury?

It would be 25%. We had a league-wide injury rate of 20%, but following career/high-inning years we would retrospectively observe a rate of 25%. Even though the injuries were distributed completely at random and completely independently, we'd see a false connection between high-inning years and injuries in subsequent years.

Even a small difference would appear statistically significant with a large data set, but it would be an illusion. The original probability of a year being an injury year was always 20%, but after looking back and removing a year in which we're virtually assured of no injury, we'd see a 25% injury rate.

Try It Yourself

If you don't believe me, you can play the game yourself. Shuffle a deck of cards and deal out 4 face up in row. Those 4 cards represent 4 years of a pitcher's career. Every time we see a diamond card, we'll call that an injury year. We'll say the highest non-diamond card is a career/high-inning year. It goes without saying that, on average, 1 in every 4 cards will be a diamond--an injury. That's our true baseline rate.

After dealing the 4 cards, remove the highest non-diamond card and set it aside. Look at the card immediately to the right of the one you removed. What is the probability it is a diamond? If you said 1 in 4 you'd be mistaken. It's 1 in 3. This is the same illusion.

Try it. I did, 78 times and got a diamond on 26 tries--exactly one third (p=0.04 for the sticklers out there). You have to re-shuffle each time for it to be completely random and independent. Also, if the high "career" card is the right-most card, you can either throw out that iteration or loop around to look at the first card. The effect is the same. In fact, just look at how many of the 3 remaining cards are diamonds, and you'll eventually see that it's 1 in 3.

I'm sure there is a name for this, but I don't know it. If anyone is familiar, fill me in. Otherwise, I'm sticking with the Monty Hall Effect. Also, the Verducci Effect may still be real, but it would have to be shown that the observed injury rates significantly exceed the rate predicted by the effect of the illusion.

Maybe I'm wrong, and that's ok. But every time I deal 4 cards and remove 1 non-diamond, I keep seeing diamonds 33% of the time. Just like the Monty Hall game, you probably won't believe it until you try it yourself.

[Edit: I am now completely certain the paradox/illusion exists as I described. However, after a good discussion with commenter Vince (see below), I'm no longer convinced that the way I set up the question applies to how the Verducci Effect is truly applied. In other words, the illusion is real if you set it up the way I did. It's just that strange, and subtle differences exist in how you look at the problem. For example, if in the pitcher example, you removed the first instance of a non-1, you would see a 1 in the following block 20% of the time, just like you'd expect. But if you select the highest number in the row, or if you remove the final non-1, you'd detect a 1 next 25% of the time.]

Entropy, Let's Make a Deal, and the Verducci Effect

Let’s Make a Deal was a 1970s game show, entropy is the second law of thermodynamics, and the Verducci Effect is an injury phenomenon named for a Sports Illustrated reporter. What do they have in common?

In Let's Make a Deal, host Monty Hall would walk the costumed audience, picking contestants on the spot to play various challenges for prizes. The central challenge was a simple game where the contestant had to choose one of three doors. Behind one of the doors was a big prize, such as a brand new Plymouth sedan. But behind the two other doors were gag prizes, such as a donkey.

Sounds simple, right? The contestant starts with 1 in 3 chance of picking the correct door. But then Monty would open one of the doors (but never the one with the real prize) and with two closed doors remaining, ask the contestant if she wanted to switch her choice. She would waffle as the audience screamed “switch!...stay!...switch.”

The answer is intuitively obvious. It doesn’t matter. She has a 1 in 3 chance when she first picked the door, and we already know one of the other two doors doesn’t have the real prize. So whether she switches or not is irrelevant. It’s still 1 in 3.

...And that would be completely wrong.

The real answer is she should always switch. If she stays, she has a 1 in 3 chance of winning, but if she switches she has a 2 in 3 chance of winning. I know, I know. This doesn’t make any sense.

Don’t fight it. It’s true. If the contestant originally picks a gag door, which will happen 2 out of 3 times, Monty has to open the only remaining gag door. In this case, switching always wins. And because this is the case 2/3 of the time, always switching wins 2/3 of the time.

(If you don’t believe me, visit this site featuring a simulation of the game. It will tally how many times you win by switching and staying. It’s the only thing that ultimately convinced me. But don’t forget to come back and find out what this has to do with the Verducci Effect.)

Baseball Prospectus defines the Verducci Effect as the phenomenon where young pitchers who have a large increase in workload compared to a previous year tend to get injured or have a decline in subsequent year performance. The concept was first noted by reporter Tom Verducci and further developed by injury guru Will Carroll.

But I'm not sure there really is an effect. First, consider why a young pitcher would have a large increase in workload. He’s probably pitching very well, and by definition he’s certainly healthy all year. Bad or injured pitchers don’t often pitch large numbers of innings.

Now, consider a 3-year span of any pitcher’s career. He’s going to have an up year, a down year, and a year in between. Pitchers also get injured fairly often. There’s a good chance he’ll suffer an injury at some point in that span.

Injuries in sports are like entropy, the inevitable reality that all matter and energy in the Universe are trending toward deterioration. Players always start out healthy and then are progressively more likely get injured. Pitchers don’t enter the Major Leagues hurt and gradually get healthier throughout their career. It just doesn’t work that way. Injuries always tend to be more probable in a subsequent year than any prior year. The second year in a 3-year span will have a greater chance of injury than the first, and the third would have a greater chance than the second.

Back to Let’s Make a Deal. Think of that three year span as the three doors. Without a Verducci Effect, the years would each have an equal chance at being an injury year. For the sake of analogy, say it’s a 1 in 3 chance. Now Monty opens one of the doors and shows you a non-injury year. The remaining doors have a significantly increased chance of being identified as an injury year. In this case, it’s a 1 in 2 chance.

I think that’s essentially what Verducci and Carroll did in their analysis. We already know a high workload season can’t be an injury season, therefore subsequent years will retrospectively appear to have higher injury rates. We would normally expect to see injuries in 1 out of 3 years, but we would actually see them 1 out of 2. It’s an illusion.

The analogy isn't perfect. Door one is always the open door without the prize, and there's no switching. Also, unlike a single prize behind one door, injuries can be somewhat independent (or more properly described in probability theory as "with replacement"). That is, a pitcher could be injured in more than just one year. But the Verducci Effect only considers two-year spans, and since one year is always a non-injury year, the analogy holds in this respect.

Ultimately, just like in Monty Hall’s game, the underlying probabilities don’t change at all. Only the chance of finding what we're seeking changes. There was always a 1 in 3 chance that one particular door would contain the prize. That never changes throughout the course of the game. But after identifying a non-prize door, we’ve increased our chances of finding the injury…err…I mean Plymouth.

I hereby name this phenomenon the Monty Hall Effect.

(PS Quite frankly, I’m not entirely confident in this. It’s hard to wrap my head around, and I keep second-guessing my logic. If someone out there, like a quantum physicist maybe, understands this stuff well, please add your two cents.)

Edit: See my comment for an alternate explanation of how the Verducci Effect may be an illusion.

An Underdog Wins with Aggressive, Risky Football

No, not that kind of football.

A couple weeks ago I wrote a post about how underdogs can increase their chances of winning by employing a high-risk, high-reward strategy. It seems that’s just what the US soccer team did in their recent upset against the globe's top team, Spain.

According to this analysis by the Journal’s Carl Bialik, the American team uses long aggressive passing, looking for fast-break scores, instead of using a more typical ball control offense. This opens up opportunities for a quick goal, but usually results in the opponent controlling the ball on the US side of the field (or pitch, if you’re a ‘football’ aficionado). As long as the goalie has a good game, and the defense gets some breaks, the strategy works.

It makes sense because the US team has nothing to lose. No one expects them to go very far in World Cup play, so they can afford to use a risky gameplan without being humiliated if they end up losing 4-0.

Epic Bulls-Celtics Series

If you haven’t been paying attention to this series you’re missing one of the most exciting 7-game series ever, even if it’s just the first round.

Going into Thursday night’s game, there were already 3 OT games and 1 last second buzzer-beater game. Then this happened: A 3OT potential elimination game that featured multiple furious comebacks for both teams and ultimately tied the series 3-3 with a 1-point win by Chicago. These two teams are as evenly matched as it gets. They square off for Game 7 tonight at 8.



Here are all 6 games of the series so far.

Compared to basketball or football, the hockey graphs aren't as compelling at first glance, at least during the game. But a quick look at a graph after the game tells a dramatic story you just won’t get with a box score. Below is Thursday night’s Game 1 between Chicago and Vancouver.

5-3? Well, that doesn’t sound terribly exciting. But check out what happened: 3-0 lead held until the third period. With 10 minutes left in the game, another goal, and with 5 min left another to tie it 3-3. Then Vancouver gets the game-winning goal with less than 2 min to go. Then in the final seconds, a garbage goal with the goalie pulled makes it 5-3.


You can check out the beta version of the hockey graphs for today's Caps-Penguins and Blackhawks-Canucks games. Power plays have not been factored in yet, but that's in progress, and it should be ready early next week.

NHL In-Game Win Probability

I was at an NHL game the other night, and with the score 2-0 someone asked me, “So Mr. Win Probability, what’s the chance the Capitals win?” I was caught off guard, and after I choked out, “I…don’t…know…,” I experienced the horror that is not knowing the exact up-to-the-second win probability of a sporting contest. Don’t let this happen to you.

The anxiety and shame lasted for two days straight. I kept blaming myself and replaying the incident over and over in my head. The only way to cure my depression was to build a win probability model for NHL hockey.

Unlike my previous models for basketball and football which were empirically based, my hockey model is theoretical. In other words, instead of being based on a massive database of actual previous games, the probabilities are calculated based on a Poisson scoring distribution. The distribution is calculated using the average goals scored per minute in the 2008-9 NHL season. It’s an extension of the model I developed in this post.

Teams score an average of 2.79 goals per 60 minutes of regulation time, which is equal to 0.0465 goals per minute. A Poisson distribution based on that per-minute scoring rate and the time remaining in the game yields the probabilities of each team scoring each number of possible goals by the end of the game. Summing up all the probabilities of all the possible combinations of final scores gives the game’s win probability.

Here’s the graph:


There are a couple wrinkles to address. First, there are power plays. When a team as a man advantage on the ice, it’s much more likely to score. About one in five power plays results in a goal for the team with the advantage. Only about 2% of the time the short-handed team will score. So at the start of a power play, a rough approximation would put the win probability a little less than one fifth of the way toward the next best curve.

For example, if the score is 2-0 with 30 minutes remaining in the game, the win probability would normally be about 13% for the trailing team (the red line). But at the beginning of a power play, the trailing team’s win probability would jump about a fifth of the way up to the ‘down by 1’ line (blue). A rough approximation puts the new win probability at 16%. Then as the power play expires and there’s no score, the win probability would gradually return to the ‘down by 2’ line.

Second, there is the ‘end-game,’ when teams down by a goal will pull their goalie in favor of an additional skater. That would increase the win probability of the trailing team slightly, but only half as much as you might expect. They’d still only be buying an opportunity in overtime. But it could still be factored in. Before I do, I’d need some data on end-game goals.

One advantage of a theoretical approach over an empirical model is that team strength can be factored in far more easily. In an empirical model, when you divide up the data by various classes of team strength, the data is sliced into tiny fragments, usually with very small and unreliable sample sizes. Theoretical formula-based models don’t suffer from that problem. I can simply adjust the mean goals scored and goals allowed for any particular opponent, then rerun the model. The resulting model would be tailored to the specific match-up instead of a generic model for the league as a whole. Home ice advantage can be factored in with a similar approach.

Remember, WPD (Win Probability Dysfunction) can happen at any time, and it’s nothing to be ashamed of. Don't analyze win probability graphs if you take nitrates, often prescribed for chest pain, as this may cause a sudden, unsafe drop in blood pressure. Discuss your health with your doctor to ensure that you are healthy enough to view win probability graphs. If you experience chest pain, nausea, or any other discomforts during a sporting contest, seek immediate medical help. In the rare event of viewing win probability graphs more than 4 hours, seek immediate medical help to avoid long-term injury.

Live NHL win probability graphs now online.