Archive for xWOBA

Advocating For A Different Type of Swing Change

When Statcast was launched, we were graced with incredible new stats such as Exit Velocity and Launch Angle, which revolutionized how we evaluate hitting. This new information confirmed obvious things like that Giancarlo Stanton hits missiles, but it also gave us a new breed of hitter. Daniel Murphy, Justin Turner, J.D. Martinez, and others looked at the data and made adjustments that started maximizing their power outputs. The standard evaluation method has become to look at EVs mixed with LAs to determine who is one tweak away from stardom. Hitting is a complex beast, with pitchers throwing 95-plus with nasty hooks to go with shifting defenses. Ultimately, a hitter is looking to produce solid contact regardless of where the ball goes. The goal of this analysis is to identify hitters who have an inefficient spray chart and see how they could optimize their profile by hitting more balls in a different direction to maximize production. Luckily with Statcast, we can now try to find these answers.

To do this analysis, I used Baseball Savant to gather 2018 Exit Velocity and xwOBA to Pull Side, Straight Away, and Oppo Side for all hitters with at least 50 plate appearances. I then used FanGraphs to pull the 2018 data for Pull%, Mid%, and Oppo% to discern how often a hitter attacks that field. I used 50 PAs as a filter since this is about where exit velocities become stable and helps weed out pitchers and other noise. This does create gaps in the data because some players didn’t register 50 PAs of a batted-ball direction. This dataset gives us the ability to look at how hard a hitter hits the ball to a field, what was their expected damage (xwOBA) to that field, and how often they went that way.

The first category I looked at was players who could use the opposite field more often. To do this, I looked at players who had an above average Oppo Side xwOBA and a below-average Oppo%. I used exit velocities to each field as a proxy to justify the directional swing change. Read the rest of this entry »


Joe Biagini’s xwOBA and RISP Spread

Would you believe me if I told you that Joe Biagini did a better job minimizing contact quality last year than Marcus Stroman?  I didn’t believe it at first, but it turns out he had slightly better contact quality control on the whole. See the arranged summary table below: (min 500 pitches, showing data for batted balls)

player_name xwOBA
1 Danny Barnes 0.306589888
2 Aaron Loup 0.307355422
3 Roberto Osuna 0.326513158
4 Joe Biagini 0.334128342
5 Ryan Tepera 0.33518593
6 J.A. Happ 0.339169336
7 Dominic Leone 0.340634286
8 Marco Estrada 0.341392086
9 Marcus Stroman 0.355793677
10 Joe Smith 0.368621951
11 Francisco Liriano 0.371231373
12 Aaron Sanchez 0.37646281
13 Mike Bolsinger 0.378370079

xwOBA in this case is a statcast proxy for contact quality, based on launch speed and angle. I’d go on a limb to say it essentially imputes an expected number of wOBA based on the quality of how the hitter squared up the ball, irrespective of what happens after that. Last year, the average xwOBA for a Blue Jays pitcher included in this sample above was 0.344.

Interesting. Biagini’s contact quality was fourth best on the team, but his ERA was third highest among the group. The two higher ERAs were Liariano and Bolsinger (also 11th and 13th highest expected wOBA). This is an example of how situational pitching can ruin you if you let it. Let’s factor in runners in scoring position and compare the same analysis. Below is the same table, except now showing one for runner in scoring position (RISP), 0 for not:

player_name RISP xwOBA
1 Joe Biagini 0 0.299694737
2 Aaron Loup 0 0.302101695
3 Roberto Osuna 0 0.309736842
4 Ryan Tepera 0 0.325156463
5 Danny Barnes 0 0.339150794
6 J.A. Happ 0 0.350223214
7 Marco Estrada 0 0.350911628
8 Marcus Stroman 0 0.352419087
9 Francisco Liriano 0 0.363611702
10 Dominic Leone 0 0.364512605
11 Aaron Sanchez 0 0.369082353
12 Mike Bolsinger 0 0.372388235
13 Joe Smith 0 0.395754386
14 Danny Barnes 1 0.227692308
15 Dominic Leone 1 0.289892857
16 J.A. Happ 1 0.30239604
17 Joe Smith 1 0.30676
18 Marco Estrada 1 0.308904762
19 Aaron Loup 1 0.320270833
20 Ryan Tepera 1 0.363538462
21 Marcus Stroman 1 0.369462185
22 Roberto Osuna 1 0.376842105
23 Mike Bolsinger 1 0.39047619
24 Francisco Liriano 1 0.39261194
25 Aaron Sanchez 1 0.393888889
26 Joe Biagini 1 0.444393258

Look at the two Biaginis! At the very top and very bottom. Without runners in scoring position, Biagini was the best pitcher on the roster in terms of limiting contact quality. Put a guy in scoring position, and he starts getting lit up. Here’s that same table, but sorted by the differences.

player_name RISP.x xwOBA.x RISP.y xwOBA.y diff
1 Danny Barnes 0 0.339150794 1 0.227692308 -0.111458486
2 Joe Smith 0 0.395754386 1 0.30676 -0.088994386
3 Dominic Leone 0 0.364512605 1 0.289892857 -0.074619748
4 J.A. Happ 0 0.350223214 1 0.30239604 -0.047827175
5 Marco Estrada 0 0.350911628 1 0.308904762 -0.042006866
6 Marcus Stroman 0 0.352419087 1 0.369462185 0.017043098
7 Mike Bolsinger 0 0.372388235 1 0.39047619 0.018087955
8 Aaron Loup 0 0.302101695 1 0.320270833 0.018169138
9 Aaron Sanchez 0 0.369082353 1 0.393888889 0.024806536
10 Francisco Liriano 0 0.363611702 1 0.39261194 0.029000238
11 Ryan Tepera 0 0.325156463 1 0.363538462 0.038381999
12 Roberto Osuna 0 0.309736842 1 0.376842105 0.067105263
13 Joe Biagini 0 0.299694737 1 0.444393258 0.144698522

Biagini was not the same person on the mound when threatened with a runner past first. To offer some perspective, that very difference is larger than that between Mike Trout (1st @ 0.437) and Kevin Pillar (130th @ 0.302). You must wonder what some possible explanations of this could be .Sign stealing? The yips? Pitch selection? Let’s look at the 2016 differences table and see if this affected him at all. (min 500 pitches)

player_name RISP.x xwOBA.x RISP.y xwOBA.y diff
1 Jason Grilli 0 0.429316 1 0.229538 -0.19978
2 Drew Storen 0 0.475368 1 0.320645 -0.15472
3 Joe Biagini 0 0.345207 1 0.270319 -0.07489
4 R.A. Dickey 0 0.39697 1 0.337583 -0.05939
5 Aaron Sanchez 0 0.366077 1 0.328248 -0.03783
6 Roberto Osuna 0 0.386113 1 0.350638 -0.03547
7 Brett Cecil 0 0.411931 1 0.379067 -0.03286
8 Marcus Stroman 0 0.362472 1 0.334 -0.02847
9 Francisco Liriano 0 0.351901 1 0.369333 0.017432
10 J.A. Happ 0 0.363058 1 0.386705 0.023646
11 Marco Estrada 0 0.335267 1 0.359102 0.023835
12 Jesse Chavez 0 0.340885 1 0.48161 0.140725

Runners on second and or third in 2016, Biagini pitched to better contact quality.  He was coming out of the bullpen, but it still leaves our question of consistency from last year unresolved. It wasn’t Biagini’s pitch selection either. Based on the table below, his distribution of pitches with and without RISP last year was more or less the same. It’s not as though he wasn’t throwing the breaking ball with RISP.

pitch_type RISP n Frequency
1 CH 0 207 0.146393
2 CU 0 278 0.196605
3 FC 0 145 0.102546
4 FF 0 784 0.554455
5 CH 1 80 0.154739
6 CU 1 138 0.266925
7 FC 1 39 0.075435
8 FF 1 260 0.502901

I don’t know what the real explanation for this is. It likely could just be chance, but I’d like to think there’s a more probable explanation for it. I say the yips! Pitchers aren’t robots, some pitchers must get phased more than others by the pressure of potential runs scoring. But last year on the whole Toronto pitching allowed very similar contact quality regardless of having runners in scoring position.

RISP xwOBA
1 0 0.343178
2 1 0.345718

P.S. first time posting! let me know what you think. had a lot of fun doing this.


Using Statcast Data to Measure Team Defense

As I’m sure you all know, Statcast allows us to measure the launch angle and velocity for each batted ball. These measurements afford us the ability to estimate precisely the expected wOBA value of every batted ball. Due to the skills of the opposing defense (as well as, admittedly, factors like luck, weather, and ballpark quirks), these estimated wOBA values are often drastically different from their actual values. That is the idea behind Expected Runs Saved (xRS), a metric that I have created to measure team defense. What follows is a discussion of the xRS methodology and some results.

The methodology: The calculation of xRS is actually quite simple. I started by downloading Statcast data from Opening Day through August 29th using Python’s pybaseball module. I then created a dataset consisting of all fair batted balls (excluding home runs) during that time frame. Conveniently, the downloaded data already has the expected wOBA value (based on exit velocity and launch angle), and the actual wOBA value (based on the outcome of the play) for each batted ball. Since we want to penalize teams for making errors, I changed the actual wOBA values for errors from 0 to 0.9 (the value of a single). Then all we have to do is take the average of each metric by team, find the difference, convert that to run values, and we have Expected Runs Saved.

Note that xRS is quite a bit more simplistic than UZR or DRS, as it doesn’t include any of the defensive value derived from keeping baserunners from taking the extra base, preventing steals, turning double plays, etc. While these surely play a role in run prevention, they are less important than converting batted balls into outs, and since I have a full-time job I decided to keep it simple and ignore them.

The results: Let’s start with the most obvious question: which team has the best defense?

It’s the Angels, and it’s not particularly close. While their pitchers have allowed a lot of hard contact (.323 batted-ball xwOBA, 28th in baseball), their actual wOBA on contact is 2nd in baseball at .291, trailing only the Dodgers (.284), who, as Jeff Sullivan recently noted, excel at inducing weak contact.

On the opposite end of the spectrum are the Blue Jays, who have been generally good at generating weak contact (.305 batted-ball xwOBA, 5th in baseball) but terrible at converting those weakly hit balls into outs (.322 batted ball wOBA, 28th in baseball).

In both cases UZR tends to agree, ranking the Angels and Blue Jays 1st and 27th, respectively. Due to (I think) the simplicity of the model, the run values for xRS are quite a bit more extreme than those of either UZR or DRS, but it ranks the teams in generally the same order. At the very least, xRS doesn’t disagree with UZR and DRS much more than the latter two disagree with each other.

Two teams that xRS likes a lot more than UZR and DRS are the Mariners (2nd in xRS, 11th in UZR, 15th in DRS) and Yankees (4th in xRS, 13th in both UZR and DRS). Meanwhile, it dislikes the Dodgers (12th in xRS, 3rd in UZR, 1st in DRS) relative to the other metrics, as well as the Reds (28th in xRS, 5th in UZR, 4th in DRS). Why is this happening? I really don’t know. Could be some defensive components I have left out of xRS, could be ballpark effects, or it could just be that defensive metrics are weird. It remains a mystery. Such is baseball, and such is life.


Introducing XRA: The New Results-Independent Pitching Stat

There are a multitude of ways that we can judge pitchers. Most people look at earned run average to gauge whether a pitcher has been successful, while many old school announcers will still cite a pitcher’s win-loss record. ERA is a nice, easy way of looking at how a pitcher has performed at limiting runs, but it doesn’t come close to telling the whole story. In the early 2000s, Voros McCracken created the idea of Defense Independent Pitching Stats or DIPS, which credited the pitcher only with what he could actually control. Fielding Independent Pitching was born from this theory and only took into account a pitcher’s strikeouts, walks and home runs allowed. It turns out that a pitcher’s home run rate is not terribly consistent, thus xFIP was created by Dave Studeman to normalize the home run aspect of the FIP equation by using the league home run per fly ball rate and the pitcher’s fly ball rate.

In 2015, a new metric was developed by Jonathan Judge, Harry Pavlidis and Dan Turkenkopf called Deserved Run Average or DRA. This new stat attempts to take into account every aspect that the pitcher has control over and control for everything that he does not, thus crediting the pitcher only for the runs that he actually deserves. DRA, however, is still dependent on the result of each batted ball. If the batter hits a ball deep in the gap and it rolls to the wall, the pitcher is charged with a double, but if the center fielder lays out and makes a remarkable catch, the pitcher is credited with an out. When evaluating pitchers, why should it matter whether they have a Gold Glove caliber defender behind them or not? It shouldn’t, and that’s where Expected Run Average comes in.

Expected Run Average or XRA gives pitchers credit for what they actually can control. FIP attempts to do this as well but assumes that pitchers have no control over batted balls. While the pitcher does not control how the fielders interact with the live ball, he does have an impact on the type of contact that he allows. XRA is based on a modified DIPS theory that the pitcher controls three things: whether he strikes the batter out, whether he walks the batter and the exit velocity, launch angle combination off the bat. After the ball leaves the batter’s bat, the play is out of the pitcher’s hands and should no longer have any effect on his statistics. The goal is to figure out a way to measure, independently of the defense and park, how each pitcher performs on balls in play. Since 2015, StatCast has tracked the exit velocity and launch angle of every batted ball in the majors. Each batted ball has a hit probability based on the velocity off of the bat and its trajectory. The probability for extra bases can also be determined. These batted ball probabilities have been linearly weighted for each event including strikeouts and walks to give each player’s xwOBA, which can be found on Baseball Savant. This is the perfect way to look specifically at how well a pitcher has performed on a per plate appearance basis.

Once xwOBA is found, then XRA can be calculated. The first objective is to find the pitcher’s weighted runs below average. To do this, I used the weighted runs above average formula from FanGraphs except I made it negative since fewer runs are better for pitchers.

wRBA = – ((xwOBA – League wOBA) / wOBA Scale) * TBF

For example, Max Scherzer has had a .228 xwOBA so far this season and has faced 487 batters. After finding the league wOBA and wOBA scale numbers at FanGraphs I can plug these numbers into the formula.

– ((.228 – .321) / 1.185) * 487 = 38.22

Max Scherzer has been 38.22 runs better than average so far this season, but now I need to figure out what the average pitcher would do while facing the same number of batters. To find this I need the league runs per plate appearance rate and multiply that number by the number of batters that Scherzer has faced.

League R/PA * TBF = Average Pitcher Runs
.122 * 487 = 59.41

So a league average pitcher would have been expected to surrender 59.41 runs facing the number of batters that Scherzer has so far this season. Now that we know how the average pitcher should have performed we can find the expected number of runs that Scherzer should have surrendered so far this season by subtracting his wRBA of 38.22 from the average pitcher’s runs.

Average Pitcher Runs – Weighted Runs Below Average = Expected Runs
59.41 – 38.22 = 21.19

Based on Scherzer’s xwOBA, he should have only given up 21.19 to this point in the season. If this sounds incredible it’s because this is the lowest mark of any starting pitcher though the first half of the season. Finally, XRA is found by using the RA/9 formula by multiplying the expected number of runs allowed by 9 and then dividing by innings pitched.

(9 * Expected Runs) / Innings Pitched = XRA
(9 * 21.19) / 128.33 = 1.49

Max Scherzer’s XRA of 1.49 is easily the lowest of any starter through the first half. The second best starter has been Chris Sale who has a 2.15 XRA. Of course these names are not surprising as they each started the All Star Game and are both currently the front runners for their leagues’ respective cy young award.

Here is a list of the top ten qualified pitchers:

Pitcher XRA
Max Scherzer 1.49
Chris Sale 2.15
Zack Greinke 2.26
Corey Kluber 2.33
Clayton Kershaw 2.34
Dan Straily 2.87
Lance McCullers 2.89
Chase Anderson 3.11
Luis Severino 3.17
Jeff Samardzija 3.23

And the bottom ten:

Pitcher XRA
Matt Moore 6.58
Kevin Gausman 6.47
Derek Holland 6.32
Matt Cain 6.26
Ricky Nolasco 6.26
Wade Miley 6.17
Johnny Cueto 6.10
Martin Perez 5.97
Jason Hammel 5.95
Jesse Chavez 5.84

Full First Half XRA List

It is interesting to see that three members of the Giants rotation rank in the bottom seven in all of baseball. In fact, AT&T Park is such a pitcher-friendly park that once you park adjust these numbers, Moore, Cain and Cueto become the three worst pitchers in baseball. It’s not surprising then why the Giants are having such a disappointing season.

One measure of a good stat is whether or not it matches your perception. Therefore, while it is interesting to see Dan Straily as one of the best pitchers in baseball and Johnny Cueto as one of the worst, it is much more assuring to see Max Scherzer, Chris Sale and Clayton Kershaw as some of the very best in the sport. The numbers for relievers also reveal how dominant Kenley Jansen and Craig Kimbrel have been. This is all good evidence that XRA is doing what it is supposed to do, accurately displaying how good pitchers have actually been, independent of all other factors.

Another important characteristic of a good stat is how well it correlates from year to year. While ERA is the most simple and popular way to look at pitchers, it is not very consistent. XRA is much more consistent than ERA and FIP and also compares favorably with xFIP. However, it is not as consistent as DRA. DRA controls for so many aspects of the game that it should be expected to be the most consistent. However, being the most predictive or most consistent stat is not necessarily the goal of XRA. The real goal is to show how well the pitcher actually did, and XRA seems to do this remarkably. While not being as consistent as a stat like DRA, the level of consistency is extremely encouraging and puts it right in line with the other run estimators.

XRA is a stat that takes luck, defense, and ballpark dimensions out of the equation. When evaluating a pitcher, he shouldn’t be penalized for giving up a 350-foot pop fly for a home run in Cincinnati while being rewarded for that same pop fly being caught for an easy out in Miami. With XRA, no longer will people have to quibble about BABIP, since it is results-independent and removes all luck from consideration. A ground ball with eyes will now be treated the same whether it squirts through for a single or is tracked down for an out. Pitching ability will no longer need to be measured with an eye on the level of the defense. It takes a good offense, a good pitching staff and a good defense to make a great team, and with XRA we can finally separate all of these important factions.


xHitting (Part 4): 2014 Fantasy Edition!

Welcome to the fourth installment of xHitting!  As always, reader comments and feedback are super encouraged and appreciated.  (Links to parts one, two, and three)

Briefly recapping the method, the gist is to estimate the expected rate of each individual hit type based on a player’s underlying peripherals, and in turn recover all the needed components to compute expected versions of wOBA, OPS, etc.  The only real change to the model since last time is that I now utilize a “hybrid” predicted home run rate, that averages between actual and (raw) predicted home run rate, with the weight given to actual HR rate increasing in the number of plate appearances.  (This is explained in part three, for those curious.)

Perhaps the more exciting change, though, is that this time I actually have results for an ongoing season, which potentially can help for fantasy purposes.  (Not that most readers need my help necessarily.)  Related to fantasy usage, there were a few requests to see a full spreadsheet of past results (2010-2013 seasons), which I have posted here.  Again feel free to take it or leave it at your leisure.

Note: I collected most of these data at the All-Star Break, so numbers may be a few weeks behind, but they’re still mostly true.  Also, for time considerations I only fetched 2014 stats for qualified leaders.  This even leaves out a few big names, but I couldn’t justify time to fetch every player.

So far, I’ve typically posted the biggest “over-” and “under”-achievers for a given season.  And I suppose I’ll continue that tradition today.  But while these lists are useful for highlighting which players seem most likely to regress, it overlooks another main use of the model, which is to assess the realness of a player’s apparent “breakout” or “decline;” at least in-sample.  (In some cases, the model may think that a player’s breakout is entirely justified, given peripherals, while others it may view more skeptically.)  Thus, today I’ll also post a second list, of players who seem to have taken a pronounced step forward/step back this season, and what the model thinks of their season-to-date performance.

Okay, time for results!  I’ll start with the list of “over-” and “underachievers.”

2014 Underachievers (1st half) 2014 Overachievers (1st half)
Name wOBA xWOBA Diff Name wOBA xWOBA Diff
Jean Segura 0.256 0.305 -0.049 Casey McGehee 0.345 0.277 0.068
Chris Davis 0.306 0.353 -0.047 Yasiel Puig 0.398 0.340 0.058
Mark Teixeira 0.352 0.397 -0.045 Matt Adams 0.376 0.324 0.052
Gerardo Parra 0.289 0.327 -0.038 Mike Trout 0.428 0.381 0.047
Brian McCann 0.298 0.330 -0.032 Marcell Ozuna 0.343 0.300 0.043
Torii Hunter 0.323 0.355 -0.032 Lonnie Chisenhall 0.396 0.359 0.037
Joe Mauer 0.308 0.340 -0.032 Scooter Gennett 0.355 0.320 0.035
Jimmy Rollins 0.320 0.352 -0.032 Marlon Byrd 0.344 0.309 0.035
Brian Roberts 0.304 0.334 -0.030 Giancarlo Stanton 0.397 0.363 0.034
Buster Posey 0.326 0.352 -0.026 Hunter Pence 0.359 0.325 0.034

A general pattern I notice is that, having worked with this model for a while now, there do seem to be players that give the model some trouble and have a disproportionate tendency to appear on this list from year to year.  A few of these players appear on this list… more on that later.

Partly for that reason, I wouldn’t necessarily say to “buy low” the guys on the left, nor “sell high” the guys on the right; although you can if you want.  I won’t address every player, but I have some scattered comments:

  • For readers who prefer OPS, .020 wOBA translates to about .050 OPS, on the margin.
  • .397 predicted for Teixeira?  Not sure where that came from…
  • Poor Segura.  All things considered, I think nobody deserves a big second half more than he does.
  • Whatever happened to Casey McGehee’s power?  The guy once hit 23 home runs in a season, but now has ISO of .073, with surprisingly low fly ball distance.
  • Although Chisenhall’s breakout is not as impressive if you take out what the model thinks is luck, it’s still a pretty impressive improvement.
  • Chris Davis is sort of the reverse of Chisenhall.  Adding back in what the model thinks has been bad luck, he’s still way down from what he did last year, but not nearly as disappointing as he probably has been to many owners thus far.

As mentioned, certain players do seem to be able to over/underperform the model somewhat consistently; the same way we think some pitchers are usually better or worse than their FIP.  With now 4.5 years of data to work with, however, I think I can make educated guesses about which players systematically deviate from the model predictions.  I’ll term this deviation the “player fixed effect.”

(Requiring at least 1000 PA from 2010 through 2014 first half)

Model loves too much Model loves too little
Name Player FE
estimate (wOBA)
Name Player FE
estimate (wOBA)
Brian Roberts -0.033 Wilson Betemit 0.032
Todd Helton -0.026 Brandon Moss 0.032
Jean Segura -0.026 Ryan Sweeney 0.028
Jose Lopez -0.025 Mike Trout 0.027
Mark Teixeira -0.025 Peter Bourjos 0.026
Russell Martin -0.024 Matt Carpenter 0.025
Darwin Barney -0.023 Brandon Belt 0.025
Chris Getz -0.023 Melky Cabrera 0.025
Jimmy Rollins -0.021 Carlos Ruiz 0.024
Jason Bay -0.020 Chris Johnson 0.024

Comments:

  • Again, .020 wOBA is equivalent to about .050 OPS, on the margin.
  • Taking out their apparent fixed effect, Teixeira is only underperforming his xWOBA by about .020, and Brian Roberts is actually doing about par.
  • On the reverse side, Mike Trout’s “adjusted” xWOBA jumps up to .408, where really it probably doesn’t surprise us that he’s outperforming even that, since he’s Mike Trout.  And although Giancarlo Stanton misses the Top 10 cutoff above, his apparent fixed effect of +.022 would be 11th; so his “adjusted” xWOBA is more like .385.
  • Yasiel Puig (.058) would also be on the list of “positive fixed effects” if we relaxed the PA requirement (he has 826 during this time).  And Matt Adams (~.040) might also be well on his way to that list; although he has fewer plate appearances still than Puig.
  • I don’t really have good explanations/know any common themes for players with negative fixed effects.  Maybe readers can help?
  • For Trout, home runs are pretty clearly the area where the model underestimates him.  In any given season (2010-2014), he hits about twice as many HR as the model thinks he should in the “raw” prediction.
  • And Trout’s not the only “HR rate defier,” either; just the most salient.  In general, the model has never done as well with home runs as it does with singles, doubles, and triples.  It seems there are other important determinants of home run hitting that really should be in the model, but currently are not.  Intuitively, I sort of would like velocity and angle of the ball off the bat, but so far have not found a good data source to actually include these.  (Maybe that will change in the coming years as MLBAM releases “Hit F/X” style data?)  Until then, reader suggestions are also super welcome here.

And now, finally, for the other usage: here’s a partial list of players who have taken either a pronounced step forward or back this season, relative to established norms.

2014 “Decliners” 2014 “Improvers”
Name Career wOBA 2014 wOBA 2014 xWOBA Name Career wOBA 2014 wOBA 2014 xWOBA
Nick Swisher 0.352 0.285 0.305 Michael Brantley 0.324 0.394 0.404
Joe Mauer 0.373 0.308 0.340 Lonnie Chisenhall 0.328 0.396 0.359
Allen Craig 0.350 0.289 0.309 Seth Smith* 0.334 0.389 0.356
Billy Butler 0.352 0.300 0.309 Victor Martinez 0.362 0.416 0.422
Evan Longoria 0.365 0.315 0.323 Jonathan Lucroy 0.342 0.383 0.354
Domonic Brown 0.315 0.267 0.267 Anthony Rizzo 0.342 0.382 0.382
Chris Davis 0.351 0.306 0.353 Nelson Cruz 0.356 0.393 0.380
Matt Holliday* 0.385 0.342 0.318 Jose Altuve 0.319 0.356 0.325
Jean Segura 0.299 0.256 0.305 Brian Dozier 0.311 0.344 0.362
David Wright 0.377 0.335 0.305 Kyle Seager 0.334 0.367 0.344
Buster Posey 0.366 0.326 0.352 Dee Gordon 0.297 0.329 0.318
Shin-Soo Choo 0.369 0.333 0.346 Alcides Escobar 0.284 0.312 0.300
Dustin Pedroia 0.356 0.325 0.337 Casey McGehee 0.321 0.345 0.277
Jed Lowrie 0.327 0.297 0.305
Jay Bruce 0.343 0.315 0.326

* – To avoid inflation from Coors Field, for these players I’ve taken the total from 2011-13 seasons only

Comments:

  • At least in-sample, Brantley’s breakout seems to be pretty much entirely justified.  Of course this doesn’t mean that he won’t regress somewhat, but if I were to guess, I’m a little more optimistic than ZiPS and Steamer (which currently project .341 and .333 RoS, respectively).  Similar deal for some others.
  • “Yikes” for Billy Butler and Domonic Brown, whose declines this season seem (at least in-sample) to be entirely justified.
  • I’m not sure why the model dislikes Casey McGehee so much.  Obviously his fly ball distance (mentioned earlier) isn’t doing him any favors, and his .369 first-half BABIP is probably unsustainable.  Still, .277 xWOBA?  Seems harsh.

As with any fantasy advice, don’t take any of this too literally…  Take it or leave it as you see fit.

Lastly, although I hyped this piece from a fantasy perspective, the overall goal remains that I would love to see more work done to de-luck hitter stats, the way people do so often for pitchers.  (FIP for pitchers, and xWOBA or xWRC+ for hitters! Is the dream.)

Reader thoughts on how to improve the model, or requests for players not already mentioned?


xHitting (Part 2): Improved Model, Now with 2013 Leaders/Laggards

Happy holidays, all.  It took me a while, but I finally have the second installment of xHitting ready.  First off, thank you to all those who read/commented on the first piece.  For those who didn’t get a chance to read it, the goal here is to devise luck-neutralized versions of popular hitter stats, like OPS or wOBA.  A main extension over existing xBABIP calculators is that this approach offers an empirical basis to recover slugging and ISO, by estimating each individual hit type.

I’ve returned today with an improved version of the model.  Highlights:

  • One more year of data (now 2010-2013)
  • Now includes batted-ball direction (all player-seasons with at least 100 PA)
  • FB distance now recorded for all player-seasons with at least 100 PA

(There’s no theoretical reason for the 100 PA cutoff, only that I was grabbing some of the new data by hand and couldn’t justify the time to fetch literally every single player.)

I have also relaxed the uniformity of peripherals used for each outcome.  At least one reader asked for this, and after thinking about it a while, I decided I agree more than I disagree.  The main advantage of imposing uniformity was that it ensures the predicted rates (when an outs model is also included) sum to 100%.  But it is true that there are certain interactions or non-linearities that are important for some outcomes, but not others.  Including these where they don’t fully belong has a cost to standard errors/precision, and to intuitive interpretation.  To ensure rates still sum to 100%, there’s no longer an explicit ‘outs’ model; outs are simply assumed to be the remainder.

For those curious, below I display regression results for each outcome and its respective peripherals.  You can otherwise skip below if these are not of direct interest.

(The sample includes all player-years with at least 100 plate appearances between the 2010 and 2013 MLB seasons.  Park factors denote outcome-specific park factors available on FanGraphs.  Robust standard errors, clustered by player, are in parentheses; *** p$<$0.01, ** p$<$0.05, * p$<$0.1)

The new variables seem to help, as each outcome is now modeled more accurately than before (by either R2 or RMSE).  For comparison, here are the R2’s of the original specification:

  • 0.367 for singles rate
  • 0.236 for doubles rate
  • 0.511 for triples rate
  • 0.631 for HR rate

Something else I noticed: for balls that stay “inside the fence,” both pull/opp and actual side of the field matter.  Consider singles: the ball needs to be thrown to 1st base (right side of infield) specifically.  Thus an otherwise-equivalent ball hit to the left side is not the same as one hit to the right side, since the defensive play is harder to make from the left side.  Similarly, hitting the ball to left field is less conducive for triples than hitting the ball to right field.

But hitting the ball to the left side as a lefty is not the same as hitting it there as a righty, since one group is “pulling” while the other group is “slapping.”  The direction x handedness interactions help account for this.

How well do the predicted rates do in forecasting?  For singles, doubles, and triples, the predicted rates do unambiguously better than realized rates in forecasting next season’s rates.  Things are a little less clear for home runs, which I will expand on below.

Although predicted HR rate shows a slight edge in Table 1, the pattern often reverses (for HR only) if you use a different sample restriction — say requiring 300 PA in the preceding season.  (For other outcomes, the qualitative pattern from Table 1 still holds even under alternative sample restrictions.)

So home runs appear to be a potential problem area.  What should we do when we need HR to compute xAVG/xSLG/xOPS/xWOBA, etc.?  Should we:

  1. Use predicted HR anyway?
  2. Use actual HR instead?
  3. Use some combo of actual and predicted HR?

Empirically there is a clear answer for which choice is best.  But before getting to that, let’s take a look at whether predicted home-run rate tells us anything at all in terms of regression.  That is, if you’ve been hitting HR’s above/below your “expected” rate, do you tend to regress toward the prediction?

The answer to this seems to be “yes,” evidenced by the negative coefficient on ‘lagged rate residual’ below.

So, although realized HR rate is sometimes a better standalone forecaster of future home runs, predicted HR rate is still highly useful in predicting regression.  Making use of both, it seems intuitively best to use some combo of actual and predicted HR rate for forecasting.

This does, in fact, seem to be the best option empirically.  And this is true whether your end outcome of interest is AVG, OBP, SLG, ISO, OPS, or wOBA.

Observations:

  • (Option 1 = predicted HR only; Option 2 = actual HR only; Option 3 = combo)
  • Whether you use option 1, 2, or 3, xAVG and xOBP make better forecasters than actual past AVG or OBP
  • Option 1 does not do well for SLG, ISO , OPS, or wOBA
  • ^This was not the case in the previous article, but results to that point had sort of a funky sample, having recorded flyball distance only for a partial list of players
  • Option 2 “saves” things for xOPS and xWOBA, but still isn’t best for SLG or ISO
  • Option 3 makes the predicted version better for any of AVG, OBP, SLG, ISO, OPS, or wOBA

End takeaways:

  • The original premise that you can use “expected hitting,” estimated from peripherals, to remove luck effects and better predict future performance seems to be true; but you might need to make a slight HR adjustment.
  • The main reason I estimate each hit type individually is for the flexibility it offers in subsequent computations.  Whether you want xAVG, xOPS, xWOBA, etc., you have the component pieces that you need.  This would not be true if I estimated just a single xWOBA, and other users prefer xOPS or xISO.
  • A major extension over existing xBABIP methods is that this offers an empirical basis to recover xSLG.  The previous piece actually provides more commentary on this.
  • Natural next steps are to test partial-season performance, and also whether projection systems like ZiPS can make use of the estimated luck residuals to become more accurate.

Finally, I promised to list the leading over- and underachievers for the 2013 season.  By xWOBA, they are as follows:

Overachievers (250+ PA) Underachievers (250+ PA)
Name 2013 wOBA 2013 xWOBA Difference Name 2013 wOBA 2013 xWOBA Difference
Jose Iglesias 0.327 0.259 0.068 Kevin Frandsen 0.286 0.335 -0.049
Yasiel Puig 0.398 0.338 0.060 Alcides Escobar 0.247 0.296 -0.049
Colby Rasmus 0.365 0.315 0.050 Todd Helton 0.322 0.369 -0.047
Ryan Braun 0.370 0.321 0.049 Ryan Hanigan 0.252 0.296 -0.044
Ryan Raburn 0.389 0.344 0.045 Darwin Barney 0.252 0.296 -0.044
Mike Trout 0.423 0.379 0.044 Edwin Encarnacion 0.388 0.429 -0.041
Junior Lake 0.335 0.292 0.043 Josh Rutledge 0.281 0.319 -0.038
Matt Adams 0.365 0.323 0.042 Wilson Ramos 0.337 0.374 -0.037
Justin Maxwell 0.336 0.295 0.041 Yuniesky Betancourt 0.257 0.294 -0.037
Chris Johnson 0.354 0.314 0.040 Brian Roberts 0.309 0.345 -0.036

Comments/suggestions?


xHitting: Going beyond xBABIP (part I)

For a few years, it’s struck me as unusual that pitching and hitting metrics are asymmetric.  If the metrics we use to evaluate one group (FIP or wRC+) are so good, why don’t we use them for the other?

One issue is that we’re not used to evaluating pitchers on an OPS-type basis, and similarly we’re not used to evaluating hitters on an ERA basis.  Fine.  But there’s a bigger issue: Why do pitching metrics put so much more emphasis on the removal of luck?

While most sabermetricians are aware of BABIP, and recognize the pervasive impacts it can have on a batting line, attempts to (precisely) adjust hitter stats for BABIP are surprisingly uncommon.  While there do exist a few xBABIP calculators, these haven’t yet caught on en masse like FIP.  And xBABIP doesn’t appear on player pages in either FanGraphs or Baseball Prospectus.

xBABIP itself isn’t even the end goal.  What you probably really want is xAVG/xOBP/xSLG, etc.  Obtaining these is a bit cumbersome when you need to do the conversions yourself.

Moreover, it strikes me that xBABIP cannot be converted to xSLG without some ad hoc assumptions.  Let’s say you conclude a player would have gained or lost 4 hits under neutral BABIP luck.  What type of hits are those?  All singles?  2 singles and 2 doubles?  1 single, 2 doubles, 1 triple?  The exact composition of hits gained/lost affects SLG.  Or maybe you assume ISO is unaffected by BABIP, but this too is ad hoc.

At least to me, whenever a hitter performs better/worse than expected, we really care to know two things:

  1. Is it driven by BABIP?
  2. If so, what is the luck-neutral level of performance?

As I’ve attempted to illustrate, answering #2 is not so easy under existing methods.  (Nor do people always even attempt to answer it, really.)  Even answering #1 correctly takes a little bit of effort.  (“True talent” BABIP changes with hitting style, so it isn’t always enough just to compare current vs. career BABIP.  And then there are players with insufficient track record for career BABIP to be taken at face value.)

Compare this to pitchers.  When a pitcher posts a surprisingly good/bad ERA, we readily consult FIP/xFIP/SIERA.  Specific values, readily provided on the site.  So why not for hitters?

Here I attempt to help fill this gap.  The approach is to map a hitter’s peripheral performance to an entire distribution of hit outcomes.  These “expected” values of singles, doubles, triples, home runs, and outs, can then be used to computed “expected” versions of AVG, OBP, SLG, OPS, wOBA, etc.

Recovering xAVG and xOBP isn’t that different from current xBABIP-based approaches.  The main extension is that, unlike xBABIP, this provides an empirical basis to recover xSLG, and also xWOBA.

Steps:

  1. Calculate players’ rates of singles, doubles, triples, home runs, and outs among balls in play.  (Unlike some other BABIP settings, I count home runs as “balls in play” to estimate an expected number.)
  2. Regress each rate separately on a common set of peripherals.  You’ll now have predicted rates of each for each player.   (Keeping the explanatory variables common throughout ensures the rates sum to 100%.)
  3. Multiply by the number of balls in play (again counting home runs) to get expected counts of singles, doubles, triples, home runs, and outs.
  4. Use these to compute expected versions of your preferred statistics.

What explanatory peripherals are appropriate?  Initially I’ve used:

  • Line drive rate, ground ball rate, flyball rate, popup rate
  • Speed score
  • Flyball distance (from BaseballHeatMaps.com), to approximate power
  • Speed * ground ball rate
  • Flyball distance * flyball rate

These explanatory variables differ somewhat from those in the xBABIP formula linked earlier.  The main distinctions are adding flyball distance (think Miguel Cabrera vs. Ben Revere) and using Speed score instead of IFH%.  (IFH% already embeds whether the ball went for a hit.  Certainly in-sample this will improve model fit, but it might not be good for out-of-sample use.)

Regression results:

Spd FB Dist/1000 FB Dist missing (Spd*GB%)/1000 (FB Dist*FB%)/10000 LD% GB% FB% IFFB%/100 Pitcher dummy Constant
Singles rate -0.0177 0.0608 0.0111 0.4882 0.0090 -0.0019 -0.0063 -0.0066 -0.0417 -0.6833 0.7296
Doubles rate 0.0076 0.6044 0.1457 -0.1059 -0.0152 -0.0058 -0.0066 -0.0061 -0.0070 -0.6700 0.5235
Triples rate 0.0040 0.0193 0.0057 -0.0279 -0.0019 -0.0077 -0.0077 -0.0077 -0.0010 -0.7695 0.7634
HR rate 0.0018 0.9392 0.2764 -0.0295 0.0283 0.0081 0.0080 0.0085 -0.0127 0.8020 -1.0790
Outs rate 0.0043 -1.6238 -0.4389 -0.3249 -0.0202 0.0073 0.0125 0.0118 0.0624 1.3205 0.0625

Technical notes:

  • These are rates among balls in play (including home runs)
  • Each observation is a player-year (e.g. 2012 Mike Trout)
  • I’ve used 2010-2012 data for these regressions
  • Currently I’ve only grabbed flyball distance for players on the leaderboard at BaseballHeatMaps.  This is usually about 300 players per year, or most of the “everyday regulars.”  (Fear not, Ben Revere/Juan Pierre/etc. are included.)  The remaining cases get an indicator for ‘FB Dist missing.’
  • LD%, GB%, FB%, and IFFB% are coded so that 50% = 50, not 0.50.
  • Pitcher proxy = 1 if LD% + GB% + FB% = 0.  Initially I haven’t thrown out cases of pitcher hitting, nor other instances of limited PA.
  • Notice the interaction terms.  The full impact of GB% depends both on GB% and Speed; the full impact of FB% depends on both FB% and FB distance; etc.  So don’t just look at Speed, GB%, FB%, or FB Distance in isolation.
  • Don’t worry that the coefficients on pitcher proxy “look” a bit funny for HR rate and Outs rate.  (Remember that these cases also have LD%=0, GB%=0, and FB%=0.)  In total the average predicted HR rate for pitchers is 0.01% and their predicted outs rate is 94%.
  • Strictly speaking, these are backwards-looking estimators (as are FIP and its variants), but they might well prove useful in forecasting.

I next calculate xAVG, xOBP, xSLG, xOPS, and xWOBA.  For now, I’ve simply taken BB and K rates as given.  (xBABIP-based approaches seem to do the same, often.)

Early results are promising, as “expected” versions of AVG, OBP, SLG, OPS, and wOBA all outperform their unadjusted versions in predicting next-year performance.  (At least for the years currently covered.)

Which players deviated most from their xWOBA?  Here are the leaders/laggards for 2012, along with their 2013 performance:

Leaders Laggards
Name 2012 wOBA 2012 xWOBA Difference 2013 wOBA Name 2012 wOBA 2012 xWOBA Difference 2013 wOBA
Brandon Moss 0.402 0.311 0.091 0.369 Josh Harrison 0.274 0.355 -0.081 0.307
Giancarlo Stanton 0.405 0.332 0.073 0.368 Ryan Raburn 0.216 0.290 -0.074 0.389
Will Middlebrooks 0.357 0.285 0.072 0.300 Nick Hundley 0.205 0.265 -0.060 0.295
Chris Carter 0.369 0.298 0.071 0.337 Jason Bay 0.240 0.299 -0.059 0.306
John Mayberry 0.303 0.238 0.065 0.298 Eric Hosmer 0.291 0.349 -0.058 0.350
Torii Hunter 0.356 0.293 0.063 0.346 Gerardo Parra 0.317 0.369 -0.052 0.326
Jamey Carroll 0.299 0.244 0.055 0.237 Daniel Descalso 0.278 0.328 -0.050 0.284
Cody Ross 0.345 0.291 0.054 0.326 Jason Kipnis 0.315 0.365 -0.050 0.357
Melky Cabrera 0.387 0.333 0.054 0.303 Rod Barajas 0.272 0.322 -0.050 –
Kendrys Morales 0.339 0.286 0.053 0.342 Cameron Maybin 0.290 0.339 -0.049 0.209

Is performance perfect?  Obviously not.  The model does quite well for some, medium-well for others, and not-so-well for some.  Obviously this is not the end-all solution for xHitting.

Some future work that I have in mind:

  • A still more complete set of hitting peripherals.  I’m thinking of park factors, batted ball direction, and possibly others.
  • Testing partial-season performance
  • Comparing results against projection systems like ZiPS and Steamer

Otherwise, my main hope from this piece is to stimulate greater discussion of evaluating hitters on a luck-neutral basis.  Simply identifying certain players’ stats as being driven by BABIP is not enough; we really should give precise estimates of the underlying level of performance based on peripherals.  We do this for pitchers, after all, with good success.

Above I’ve contributed my two cents for a concrete method to do this.  A major extension to xBABIP-based approaches is that this offers an empirical basis to recover xSLG and xWOBA.  While the model is far from perfect, even in its current form it generates “expected” versions of AVG, OBP, SLG, OPS, and wOBA that outperform their unadjusted versions in predicting subsequent-year performance.  (Not just for leaders/laggards.)

Comments and suggestions are obviously welcome!