Archive for Pitcher

CAIN: Counting a Pitcher’s HR/FB Out-Performance

Dave Cameron recently posted an interesting article about Jhoulys Chacin. It’s all about how Jhoulys Chacin is defying the rules of HR/FB rates. His HR/FB rate this year is a mere 2.8%. Jhoulys Chacin has pitched 120 innings, had 106 fly balls, and allowed just three home runs. Very impressive. But it makes you wonder if there are other pitchers who are maintaining low rates while allowing more fly balls overall. Because while Jhoulys Chacin is obviously benefiting from his HR/FB ratio, it’s possible for a pitcher to have more fly balls while maintaining a slightly higher HR/FB and benefit more. So I invented CAIN, a counting stat to help measure that.

CAIN does not stand for anything. I’m just paying homage to a famous outlier.

CAIN = FB – (9.34 x HR)

To explain, the Fangraphs Glossary says that the league-average fly ball rate is “~9-10% depending on the year”. In fact, of the 91 qualified pitchers in Fangraphs database for 2013, the average HR/FB ratio is 10.7 percent. So there are 9.34 fly balls for every homer. So we can say that for most pitchers, if they had ten homers at this point in the season, they would have about 93.4 fly balls.  Ten homers and 93.4 fly balls would give you a CAIN of exactly 0. Make sense?

Now for what you came here for. Here are the top ten in CAIN this year:

Note that I’m not saying any players might actually be able to sustain their CAIN, I just think it’s an interesting little tidbit, and perhaps a nice follow on to Dave Cameron’s article.

Name Team IP HR FB CAIN HR/FB
Eric Stults Padres 133 8 163 88.3 4.90%
Jhoulys Chacin Rockies 120 3 106 78 2.80%
Bartolo Colon Athletics 135.2 9 161 76.9 5.60%
Travis Wood Cubs 128.1 10 159 65.6 6.30%
Adam Wainwright Cardinals 154.2 6 113 57 5.30%
Bud Norris Astros 119.2 10 150 56.6 6.70%
Lance Lynn Cardinals 122 7 121 55.6 5.80%
Matt Moore Rays 116.1 8 130 55.3 6.20%
Derek Holland Rangers 133.2 9 137 52.9 6.60%
Clayton Kershaw Dodgers 152.1 9 136 51.9 6.60%

And Jhoulys Chacin is not #1. It turns out that Eric Stults is in fact benefiting more from his HR/FB rate outlier this year. Of course, that’s partially happening in Petco. Petco is not Coors.

Name Team IP HR FB CAIN HR/FB
Joe Blanton Angels 116 24 133 -91.2 18.0%
Roberto Hernandez Rays 113.1 18 91 -77.1 19.8%
Jason Marquis Padres 117.2 18 99 -69.1 18.2%
CC Sabathia Yankees 142 23 150 -64.8 15.3%
Chris Tillman Orioles 119.2 21 135 -61.1 15.6%
Ryan Dempster Red Sox 115.2 20 130 -56.8 15.4%
R.A. Dickey Blue Jays 134.2 23 163 -51.8 14.1%
Jeremy Guthrie Royals 126.2 22 155 -50.5 14.2%
Hisashi Iwakuma Mariners 138.1 21 146 -50.1 14.4%
Lucas Harrell Astros 112 15 96 -44.1 15.6%

Poor Joe Blanton. His peripherals aren’t that bad this year. But he’s been posting some pretty high HR/FB rates for the last five years or so. I’ll leave it to someone else to puzzle that out.

After doing this analysis I wanted to know about exceptional seasons in the “UZR era” for pitchers’ CAINs. I am continuing to use 9.34 as the FB/HR value, not adjusted for year. If I was being very scientific I would probably break that constant out for league AND year, but I’m lazy and unpaid. Anyway, here, unsurprisingly, is Matt Cain:

Season Name Team IP HR FB HR/FB CAIN
2011 Matt Cain Giants 221.2 9 246 3.70% 161.94
2007 Chris Young Padres 173 10 243 4.10% 149.6
2002 Jarrod Washburn Angels 206 19 317 6.00% 139.54
2009 Zack Greinke Royals 229.1 11 242 4.50% 139.26
2002 Mark Redman Tigers 203 15 273 5.50% 132.9
2011 Jered Weaver Angels 235.2 20 319 6.30% 132.2
2010 Anibal Sanchez Marlins 195 10 222 4.50% 128.6
2010 Livan Hernandez Nationals 211.2 16 278 5.80% 128.56
2010 Jason Vargas Mariners 192.2 18 295 6.10% 126.88
2007 Matt Cain Giants 200 14 255 5.50% 124.24

So in summary, CAIN is a nice little tool if you are interested in seeing just how much a HR/FB rate is affecting a pitcher’s performance. If anyone can think of a better acronym, like one that actually is an acronym, please leave a comment.


The True Dickey Effect

Most people that try to analyze this Dickey effect tend to group all the pitchers that follow in to one grouping with one ERA and compare to the total ERA of the bullpen or rotation. This is a simplistic and non-descriptive way of analyzing the effect and does not look at the how often the pitchers are pitching not after Dickey.

I decided to determine if there truly is an effect on pitchers’ statistics (ERA, WHIP, K%, BB%) who follow Dickey in relief and the starters of the next game against the same team. I went through every game that Dickey has pitched and recorded the stats (IP, TBF, H, ER, BB, K) of each reliever individually and the stats of the next starting pitcher if the next game was against the same team. I did this for each season. I then took the pitchers’ stats for the whole year and subtracted their stats from their following Dickey stats to have their stats when they did not follow Dickey. I summed the stats for following Dickey and weighted each pitcher based on the batters he faced over the total batters faced after Dickey. I then calculated the rate stats from the total. This weight was then applied to the not after Dickey stats. So for example if Francisco faced 19.11% of batters after Dickey, it was adjusted so that he also faced 19.11% of the batters not after Dickey. This gives an effective way of comparing the statistics and an accurate relationship can be determined. The not after Dickey stats were then summed and the rate stats were calculated as well. The two rate stats after Dickey and not after Dickey were compared using this formula (afterDickeySTAT-notafterDickeySTAT)/notafterDickeySTAT. This tells me how much better or worse relievers or starters did when following Dickey in the form of a percentage.

I then added the stats after Dickey for starters and relievers from all three years and the stats not after Dickey and I applied the same technique of weighting the sample so that if Niese’12 faced 10.9% of all starter batters faced following a Dickey start against the same team, it was adjusted so that he faced 10.9% of the batters faced by starters not after Dickey (only the starters that pitched after Dickey that season). The same technique was used from the year to year technique and a total % for each stat was calculated.

Here is the weighted year by year breakdown of the starters’ statistics following Dickey and a total (- indicates a decrease which is desired for all stats except K%):

2012:
ERA: -46.94%  with 5/5 starters seeing a decrease
WHIP: -16.16% with 4/5 seeing a decrease
K%: 47.04% with 4/5 seeing an increase
BB%: 6.50% with 3/5 seeing a decrease
HR%: -50.53% with 5/5 seeing a decrease
BABIP: -14.08% with 4/5 seeing a decrease
FIP: -25.17% with 5/5 seeing a decrease

2011:
ERA: 17.92%  with 0/3 seeing a decrease
WHIP: -9.63% with 2/3 seeing a decrease
K%: -2.64% with 2/3 seeing an increase
BB%: -15.94% with 2/3 seeing a decrease
HR%: -9.21% with 2/3 seeing a decrease
BABIP: -15.14% with 2/3 seeing a decrease
FIP: -5.58% with 2/3 seeing a decrease

2010:
ERA: -23.82%  with 5/7 seeing a decrease
WHIP: 1.68% with 5/7 seeing a decrease
K%: -22.91% with 1/7 seeing an increase
BB%: -2.34% with 5/7 seeing a decrease
HR%: -43.61% with 5/7 seeing a decrease
BABIP: -3.61% with 4/7 seeing a decrease
FIP: -10.61% with 5/7 seeing a decrease

Total:
ERA: -17.21%  with 10/15 seeing a decrease
WHIP: -8.10% with 11/15 seeing a decrease
K%: -3.38% with 7/15 seeing an increase
BB%: -5.17% with 10/15 seeing a decrease
HR%: -32.96% with 12/15 seeing a decrease
BABIP: -11.04% with 10/15 seeing a decrease
FIP: -13.34% with 12/15 seeing a decrease

So for starters that pitch in games following Dickey against the same team, it can be concluded that there is an effect on ERA, WHIP, BABIP, and FIP and a slight effect on BB% and on K%. There is also a large effect on HR rates which we can attribute the ERA effect to. This also tells us that batters are making worse contact the day after Dickey.

So a starter (like Morrow) who follows Dickey against the same team can expect to see around a 17.2% reduction in his ERA that game compared to if he was not following Dickey against the same opponent. For example if Morrow had a 3.00 ERA in games not after Dickey he can expect a 2.48 ERA in games after Dickey.

So if in a full season where Morrow follows Dickey against the same team 66% of the time (games 2 and 3 of a series) in which he normally would have a 3.00 ERA without Dickey ahead of him, he could expect a 2.66 ERA for the season. This seams to be a significant improvement and would equate to a 7.6 run difference (or 0.8 WAR) over 200 innings.

Here is a year by year breakdown of relievers after Dickey (these are smaller sample sizes so I will not include how many relievers saw an increase or decrease):

2012:
ERA: -25.51%
WHIP: -1.57%
K%: 27.04%
BB%: -49.25%
HR%: -34.66%
BABIP: 30.23%
FIP: -38.34%

2011:
ERA: -17.43%
WHIP: 8.45%
K%: 6.74%
BB%: -5.14%
HR%: 7.34%
BABIP: 9.75%
FIP: -2.05%

2010:
ERA: -2.55%
WHIP: 7.69%
K%: -9.28%
BB%: 10.84%
HR%: 2.11%
BABIP: 4.23%
FIP: 9.43%

Total:
ERA: -16.61%
WHIP: 5.38%
K%: 7.50%
BB%: -12.65%
HR%: -8.53%
BABIP: 13.38%
FIP: -10.40%

As expected there was a good effect on the relievers’ ERA, FIP, K%, and BB%, but the WHIP and BABIP were affected negatively. This tells me that the batters were more free swinging after just seeing Dickey (more hits, less walks, more strikeouts).

So in a season where there are 55 IP after Dickey in games (like in 2012) there would be a 16.6% reduction in runs given up in those 55 innings. If the bullpen’s ERA is 4.20 without Dickey it can be expected to be 3.50 after Dickey. Over 55 IP this difference would save 4.3 runs (or 0.4 WAR).

Combine this with the saved starter runs and you get 11.9 runs saved or (1.2 WAR). This is Dickey’s underlying value with the team that he creates by baffling hitters. This 1.2 WAR is if Morrow has a 3.00 ERA normally and the bullpen has a 4.00 ERA. If Morrow normally had a 4.00 ERA than his ERA would reduce to 3.54 over the season with 10.2 runs saved for 200 innings (1.0 WAR) and if the bullpen has a 4.00 ERA normally as well, 4.1 runs would be saved there, equating to 14.3 runs saved or a 1.4 WAR over a season.


BABIP and Innings Pitched (Plus, Explaining Popups)

In my last post on explaining pitchers’ BABIPs by way of their batted ball rates, I was very careful to say that it was applicable in the long run, as it’s hard to be accurate over a short number of innings pitched, due to all the “noise” in BABIP (Batting Average on Balls In Play).  I only used pitchers with a qualifying number of innings pitched (IP) in the calculations, for that reason.  After writing the post, I did some messing around with the data, to find out just how much of an effect IP had on the predictability of BABIP.

Hold on to your propeller beanies, fellow stat geeks: the correlation between xBABIP and BABIP went from 0.805 when the minimum IP was set to 1500, to 0.632 at a 200 IP minimum, down to 0.518 at 50 IP.  OK, maybe it’s not that surprising.  Still, I thought I’d better show you how confident you can be in my xBABIP formula’s accuracy when you take the pitcher’s innings pitched into account.

The formula, again: xBABIP = 0.4*LD% – 0.6*FB%*IFFB% + 0.235

And remember, that formula is primarily meant to be a backwards-looking estimator of “true,” defense-neutral BABIP.  My next article will (probably) discuss another formula I’ve come up with that’s more forward-looking.

Read the rest of this entry »


Projecting BABIP Using Batted Ball Data

Hi everybody, this is my first post here. Today, I’ll be sharing some of my BABIP research with you. There will probably be several more in the near future.

Now, I don’t know about you, but Voros McCracken’s famous thesis stating that pitchers have practically no control over their batting average on balls in play (BABIP) always seemed counterintuitive to me, ever since I heard it about 10 years ago. Basically, my thought this whole time was that if an Average Joe were pitching to an MLB lineup, the hitters would rarely be fooled by the pitches, and would be crushing most of them, making it very tough on the fielders. Think Home Run Derby (only with a lot more walks). Now, the worst MLB pitcher is a lot closer in ability to the best pitcher than he is to an Average Joe, but there still must be a spectrum amongst MLB pitchers relating to their BABIP, I figured. After crunching some numbers, I have to say that intuition hasn’t completely failed me.

This is going to be a long article, so if you want the main point right here, right now, it’s this: in the long run, about 40% or more of the difference in pitchers’ BABIPs can be explained by two factors that are independent of their team’s defense: how often batters hit infield fly balls and line drives off of them. It is more difficult to predict on a yearly basis, where I can only say that those factors can predict over 22% of the difference. Line drive rates are fairly inconsistent, but pop fly rates are among the more predictable pitching stats (about as much as K/BB). I’ll explain the formula at the very end of the article.

Read the rest of this entry »