Archive for ERA

The Impact of Defensive Prowess on a Pitcher’s Earned Runs Average

EXECUTIVE SUMMARY

  • This study attempts to determine how much the fielders’ prowess, measured by the metric UZR (Ultimate Zone Range), affects a pitcher’s Earned Runs Average.
  • The data used for the regression (collected from FanGraphs.com) includes collective ERA, BABIP, HR/9, BB/9, K/9 and UZR for every Major League Baseball team for the past three years.
  • ERA (Earned Runs Average) is the amount of earned runs a pitcher allows per nine innings pitched. BABIP (Batting Average per Balls in Play) is the batting average against any given pitcher, but only including the at bats where the hitter puts the ball in play. HR/9 is home runs allowed per nine innings pitched. BB/9 is walks allowed per nine innings pitched. K/9 is batter struck out per nine innings pitched. UZR (Ultimate Zone Range) is a widely used metric to evaluate defense. It summarizes how many runs any given fielder saved or gave up during a season compared to the league average in that position.
  • The model passed the F-test, the adjusted “R” squared came out at 91.2 percent and every one of the independent variables passed their respective t-test.
  • The model tested negative for both Multicollinearity (using Variance Inflation Factors) and Heteroskedasticity (using the second version of the White’s test).
  • The regression equation looks like this: ERA = -2.55 – 0.187 K/9 + 0.413 BB/9 +16.9 BABIP + 1.72 HR/9 – 0.00157 UZR. Even though the independent variable UZR has a low coefficient, it definitely affects a pitcher’s ERA, and in the way it was suspected. As the UZR goes up the ERA goes down.

INTRODUCTION

Since Bill James started to write about baseball in the late 1970’s and started to defy the traditional stats used to evaluate players, hundreds of baseball fans have tried to follow his footsteps creating new ways to evaluate players and defy the existing ones. One of the stats that has been brought to light lately is Earned Runs Average (ERA).

According to several baseball analysts ERA is not an efficient way to evaluate how good or bad a pitcher performs. The rationale behind this thinking is pretty simple; ERA is the amount of earned runs that any given pitcher allows per nine innings pitched, but the pitcher is not always 100 percent responsible for every earned run allowed. Sometimes, a fielder’s lack of defensive prowess will allow hitters to reach base safely (I am not talking about errors), and when it happens, rather often, those hits will translate into earned runs, thus affecting the pitcher’s ERA.

One of the metrics that has been used to determine any given fielder’s prowess is UZR (Ultimate Zone Range). UZR compiles data on the outfielders arms, fielder range and errors and summarizes the amount of runs those fielders saved or gave up during a season compared to the league average in that position. Using that metric along with other metrics that affect the ERA, we can answer the question “How much does defensive prowess impacts a pitcher’s ERA?”

If in fact defensive prowess affects ERA, we could also determine how much it affects it. With that kind of information, cost-effective teams (Tampa Bay Rays and Oakland Athletics) can help improve their pitching staff without investing heavily on new pitchers.

DATA

The unit of observation for this study is one Major League Baseball team. And the number of observations is 90. Currently, there are 30 Major League Baseball teams, so data was collected for the past three Major League Baseball seasons. So the time period covered goes from 2010 to 2012, including both seasons.

The dependent variable used in this project was Earned Runs Average, and the independent variables are as follow:

  • BABIP: Batting average per balls in play
  • HR/9: Homeruns allowed per nine innings pitched
  • BB/9: Walks allowed per nine innings pitched
  • K/9: Hitters struck out per nine innings pitched
  • UZR: Runs saved or given up by any given fielder during a season

All the data for this study is cross-sectional because all the observations have been collected at the same point of time.

All the data for this study was collected from the baseball website FanGraphs.com. FanGraphs is a widely known source of baseball stats and news, but the data they publish on their website is collected by another company called Baseball Info Solutions.

REGRESSION ESTIMATIONS

            Regression Analysis: ERA versus BABIP, HR/9, BB/9, K/9 and UZR

The regression equation is

ERA = – 2.55 – 0.187 K/9 + 0.413 BB/9 + 16.9 BABIP + 1.72 HR/9 – 0.00157 UZR

 

Predictor       Coef         SE Coef              T           P             VIF

Constant      -2.5474     0.5594        -4.55    0.000

K/9              -0.18718    0.02428     -7.71     0.000    1.099

BB/9            0.41261     0.04671        8.83     0.000    1.052

BABIP          16.914        1.876             9.02     0.000     1.741

HR/9            1.7222       0.1105          15.58    0.000    1.180

UZR        -0.0015743  0.0006219  -2.53  0.013       1.669

 

S = 0.133650   R-Sq = 91.7%   R-Sq(adj) = 91.2%

 

Analysis of Variance

 

Source                  DF        SS            MS              F             P

Regression          5     16.5663   3.3133   185.49   0.000

Residual Error  84   1.5004     0.0179

Total                     89   18.0668

The first step used to evaluate the model was the F-test, and since the model has a p-value less than 0.05, it is safe to say that the model passed the F-test. The adjusted “R” squared for the model was 91.2 percent, which means that 91.2 percent of the variation in ERA is explained by at least one of the independent variables used in this model. The method used to evaluate the relevance of the independent variables was the t-test, and each one of them, as mentioned earlier, had a p-value below 0.05, so in conclusion, they all passed the t-test. The p-value for K/9, BB/9, BABIP and HR/9 was 0.000 for each one of them, and the p-value for UZR was 0.013.

MODEL ESTIMATION SEQUENCE

  1. Correct functional form: To check for correct functional form, each one of the independent variables was plotted against the dependent variable. The scatter plots that resulted from this check show a linear relationship between each one of the independent variables and the dependent variable.
  2. Test for Heteroskedasticity: The data for this study is cross-sectional, so it was necessary to test for Heteroskedasticity, and such test was conducted by the second version of White’s test. To do so, the residuals for the original regression were stored. Those squared residuals were regressed against the Independent variables and the independent variables squared. After running the regression, an the F-test was applied to it and since the p-value was over 0.05, it can be concluded that the regression fails the F-test, therefore Heteroskedasticity does not exist in the initial model.
  3. Multicollinearity: This model also tested for Multicollinearity and it is done by using the correlation matrix and the Variance Inflation Factors, observed in the initial regression.
    1. Since none of the VIF’s is larger than 10, it can be concluded that Multicollinearity does not exist and the p-values from the t-tests can be trusted.
    2. A correlation matrix was calculated using all the independent variables but since every one of them passed the t-test, none will be dropped from the model.
  • K/9: p-value (0.000), VIF (1.099), rho (0.252)
  • BB/9: p-value (0.000), VIF (1.052), rho (0.195)
  • BABIP: p-value (0.000), VIF (1.741), rho(0.604)
  • UZR: p-value (0.013), VIF (1.669), rho (0.604)
  1. Drop any irrelevant variable from the model: Since all the independent variables in this model are relevant, none of them will be dropped from the model.

FINAL MODEL

The final model is exactly the same as the initial model because the it passed the F-test, all of the independent variables passed their t-tests and neither Heteroskedasticity or Multicollinearity are present in the model, so it was not necessary to run another regression or drop any variable.

COEFFICIENT INTERPRETATION

  • K/9: When the team strikes out one extra batter per nine innings, the team’s ERA should go down by 0.187 runs per nine innings holding everything else constant.
  • BB/9: When the team walks one extra batter per nine innings, the team’s ERA should go up by 0.413 runs per nine innings holding everything else constant.
  • BABIP: If every time a batter puts the ball in play he records a hit, the ERA will go up by 16.9 runs per nine innings. This variable is hard to explain since it will never go up by 1, it will go up or down depending on how many hits the team allows in any given number of at-bats where the batter puts the ball in play. For example, if a team averages eight hits every 27 outs, the BABIP will be 0.296 throughout the entire season. Taking into account that every batter put the ball in play (no strikeouts). The expected increase in ERA given a 0.296 BABIP during a season, and holding everything else constant, would be 5.00.
  • HR/9: When the team allows one more homerun per nine innings, ERA should go up by 1.72 runs per nine innings holding everything else constant.
  • UZR: When the team saves one extra run defensively, ERA should go down by 0.00157 runs per nine innings holding everything else constant.

SUMMARY

The null hypothesis for this project stated that defensive prowess didn’t affect ERA, but the results showed otherwise, so it is safe to reject the null hypothesis. Defensive prowess appears to affect ERA although in a small scale. This might not seem like much, but cost-effective teams like the Rays and Athletics can acquire premium defensive players at a much cheaper cost than a premium pitcher, and although they won’t be “game changers,” they will definitely improve the team’s ERA.

Baseball is a game of numbers, and these numbers don’t lie. A good defender will help his team save runs; a lot of good defenders will help their team save multitude of runs. Is this enough to get to the postseason or win a World Series? Absolutely not, but it has been proven already that finding edges in the game, as little as they might be, will help a team in the long run. The findings in this study are a concise proof that taking advantage of defense is an edge that can be exploited for the betterment of the organization.


The R.A. Dickey Effect – 2013 Edition

It is widely talked about by announcers and baseball fans alike, that knuckleball pitchers can throw hitters off their game and leave them in funks for days. Some managers even sit certain players to avoid this effect. I decided to analyze to determine if there really is an effect and what its value is. R.A. Dickey is the main knuckleballer in the game today, and he is a special breed with the extra velocity he has.

Most people that try to analyze this Dickey effect tend to group all the pitchers that follow in to one grouping with one ERA and compare to the total ERA of the bullpen or rotation. This is a simplistic and non-descriptive way of analyzing the effect and does not look at the how often the pitchers are pitching not after Dickey.

Dickey's Dancing Knuckleball
Dickey’s Dancing Knuckleball (@DShep25)

I decided to determine if there truly is an effect on pitchers’ statistics (ERA, WHIP, K%, BB%, HR%, and FIP) who follow Dickey in relief and the starters of the next game against the same team. I went through every game that Dickey has pitched and recorded the stats (IP, TBF, H, ER, BB, K) of each reliever individually and the stats of the next starting pitcher, if the next game was against the same team. I did this for each season. I then took the pitchers’ stats for the whole year and subtracted their stats from their following Dickey stats to have their stats when they did not follow Dickey. I summed the stats for following Dickey and weighted each pitcher based on the batters he faced over the total batters faced after Dickey. I then calculated the rate stats from the total. This weight was then applied to the not after Dickey stats. So for example if Janssen faced 19.11% of batters after Dickey, it was adjusted so that he also faced 19.11% of the batters not after Dickey. This gives an effective way of comparing the statistics and an accurate relationship can be determined. The not after Dickey stats were then summed and the rate stats were calculated as well. The two rate stats after Dickey and not after Dickey were compared using this formula (afterDickeySTAT-notafterDickeySTAT)/notafterDickeySTAT. This tells me how much better or worse relievers or starters did when following Dickey in the form of a percentage.

I then added the stats after Dickey for starters and relievers from all four years and the stats not after Dickey and I applied the same technique of weighting the sample so that if Niese’12 faced 10.9% of all starter batters faced following a Dickey start against the same team, it was adjusted so that he faced 10.9% of the batters faced by starters not after Dickey (only the starters that pitched after Dickey that season). The same technique was used from the year to year technique and a total % for each stat was calculated.

The most important stat to look at is FIP. This gives a more accurate value of the effect. Also make note of the BABIP and ERA, and you can decide for yourself if the BABIP is just luck, or actually better/worse contact. Normally I would regress the results based on BABIP and HR/FB, but FIP does not include BABIP and I do not have the fly ball numbers.

The size of the sample was also included, aD means after Dickey and naD is not after Dickey. Here are the results for starters following Dickey against the same team.

Dickey Starters

It can be concluded that starters after Dickey see an improvement across the board. Like I said, it is probably better to use FIP rather than ERA. Starters see an approximate 18.9% decrease in their FIP when they follow Dickey over the past 4 years. So assuming 130 IP are pitched after Dickey by a league average set of pitchers (~4.00 FIP), this would decrease their FIP to around 3.25. 130 IP was selected assuming ⅔ of starter innings (200) against the same team. Over 130 IP this would be a 10.8 run difference or around 1.1 WAR! This is amazingly significant and appears to be coming mainly from a reduction in HR%. If we regress the HR% down to -10% (seems more than fair), this would reduce the FIP reduction down to around 7%. A 7% reduction would reduce a 4.00 FIP down to 3.72, and save 4.0 runs or 0.4 WAR.

Here are the numbers for relievers following Dickey in the same game.

Dickey Bullpen

Relievers see a more consistent improvement in the FIP components (K, BB, HR) between each other (11.4, 8.1, 4.9). FIP was reduced 10.3%. Assuming 65 IP (in between 2012 and 2013) innings after Dickey of an average bullpen (or slightly above average, since Dickey will likely have setup men and closers after him) with a 3.75 FIP, FIP would get reduced to 3.36 and save 3 runs or 0.3 WAR.

Combining the un-regressed results, by having pitchers pitch after him, Dickey would contribute around 1.4 WAR over a full season. If you assume the effect is just 10% reduction in FIP for both groups, this number comes down to around 0.9 WAR, which is not crazy to think at all based off the results. I can say with great confidence, that if Dickey pitches over 200 innings again next year, he will contribute above 1.0 WAR just from baffling hitters for the next guys. If we take the un-regressed 1.4 WAR and add it to his 2013 WAR (2.0) we get 3.4 WAR, if we add in his defence (7 DRS), we get 4.1 WAR. Even though we all were disappointed with Dickey’s season, with the effect he provides and his defence, he is still all-star calibre.

Just for fun, lets apply this to his 2012. He had 4.5 WAR in 2012, add on the 1.4 and his 6 DRS we get 6.5 WAR, wow! Using his RA9 WAR (6.2) instead (commonly used for knucklers instead of fWAR) we get 7.6 WAR! That’s Miguel Cabrera value! We can’t include his DRS when using RA9 WAR though, as it should already be incorporated.

This effect may even be applied further, relievers may (and likely do) get a boost the following day as well as starters. Assuming it is the same boost, that’s around another 2.5 runs or 0.25 WAR. Maybe the second day after Dickey also sees a boost? (A lot smaller sample size since Dickey would have to pitch first game of series). We could assume the effect is cut in half the next day, and that’d still be another 2 runs (90 IP of starters and relievers). So under these assumptions, Dickey could effectively have a 1.8 WAR after effect over a full season! This WAR is not easy to place, however, and cannot just be added onto the teams WAR, it is hidden among all the other pitchers’ WARs (just like catcher framing).

You may be disappointed with Dickey’s 2013, but he is still well worth his money. He is projected for 2.8 WAR next year by Steamer, and adding on the 1.4 WAR Dickey Effect and his defence, he could be projected to really have a true underlying value of almost 5 WAR. That is well worth the $12.5M he will earn in 2014.

For more of my articles, head over to Breaking Blue where we give a sabermetric view on the Blue Jays, and MLB. Follow on twitter @BreakingBlueMLB and follow me directly @CCBreakingBlue.


TIPS, A New ERA Estimator

FIP, xFIP, SIERA are all very good ERA estimators, and their predictability is well documented. It is well known that SIERA is the best ERA estimator over samples that occur from season to season, followed very close by xFIP, with FIP lagging behind. FIP is best at showing actual performance though, because is uses all real events (K, BB, HR). Skill is commonly best attributed to either xFIP or SIERA. ERA is also well known to be the worst metric at predicting future performance, unless the sample size is very large <500IP with the pitcher remaining in the same or a very similar pitching environment.

FIP, xFIP, and SIERA are supposed to be Defense Independent Metrics, and they are. Well, they are independent of field defense, but there is one small error in the claim of defense independent. K’s and BB’s are not completely independent of defense. Catcher pitch framing plays a role in K’s and BB’s. Catchers can be good or bad at changing balls into strikes and this affects K’s and BB’s. Umpire randomness and umpire bias also play a role in K’s and BB’s. It is unknown how much of getting umpires to call more strikes is a skill for a pitcher or not. Some pitchers are consistent at getting more strike calls (Buehrle, Janssen) or less strike calls (Dickey, Delabar), but for most pitchers it is very random (especially in small sample sizes). For example Jason Grilli was in the top 5% in 2013 but was in bottom 10% in 2012.

I wanted to come up with another ERA estimator that eliminates catcher framing, umpire randomness and bias, and eliminates defense. I took the sample of pitchers who have pitched at least 200IP since 2008 (N=410) and analyze how different statistics that meet this criteria affect ERA-. I used ERA- since it takes out park factors and adjusts for the changes in the league from year to year. I looked at the plate discipline pitchf/x numbers (O-Swing, Z-Swing, O-Contact, Z-Contact, Swing, Contact, Zone, SwStr), the six different results based off plate discipline (zone or o-zone, swing or looking, contact or miss for ZSC%, ZSM%, ZL%, OSC%, OSM%, OL%), and batted ball profiles (GB%, LD%, FB%, IFFB%). *Please note that all plate discipline data is PitchF/X data, not the the other plate discipline on FanGraphs, this is important as the values differ*

The stats with very little to absolutely no correlation (R^2<0.01) were: Z-Swing%, Zone%, OSC%, ZSC%, ZL% (was a bit surprised as this would/should be looking strike%), GB%, and FB%. These guys are obviously a no-no to include in my estimator.

The stats with little correlation (R^2<0.1) were: Swing%, LD%, and IFFB%. I shouldn’t use these either.

O-Contact% (0.17), Z-Contact%, (.302), Contact% (.319), OSM% (0.206), and ZSM% (.248) are all obviously directly related to SwStr%. SwStr% had the highest correlation (.345) out of any of these stats. There is obviously no need to include all of the sub stats when I can just use SwStr%. SwStr% will be used in my metric.

OL% (0.105) is an obvious component of O-Swing% (0.192). O-Swing had the second highest correlation of the metrics (other than the components of SwStr%). I will use it as well. The theory behind using O-Swing% is that when the batter doesn’t swing it should almost always be a ball (which is bad), but when the batter swings, there are a two outcomes, a swing and miss (which is a for sure strike) or contact. Intuitively, you could say that contact on pitches outside the zone is not as harmful to pitchers as pitches inside the zone, as the batter should get worse contact. This is partially supported in the lower R^2 for O-Contact% to Z-Contact%. It is more harmful for a pitcher to have a batter make contact on a pitch in the zone, than a pitch out of the zone. This is why O-Swing is important and I will use it.

Using just SwStr% and O-Swing%, I came up with a formula to estimate (with the help of Excel) ERA-. I ran this formula through different samples and different tests, but it just didn’t come up with the results I was looking for. The standard deviation was way too small compared to the other estimators, and the root mean square error was just not good enough for predicting future ERA-.

I did not expect/want this estimator to be more predictive than xFIP or SIERA. This is because xFIP and SIERA have more environmental impacts in them that remain fairly constant. K% is always a better predictor of future K% than any xK% that you can come up with. Same with BB% Why? Probably because the environment of catcher framing, and umpire bias remain somewhat constant. Also (just speculation) pitchers who have good control can throw a pitch well out of the zone when they are ahead in the count, just to try and get the batter to swing or to “set-up” a pitch. They would get minus points for this from O-Swing, depending on how far the pitch is off the plate, but it may not affect their K% or BB% if they come back and still strike out the batter.

So I didn’t expect my statistic to be more predictive, but the standard deviation coupled with not that great of RMSE (was still better than ERA and FIP with a min of 40IP), caused me to be unhappy with my stat.

I then started to think about if there were any stats that were only dependent on the reaction between batter an pitcher that are skill based that FanGraphs does not have readily available? I started thinking about foul balls and wondered if foul ball rates were skill based and if they were related to ERA-. I then calculated the number of foul balls that each pitcher had induced. To find this I subtracted BIP (balls in play or FB+GB+LD+BU+IFFB) from contacts (Contact%*Swing%*Pitches). This gave me the number of fouls. I then calculated the rates of fouls/pitch and foul/contacts and compared these to ERA-. Foul/Contact or what I’m calling Foul%, had an R^2 of .239. That’s 2nd to only SwStr%. This got me excited, but I needed to know if Foul% is skill based and see what else it correlates with.

This article from 2008 gave me some insight into Foul%. Foul% correlates well to K% (obviously) and to BB% (negative relationship), since a foul is a strike. Foul% had some correlation to SwStr%, this is good as it means pitchers who are good at getting whiffs are also usually good at getting fouls. Foul% also had some correlation to FB% and GB%. The more fouls you give up, the more fly balls you give up (and less GB). This doesn’t matter however, as GB% and FB% had no correlation to ERA-. Foul% is also fairly repeatable year to year as evidenced in the article, so it is a skill. I will come up with a new estimator that includes Foul% as well.

I decided to use O-Looking% instead of O-Swing%, just to get a value that has a positive relationship to ERA (more O-looking means higher ERA), because SwStr% and O-Swing are negatively related. O-Looking is just the opposite of O-Swing and is calculated as (1 – O-Swing%).

The formula that Excel and I came up with is this: (I am calling the metric TIPS, for True Independent Pitching Skill)

TIPS = 6.5*O-Looking(PitchF/x)% – 9.5*SwStr% – 5.25*Foul% + C

C is a constant that changes from year to year to adjust to the ERA scale (to make an average TIPS = average ERA). For 2013 this constant was 2.68.

I converted this to TIPS- to better analyze the statistic. FIP, xFIP, and SIERA were also converted to FIP-, xFIP-, and SIERA-. I took all pitchers’ seasons from 2008-2013 to analyze. The sample varied in IP from 0.1 IP to 253 IP. I found the following season’s ERA- for each pitcher if they pitched more than 20 IP the next year and eliminated any huge outliers. Here were the results with no min IP. RMSE is root mean square error (smaller is better), AVG is the average difference (smaller is better), R^2 is self explanatory (larger is better), and SD is the standard deviation.

N=2316 ERA- FIP- xFIP- SIERA- TIPS-
RMSE 77.005 51.647 43.650 43.453 40.767
AVG 43.941 34.444 30.956 30.835 30.153
R^2 0.021 0.045 0.068 0.147 0.169
SD 69.581 38.654 24.689 24.669 15.751

Wow TIPS- beats everyone! But why? Most likely because I have included small samples and TIPS- is based off per pitch, as opposed to per batter (SIERA) or per inning (xFIP and FIP). There are far more pitches than AB or IP so TIPS will stabilize very fast. Let’s eliminate small sample sizes and look again.

Min 40 IP
N=1619 ERA- FIP- xFIP- SIERA- TIPS-
RMS 40.641 36.214 34.962 35.634 35.287
AVG 29.998 26.770 25.660 25.835 26.115
R^2 0.063 0.105 0.120 0.131 0.101
SD 26.980 19.811 15.075 17.316 13.843

 

Min 100 IP
N=654 ERA- FIP- xFIP- SIERA- TIPS-
RMSE 32.270 29.949 29.082 28.848 29.298
AVGE 24.294 22.283 21.482 21.351 22.038
R^2 0.080 0.118 0.143 0.145 0.095
SD 20.580 16.025 12.286 12.630 10.985

Now, TIPS is beaten out by xFIP and SIERA, but beats ERA and and is close to FIP (wins in RMSE, loses in R^2). This is what I expected, as I explained earlier K% and BB% are always better at predicting future K% and BB% and they are included in SIERA and xFIP. SIERA and xFIP take more concrete events (K, BB, GB) than TIPS. I didn’t want to beat these estimators, but instead wanted a estimator that is independent of everything except for pitcher-batter reaction.

TIPS won when there was no IP limit, so it obviously is the best to use in smaller sample sizes, but when is it better than xFIP and SIERA, and where does it start falling behind? I plotted the RMSE for my entire sample at each IP. Theoretically these should be an inverse relationship. After 150 IP it gets a bit iffy, as most of my sample is less than 100 IP. I’m more interested in IP under 100 anyhow.

Orange is TIPS, Blue is ERA, Red is FIP, Green is xFIP, and Purple is SIERA. If you can’t see xFIP, it’s because it is directly underneath SIERA (they are almost identical). This is roughly what the graph should look like to 100 IP:

Looking at the graph, at what IPs is TIPS better than predicting future ERA than xFIP and SIERA? It appears to be from 0 IP to around 70 IP.

Here is the graph for 1/RMSE (higher R^2). Higher number is better. This is the most accurate graph as the relationship should be inverse.

The 70-80 IP mark is clear here as well.

I’m not suggesting my estimator is better than xFIP or SIERA, it isn’t in samples over 75 IP, but I think it is, and can be, a very powerful tool. Most bullpen pitchers stay under 75 IP in a season. This means that my unnamed estimator would be very useful for bullpen arms in predicting future ERA. I also believe and feel that my estimator is a very good indicator of the raw skill of a pitcher. It would probably be even more predictive if we had robo-umps that eliminated umpire bias and randomness and pitch framing.

2013 TIPS Leaders with 100+IP

Name ERA FIP xFIP SIERA TIPS
Cole Hamels 3.6 3.26 3.44 3.48 3.02
Matt Harvey 2.27 2 2.63 2.71 3.09
Anibal Sanchez 2.57 2.39 2.91 3.1 3.23
Yu Darvish 2.83 3.28 2.84 2.83 3.23
Homer Bailey 3.49 3.31 3.34 3.39 3.26
Clayton Kershaw 1.83 2.39 2.88 3.06 3.32
Francisco Liriano 3.02 2.92 3.12 3.5 3.34
Max Scherzer 2.9 2.74 3.16 2.98 3.36
Felix Hernandez 3.04 2.61 2.66 2.84 3.37
Jose Fernandez 2.19 2.73 3.08 3.22 3.42

 

And Leaders from 40IP to 100IP

Name ERA FIP xFIP SIERA TIPS
Koji Uehara 1.09 1.61 2.08 1.36 1.87
Aroldis Chapman 2.54 2.47 2.07 1.73 2.03
Greg Holland 1.21 1.36 1.68 1.5 2.29
Jason Grilli 2.7 1.97 2.21 1.79 2.36
Trevor Rosenthal 2.63 1.91 2.34 1.93 2.42
Ernesto Frieri 3.8 3.72 3.49 2.7 2.45
Paco Rodriguez 2.32 3.08 2.92 2.65 2.50
Kenley Jansen 1.88 1.99 2.06 1.62 2.50
Glen Perkins 2.3 2.49 2.61 2.19 2.54
Edward Mujica 2.78 3.71 3.53 3.25 2.54

 


The True Dickey Effect

Most people that try to analyze this Dickey effect tend to group all the pitchers that follow in to one grouping with one ERA and compare to the total ERA of the bullpen or rotation. This is a simplistic and non-descriptive way of analyzing the effect and does not look at the how often the pitchers are pitching not after Dickey.

I decided to determine if there truly is an effect on pitchers’ statistics (ERA, WHIP, K%, BB%) who follow Dickey in relief and the starters of the next game against the same team. I went through every game that Dickey has pitched and recorded the stats (IP, TBF, H, ER, BB, K) of each reliever individually and the stats of the next starting pitcher if the next game was against the same team. I did this for each season. I then took the pitchers’ stats for the whole year and subtracted their stats from their following Dickey stats to have their stats when they did not follow Dickey. I summed the stats for following Dickey and weighted each pitcher based on the batters he faced over the total batters faced after Dickey. I then calculated the rate stats from the total. This weight was then applied to the not after Dickey stats. So for example if Francisco faced 19.11% of batters after Dickey, it was adjusted so that he also faced 19.11% of the batters not after Dickey. This gives an effective way of comparing the statistics and an accurate relationship can be determined. The not after Dickey stats were then summed and the rate stats were calculated as well. The two rate stats after Dickey and not after Dickey were compared using this formula (afterDickeySTAT-notafterDickeySTAT)/notafterDickeySTAT. This tells me how much better or worse relievers or starters did when following Dickey in the form of a percentage.

I then added the stats after Dickey for starters and relievers from all three years and the stats not after Dickey and I applied the same technique of weighting the sample so that if Niese’12 faced 10.9% of all starter batters faced following a Dickey start against the same team, it was adjusted so that he faced 10.9% of the batters faced by starters not after Dickey (only the starters that pitched after Dickey that season). The same technique was used from the year to year technique and a total % for each stat was calculated.

Here is the weighted year by year breakdown of the starters’ statistics following Dickey and a total (- indicates a decrease which is desired for all stats except K%):

2012:
ERA: -46.94%  with 5/5 starters seeing a decrease
WHIP: -16.16% with 4/5 seeing a decrease
K%: 47.04% with 4/5 seeing an increase
BB%: 6.50% with 3/5 seeing a decrease
HR%: -50.53% with 5/5 seeing a decrease
BABIP: -14.08% with 4/5 seeing a decrease
FIP: -25.17% with 5/5 seeing a decrease

2011:
ERA: 17.92%  with 0/3 seeing a decrease
WHIP: -9.63% with 2/3 seeing a decrease
K%: -2.64% with 2/3 seeing an increase
BB%: -15.94% with 2/3 seeing a decrease
HR%: -9.21% with 2/3 seeing a decrease
BABIP: -15.14% with 2/3 seeing a decrease
FIP: -5.58% with 2/3 seeing a decrease

2010:
ERA: -23.82%  with 5/7 seeing a decrease
WHIP: 1.68% with 5/7 seeing a decrease
K%: -22.91% with 1/7 seeing an increase
BB%: -2.34% with 5/7 seeing a decrease
HR%: -43.61% with 5/7 seeing a decrease
BABIP: -3.61% with 4/7 seeing a decrease
FIP: -10.61% with 5/7 seeing a decrease

Total:
ERA: -17.21%  with 10/15 seeing a decrease
WHIP: -8.10% with 11/15 seeing a decrease
K%: -3.38% with 7/15 seeing an increase
BB%: -5.17% with 10/15 seeing a decrease
HR%: -32.96% with 12/15 seeing a decrease
BABIP: -11.04% with 10/15 seeing a decrease
FIP: -13.34% with 12/15 seeing a decrease

So for starters that pitch in games following Dickey against the same team, it can be concluded that there is an effect on ERA, WHIP, BABIP, and FIP and a slight effect on BB% and on K%. There is also a large effect on HR rates which we can attribute the ERA effect to. This also tells us that batters are making worse contact the day after Dickey.

So a starter (like Morrow) who follows Dickey against the same team can expect to see around a 17.2% reduction in his ERA that game compared to if he was not following Dickey against the same opponent. For example if Morrow had a 3.00 ERA in games not after Dickey he can expect a 2.48 ERA in games after Dickey.

So if in a full season where Morrow follows Dickey against the same team 66% of the time (games 2 and 3 of a series) in which he normally would have a 3.00 ERA without Dickey ahead of him, he could expect a 2.66 ERA for the season. This seams to be a significant improvement and would equate to a 7.6 run difference (or 0.8 WAR) over 200 innings.

Here is a year by year breakdown of relievers after Dickey (these are smaller sample sizes so I will not include how many relievers saw an increase or decrease):

2012:
ERA: -25.51%
WHIP: -1.57%
K%: 27.04%
BB%: -49.25%
HR%: -34.66%
BABIP: 30.23%
FIP: -38.34%

2011:
ERA: -17.43%
WHIP: 8.45%
K%: 6.74%
BB%: -5.14%
HR%: 7.34%
BABIP: 9.75%
FIP: -2.05%

2010:
ERA: -2.55%
WHIP: 7.69%
K%: -9.28%
BB%: 10.84%
HR%: 2.11%
BABIP: 4.23%
FIP: 9.43%

Total:
ERA: -16.61%
WHIP: 5.38%
K%: 7.50%
BB%: -12.65%
HR%: -8.53%
BABIP: 13.38%
FIP: -10.40%

As expected there was a good effect on the relievers’ ERA, FIP, K%, and BB%, but the WHIP and BABIP were affected negatively. This tells me that the batters were more free swinging after just seeing Dickey (more hits, less walks, more strikeouts).

So in a season where there are 55 IP after Dickey in games (like in 2012) there would be a 16.6% reduction in runs given up in those 55 innings. If the bullpen’s ERA is 4.20 without Dickey it can be expected to be 3.50 after Dickey. Over 55 IP this difference would save 4.3 runs (or 0.4 WAR).

Combine this with the saved starter runs and you get 11.9 runs saved or (1.2 WAR). This is Dickey’s underlying value with the team that he creates by baffling hitters. This 1.2 WAR is if Morrow has a 3.00 ERA normally and the bullpen has a 4.00 ERA. If Morrow normally had a 4.00 ERA than his ERA would reduce to 3.54 over the season with 10.2 runs saved for 200 innings (1.0 WAR) and if the bullpen has a 4.00 ERA normally as well, 4.1 runs would be saved there, equating to 14.3 runs saved or a 1.4 WAR over a season.


Johnny B. Goode

Controlling the run game, pitcher fielding and ERA

Run & Glove

Johnny Cueto has been mocking his peripherals ever since his big league debut.  For the most part FIP serves as a terrific gauge for pitcher performance, but in 2011 Cueto made FIP look like a heart monitor trying to explain the weather.  On what most consider a separate note, base runners have a healthy and robust fear of Cueto’s pickoff move, which is one of the best in the show.

FIP measures outcomes a pitcher can control (home runs, walks and strikeouts) and chalks the rest up to random variation.  Studies have shown that stolen bases contribute relatively little to run creation and perhaps on that basis the ability to control the run game has generally been ignored or deemed overrated.

It is difficult, however, to ignore the six runs Cueto saved the Reds via his contributions to controlling the run game in 2012.  By contrast, A.J. Burnett’s inability to control runners cost the Pirates four runs. The typical scale is that 10 runs amount to one team win – and teams will pay about $5 million per win.

Acknowledging run game control cannot fully explain how Cueto has routinely outperformed his peripherals, just as it cannot wholly explain Pittsburgh’s inability to keep pace with Cincinnati in the NL Central last season.  It does, however, get us closer.

Incorporating a pitcher’s fielding ability proved of comparative importance in explaining and predicting performance.  Here we’ll turn to Mark Buehrle, whose glove has saved four runs per year since 2004, and among fellow hurlers the fast-working lefty has been one of the decade’s most steadily superb fielders.  FIP underestimated Buehrle in eight of the past nine seasons, slighting his ERA by an average of .30 per year over that span.

Numbers

 The numbers indicate that a pitcher’s defense and ability to control the run game should both be considered in assessing and forecasting the pitcher’s value.

Focusing on seasons in which pitchers hurled 100 same-league (AL or NL) innings from 2003-2012 (n=1400), I ran a multiple linear regression to create a formula (“MBRA”) incorporating run control (rSB) and pitcher fielding (rPM) on top of line drive and infield fly ball percentages (credit to BABIP guru Steve Staude) and a regressed take on FIP.

MBRA = (55.25*HR + 14.05*BB – 8.57*K)/TBF – .041*rPM – .056*rSB + (5.71*LD – 8.27*IFFB)/(LD+GB+FB) + 2.34 

Correlation

Mean Absolute Error

MBRA

.7750

.4570

FIP

.7647

.4697

BERA

.7477

.4922

tERA

.7472

.5616

MBRAT

.7216

.5394

xFIP

.6451

.5649

SIERA

.6290

.5768

 

MBRA is engineered to properly credit pitchers who can field and control the run game.  When I subtracted MBRA from FIP to locate the pitcher-seasons that benifitted most from my formula, I was encouraged seeing Buehrle show up twice in the top ten, and five times in the top 100 (again, that is out of 1400).

Next, I looked at seasons in which pitchers threw 100 same-league innings in consecutive seasons from 2003 to 2012 (n=791).  This time I ran a regression to create a model suited to predict a pitcher’s ERA based on his previous year’s statistics.

MBRAT = (20.12*HR + 7.13*BB – 6.7*K)/TBF -.025*rPM -.034*rSB + 2.37*ZC% + 2.22

 

Correlation

Mean Absolute Error

MBRAT

.4526

.6498

SBERA

.4398

.6582

BERA

.4347

.6634

xFIP

.4220

.6803

MBRA

.4198

.6987

FIP

.4162

.7024

ERA

.3630

.7920

MBRAT stands tall on the lofty pinnacle of public forward-looking ERA estimators, and if you factor in the percentage at which pitchers throw over the edge of the plate (EDGE%) its correlation jumps even higher (.4621).  Unfortunately, I only have Edge% data from 2008 to 2012 (n=362) and cannot yet justify its inclusion.

On Deck

I will create expectations for pitchers with fewer innings pitched and convert my findings to a WAR measure that may serve as a middle-ground between fWAR and rWAR.  I also stumbled on a potentially significant relationship between pick-off attempts and strand rates that may work its way into future formulas.


Introducing BERA: Another ERA Estimator to Confuse You All

Coming up with BERA… like its [almost] namesake might say, it was 90% mental, and the other half was physical.  OK, maybe he’d say something more along the lines of “what the hell is this…” but that’s beside the point.    By BERA, I mean BABIP-estimating ERA (or something like that… maybe one of you can come up with something fancier).  It’s an ERA estimator that’s along the lines of SIERA, only it’s simpler, and—dare I say—better.

You know, I started out not knowing where I was going, so I was worried I might not get there.  As you may recall, I’ve been pondering pitcher BABIPs for a little while here (see article 1 and article 2), and whereas my focus thus far had been on explaining big-picture, long-term BABIP stuff in terms of batted ball data, one question that remained was how well this info could be used to predict future BABIPs.  After monkeying around with answering that question, though, I saw that SIERA’s BABIP component could be improved upon, so I set to work in coming up with BERA.  In doing so, I definitely piggybacked off of FIP and a little of what SIERA had already done.  You can observe a lot just by watching, you know.   I’m also a believer in “less is more” (except for when it comes to the size of my articles, obviously), so I tried to go for the best compromise of simplicity and accuracy that I could.

Read the rest of this entry »