Archive for Pitcher

Reworking and Improving the Outcome Machine

This post was inspired by a couple of articles that I remembered reading from Jonah Pemstein back in 2014. The intention of those posts was to predict the result of any given batter/pitcher matchup, dubbed the “Outcome Machine.” Have you ever wondered what the probability Mike Trout strikes out when he steps into the box against Justin Verlander? Of course, there are variables that are specific to any plate appearance (umpires/situation/stadium/etc.) that are harder to quantify, but it set out to predict the outcome in a vacuum. Trout vs. Verlander and nothing else (For the record, in 2020, I would estimate the answer is about 27.5%).

Being able to predict the outcomes in sports would take most of the fun out of being a spectator, sure, but I still found myself coming back to those articles. While reading and re-reading in an attempt to understand the logic and fool around with the equations, I came to a few questions of my own:

  • With all of the hubbub of juiced balls and increased launch angles, do equations that were based on data from 2003-13 still apply to the game today?
  • The regression equations were composed of the at-bat result and the stats of the batter and pitcher from the same year. This stuck out to me as an issue because it means the player’s performance later in the season, say in July, influences the prediction of an at-bat in May, and to a lesser extent, the result of that specific at-bat is already baked into that season’s performance. Shouldn’t you use data exclusively before a given at-bat to predict the outcome? Hindsight is 20/20, after all.

Eventually curiosity got the best of me and I decided to emulate the original exercise. Before I really start to nerd out on the inner workings, you can find this iteration of the Outcome Machine as a Google Sheet here. You can either select a pitcher/batter combination through the dropdown or hard key in the rates in a custom, hypothetical matchup below that. League average is set by default to projections for 2020 but can be updated as desired in the custom matchup. I would note that the preset statistics in this tool are total projections for 2020 but not broken out into L/R splits, as to my knowledge that data is currently behind a paywall. Read the rest of this entry »


How Long Before Things Go Bad?

Spring is a time for optimism, in baseball and in life. Teams are starting to think about their opening day starters and more broadly, their starting rotations. Some rotations look “set” while some have a “battle for the 5th spot”. Some are toying with the idea of a 6-man rotation.

But here’s the thing: we know that (almost) every team will end up using a 6-man rotation, whether they like it or not. Eventually, your favorite team will need to call in reinforcements. This can happen because of poor performance or injury. But hey, we’ll cross that bridge when we come to it, right?

… when do you think we might come to it?

We know, as do those in charge that teams use something like 11 starters per year (in 2017: 11.3). In a six-month season, how long does it take before the first reinforcements arrive?

Cumulative Starters Used, 2017

In a few words, not very long. Some pitchers have injuries, some get moved to the bullpen, some sent to the minors. Either way, at least one of them will be gone pretty soon, so don’t name the puppy.

Of course, fate comes at different paces. In 2017, the Cardinals didn’t use a sixth starter until June 13th. And even then, Marco Gonzales only pitched because they had a double-header. In contrast, Junior Guerra, the Brewers’ opening day starter, was injured that same opening day. He wouldn’t pitch in the majors for another seven weeks (and it turns out, not very well either).

Half of teams used a sixth starter before April 25th. 90% of teams used a sixth starter before their 50th game.

Some of those sixth starters, along with their full-season WAR: Alex Wood (3.4), Mat Latos (-0.3), Mike Clevinger (2.2), Mike Pelfrey (-1.0).

We know that teams need depth. Not only that, but life comes at you fast.

Data: Baseball Savant


Are the Mets in Rebuilding Mode Once Again?

The Mets are the talk of the town…for all the wrong reasons. They currently sit at a 31-41 record and are 12 games behind the Washington Nationals in the NL East, which as of now seems to be theirs for the taking. The Mets boast one of the worst bullpens in the majors and have been plagued by injuries as well as underperformance from the bulk of their lineup. With the results of this season, many are beginning to wonder if it’s time to turn the page on this current pack of Mets players, many of whom were on the 2015 team that lost to the feisty Kansas City Royals in the World Series. I will attempt to go group by group in an effort to determine whether or not the Mets should begin a new rebuilding process, the most dreaded phrase in sports.

Starting with the outfield, Yoenis Cespedes is locked in for three more years in his current contract. It’s understandable why the Mets were looking to sign him in the offseason based on his performance in 2015 and 2016. However, injuries and poor performance have contributed to the current record that the Mets have. Cespedes still won’t lose his spot in left. Curtis Granderson, due to his age, will most likely not be re-signed, as well as Jay Bruce who, if he is not traded before the deadline, will most certainly test free agency. Juan Lagares has been injury-prone the last couple years but the one piece of good news is that Michael Conforto has seen a resurgence since coming back from Triple-A Las Vegas. Also, one of their top prospects, Brandon Nimmo, should receive regular playing time in the outfield, if not this season, then definitely in 2018.

Next, we have the infield, which has been decimated by injuries. Neil Walker and Asdrubal Cabrera have struggled through injuries (and who knows if/when David Wright will ever step on a baseball field again). Jose Reyes and Lucas Duda have mightily underperformed. The good news for the Mets is that Cabrera, Walker, and Reyes will be gone after the season, which means that the infield can get much younger. Top prospects Dominic Smith and Amed Rosario will be September call-ups and, if all goes well, can be regulars in the lineup next year. T.J. Rivera and Wilmer Flores have proven to be reliable pieces in the lineup. Despite some injuries from Flores, he has made up for it with his versatility in both the field and in the lineup, giving manager Terry Collins options to choose from. While Flores and Rivera may not be long-term solutions, they are the best options that the Mets have at the moment. As far as catching is concerned, Travis d’Arnaud is probably the Mets’ best option right now, although he has severely underperformed since being traded to them. The Mets should try to get another catcher in free agency.

Finally, the best pitching staff is a huge question mark, but also a big concern among scouts. Matt Harvey clearly no longer has any interest in remaining with the team and Noah Syndergaard, Zack Wheeler, and Steven Matz are just injuries waiting to happen. Even Jacob deGrom, who has been I believe the best starter this season, has a history of arm injuries that makes Mets front-office personnel nervous. Even Robert Gsellman and Seth Lugo are recovering from injuries sustained during this season. The bullpen has been just as bad. The bullpen so far has logged 257 innings to the tune of a 4.97 ERA. Not to mention they have not had a reliable closer since Jeurys Familia has been both suspended and injured this season, and the rest of the bullpen outside of Addison Reed and Jerry Blevins has been downright horrendous.

Overall, the Mets need to begin the next phase of the rebuilding process. With aging veterans and current players underperforming, it’s clear that the time for a championship has come and gone for this group. The Mets need to get younger and it starts with the old addition-by-subtraction technique. By dumping aging veterans with big contracts, the Mets will be able to allocate their resources and maybe pick up some pieces in free agency while simultaneously giving their top prospects playing time and allowing them to develop. As the great Cosmo Kramer once said on Seinfeld, “I think it’s time that we shut down and re-tool.”


Tyler Wilson and His Five Plus Pitches

Let me preface this article by saying that I watch A LOT of baseball.  I also have an extensive analytical background and am always analyzing baseball stats looking for value in players.  Last week, I was watching an Orioles game and the starting pitcher was a player I have never heard of.  His name is Tyler Wilson.  While watching the game, I was very impressed with his overall make-up and the confidence he displayed in each one of his pitches.  Many times what separates a pitcher from being able to start at the big-league level versus being destined for the bullpen is the ability to throw multiple pitches.  The ability to throw each of those pitches effectively, however, can be what separates a good starting pitcher from a great starting pitcher.  The more I watched of Wilson, the more intrigued I became about his future outlook, and the more motivated I became to write this article.  (I went back and watched all of Wilson’s starts this year before writing this article.)

To give you a little background, Tyler Wilson has never been an elite prospect.  He attended college at the University of Virginia, where he was overlooked by fellow staff-mate, and future 1st round pick, Danny Hultzen.  Wilson was drafted by the Orioles in the 10th round of the 2011 MLB Draft.  Ever since being drafted, he has quietly excelled at every level.  He doesn’t have the dominant strikeout numbers that you look for in pitching prospects, which is a big reason he has gone overlooked for much of his career.

After climbing his way through the organizational ladder, Wilson made his major league debut with the Orioles last year and eventually made the team this year out of spring training.  Although he made the team in a bullpen role, early season injuries to the Orioles pitching staff opened up an opportunity and Wilson has really taken advantage of it.  Enough of the background though.  Let’s move on to what I saw while actually watching him pitch.

Tyler Wilson features a cutter and a two-seam fastball.  Each of these pitches sit in the 89-91 mph range and both show a great amount of movement.  The cutter is most effective against right-handed batters when thrown on the outside portion of the plate.  Check out the video below to watch him fool Kansas City Royals outfielder Lorenzo Cain with three straight cutters:

He essentially gave Cain, a very good hitter, three of the exact same pitches in a row…and Cain couldn’t touch them.  In every start this year, Wilson has pounded the outside corner with this cutter and has had fantastic results.  Don’t think by any means though that he is a one trick pony.  As soon as you start to expect that cutter on the outside corner, Wilson will come right back in on you with a two-seam fastball:

Look at the horizontal movement on that pitch!  Absolutely filthy!  Wilson has showed a ton of confidence in both of those pitches so far this season as he uses them to pound both sides of the strike zone and his command of them has been exceptional.  He is not afraid to throw them in any count and they are equally effective vs both left-handed and right-handed batters.

While his fastballs both seemed to be plus pitches upon first glance, I started to have thoughts that this guy might be for real as soon as he started throwing his curveball.  Wilson’s breaking ball sits in the 77-79 mph range.  I was astonished by how well he was able to locate his curve and the amount of movement on each and every one he threw.  Watch him send White Sox slugger Jose Abreu down swinging in the video below:

Abreu had no chance.  In his most recent start against the Twins, Wilson’s curve looked even better.  Check out the one he threw to Byung-Ho Park:

Both of those pitches came in a 2-2 count.  Many pitchers are scared to throw a breaking ball in a 2-2 count, especially to players with plus power such as Abreu and Park.  If you miss your target, two things can happen.  One — you leave the ball up in the zone and it gets hit out of the stadium.  Two — you throw it in the dirt; the hitter lays off; and now you have to pitch to this slugger with a full count.  Wilson isn’t scared to throw his curveball in any count and that is what makes him so dangerous.  You never know when to expect it, but at the same time you have to expect that he can throw it at any moment.

The last pitch in Wilson’s arsenal is his changeup.  This pitch has a ton of downward movement and produces a lot of groundballs.  While there were many better examples that I could have shown you of his change-up in action, I wanted to show one of his bad ones.  Even when he missed his target, the batter was still fooled by the amount of movement on this pitch.  Check out the following pitch to Royals SS Alcides Escobar:

The catcher set up down in the zone and Wilson clearly misses his target.  Luckily it didn’t seem to matter as the pitch had an insane amount of horizontal movement, running in on Escobar and jamming him.

Take a look at the chart below, showing the vertical and horizontal movement on each of Wilson’s pitches:

Tyler Wilson Movement

The middle portion of this chart is empty.  All five of his pitches have a tremendous amount of movement, and none of them move in the same direction.  The fact that he is able to command each of these pitches so well and keep hitters guessing with which one will come next is the reason why he has had so much success.  A big reason why hitters are having trouble guessing his pitches is because of how well Wilson is able to repeat his delivery.  The chart below shows Wilson’s release point for each type of pitch:

Tyler Wilson Release Point
As you can see, his release point is almost identical with all five of his pitches.  At this point, I have watched all of his starts from this season and was very impressed.   I then decided to do some research and was immediately impressed with stats such as his career BB rate and low WHIP, but wanted to dig further.  I began to look through the PITCHf/x data because I was curious to see how effective each of his pitches actually were.  Based on the PITCHf/x value metric, all of his pitches so far this year have graded as above average.  If you are not familiar with the PITCHf/x value scale, someone who has a fastball ranking of zero means that he possesses an average fastball.  Any value above zero means that pitch is above average.  Obviously the higher the number, the better the pitch.  The same goes for negative numbers and pitches being below average.  See the table below for the breakdown of Wilson’s arsenal:

Screen Shot 2016-05-15 at 1.19.17 AM

Based on the above values, the change-up has been Wilson’s most valuable pitch this season with his curveball close behind.  Obviously it is very early in the season and we are working with a small sample size…but that doesn’t mean we can’t have fun!  While doing this research, I set out the goal to find every starting pitcher who throws five or more above-average pitches.  Below is the list of players who fit that description:

Screen Shot 2016-05-15 at 1.41.09 AM
IP = Innings Pitched
FA = Fastball
FT = Two-Seam Fastball
FC = Cut Fastball
SI = Sinker
SL = Slider
CU = Curveball
CH = Change-up
KC = Knuckle Curveball
EP = Eephus

There are only five pitchers who have thrown five or more pitches above average so far this season!  Wilson is in great company, as the other four pitchers are all All-Star-caliber players and borderline household names.  Being that this is such a small sample size, I decided to look back at last year’s stats to see how many players fit this description over a full season.  Using the same parameters and setting the minimum IP to 100, the following table was produced:

Screen Shot 2016-05-15 at 2.05.17 AM

Once again, the names on this list are some of the top pitchers in baseball.  A few of these pitchers have a pitch that graded out as below average, but since they had five or more different pitches all individually grade as above average, they made the final cut.

As you can see, it is very rare to have a pitcher who has five legitimate plus pitches.  I am very interested to see if Tyler Wilson can maintain these results over the course of a full season, and I really hope he is given the opportunity to do so.  If he continues to pitch the way he has been, the Orioles will have no choice but to leave him in the rotation.  Although he has had limited success, Wilson has struggled in each of his starts when facing the lineup the third time around.  This could be due to the fact that he is still in the process of being stretched out from his bullpen role.  When in the bullpen, you don’t have to prepare to face the same hitter three times.  I am hopeful that once he is fully stretched out and back into his starter mentality, he will be able to make the necessary adjustments and continue to throw all of his pitches with confidence.  If he can continue to make quality pitches as he faces the lineup for a third time, I believe Tyler Wilson has the chance to become a very special pitcher.

Memorable quotes I heard during the TV broadcasts:

“Everyone thinks that I pitch with a chip on my shoulder but I really don’t.  I just go out and compete.  I don’t think of it that way.” – Tyler Wilson

“I think he understands himself.  He can maintain his game-plan throughout the game.  He’s going to keep us in the game and give us a chance to win.  What more can you ask for?” – Pitching Coach Dave Wallace

“I love that he can make the ball run in and then cut away.  He pitches to both sides of the plate.  Not a lot of young pitchers can do that.” – Manager Buck Showalter

…no Buck, not a lot of young pitchers can do that.

Twitter – @mtamburri922


How Game Theory Is Applied to Pitch Optimization

The timeless struggle between pitcher and batter is one of dominance — who holds it and how. Both players use a repertoire of techniques to adapt to each other’s strategies in order to gain advantage, thereby winning the at-bat and, ultimately, the game.

These strategies can rely on everything from experience to data. In fact, baseball players rely heavily on data analytics in order to tell them how they’re swinging their bats, how well they’ll do in college, how they’ll perform at Wrigley versus Miller.

Big data has been used in baseball for decades — as early as the 60s. Bill James, however, was the first prominent sabermetrician, writing about the field in his Bill James Baseball Abstracts during the 80s. Sabermetrics are used to measure in-game performance and are often used by teams to prospect players.

Baseball fans familiar with sabermetrics, the A’s, and Brad Pitt have likely seen Moneyball, the Hollywood adaptation of Michael Lewis’ book. The book told the story of As manager Billy Beane’s use of sabermetrics to amass a winning team.

Sabermetrics is one way baseball teams use big data to leverage game theory in baseball — on a team-wide scale. However, by leveraging their data through the concepts of game theory on a smaller scale, baseball teams can help their men on mound out-duel those at the plate.

Game theory studies strategic decision making, not just in sports or games, but in any situation in which a decision must be made against another decision maker. In other words, it is the study of conflict.

Game theory uses mathematical models to analyze decisions. Most sports are zero-sum games, in which the decisions of one player (or team) will have a direct effect on the opposing player (or team). This creates an equilibrium which is known as the Nash equilibrium, named for the mathematician John Forbes Nash. What this means is that if a team scores a run, it is usually at the expense of the opposing team — likely based on an error by a fielder or a hit off a pitcher.

In the case of pitching, game theory — especially the use of the Nash equilibrium — can be used to predict pitch optimization for strategic purposes. Neil Paine of FiveThirtyEight advocates using big data and sabermetrics to analyze each pitch in a hurler’s armory, then cultivating the pitcher’s equilibrium — the perfect blend of pitches that will result in the highest number of strikeouts, etc.

Paine has gone so far as to create his own formula, the Nash Score, to predict which pitcher should throw which pitches in order to outwit batters.

In perfect game theory, the Nash equilibrium states that each game player uses a mix of strategies that is so effective, neither has incentive to change strategies. For pitchers, Paine’s Nash Score uses their data to find the optimal combination of pitches to combat batters, including frequency.

Paine does point out that creating this kind of equilibrium in baseball can be detrimental to a pitcher. He is, after all, playing against another human being who is just as capable of using game theory to adapt strategies to upset the equilibrium.

If a pitcher’s fastball is his best, and his Nash Score shows that he should be using it more often, savvy hitters are going to notice. “ . . . In time, the fastball will lose its effectiveness if it’s not balanced against, say, a change-up — even if the fastball is a far better pitch on paper,” writes Paine.

In this case, a mixed strategy is the best — in game theory, mixed strategies are best used when a player intends to keep his opponent guessing. Though pitch optimization using Paine’s Nash Score could lead to efficiency, allowing pitchers to throw fewer pitches for more innings, it could also lead to batters adapting much quicker to patterns, thus negating all the work.


Stephen Strasburg Is Better Than You Think

To a casual baseball fan, Stephen Strasburg’s numbers are not pretty. The owner of a 4.76 ERA and a 1.38 WHIP, Strasburg is clearly having the worst season of his career. But how bad has he been, really? Not as bad as you think. Take a look at these 2015 stats:

Player A: 3.48 xFIP, 22.8 K%, 5.5 BB%
Player B: 3.31 xFIP, 24.1 K%, 5.3 BB%
Player C: 3.18 xFIP, 24.9 K%, 6.0 BB%

Player A is none other than Johny Cueto, recently traded to the Kansas City Royals. 12th in ERA among qualified pitchers, Cueto is widely considered among the best, and perhaps deservedly so with five straight years of a sub-3 ERA. While he has consistently outperformed the above metrics, they are still indicative of general pitcher performance and should not be overlooked when comparing the quality of different pitchers.

Player B actually has the fifth lowest ERA among qualified pitchers and was also traded at the deadline. He’s been one of the most reliable pitchers over the past five years and has been an ace on every staff for which he’s pitched. Player B is David Price.

Player C is obviously Stephen Strasburg, and as you can see, his peripheral stats stack up against the best in the game. In addition to these 2 players, Strasburg also compares positively to others like Sonny Gray and Scott Kazmir, both of whom have better ERAs but a worse xFIP, K%, and BB%.  Strasburg is pitching like an ace, and xFIP shows that, so why have his results been so poor?

Well, first of all, there’s his .345 BABIP. Not only is this high compared to the league average (.296), it’s well above his career mark of .302. Considering he’s not giving up any more line drives or hard contact than usual, his BABIP should fall back to around the .300 mark and bring his ERA down with it.

Not only is his BABIP at an all-time high, his LOB% is at an all-time low. Currently at 65.3%, it figures to inch back up to his career 73.2% mark, or at least to the league average of 72.4%. Considering his strikeouts have not dropped off, there’s no reason for his drop on LOB%, and it can simply be chalked up to bad luck, something that he’s had plenty of this year.

Looking at these stats, there’s nothing that suggests Strasburg is anything but unlucky. However, as Jeff Sullivan pointed out here, Strasburg’s problem could stem from the injury he suffered in the spring. He had apparently adjusted his mechanics to compensate for the discomfort, and even though it appears as though he has fixed this, it’s possible that when pitching from the stretch and in higher leverage situations, he returns to this altered motion by default. When looking at the difference in Strasburg’s stats between pitching from the windup and the stretch, this is what we see:

K% xFIP
Bases Empty 30.1 2.73
Runners on Base 17.0 3.98

Evidently, this claim has some ground. Strasburg is clearly having some problems with runners on base, particularly in striking batters out. Before we deal with the strikeout numbers, let’s take a look to make sure that he’s not just getting killed during the at bats that don’t end in strikeouts.

GB/FB Batted Ball Velocity (mph) Hard Hit % Infield Hit %
Bases Empty .98 89 29.7 4.5
Runners on Base 2.05 88 28.7 12.2

Strasburg is actually generating more ground balls and weaker contact with runners on base. His infield hit percentage is triple what it is when the bases are empty, something that can be attributed to luck. With such weak contact, it’s safe to say this isn’t the problem. So it must be the strikeouts. If we take a look at his whiff rates, the results are intriguing:

2010-2014 2015
Bases Empty 20.1% 17.5%
Runners On Base 17.9% 8.6%

OK, so there’s definitely a problem here. With runners on base, he’s only whiffing batters at half the rate he’s done previously in his career, as well as half the rate that he does with the bases empty. So what’s the issue? Well, it’s not his pitch velocity:

4 Seam 2 Seam Changeup Curve Slider
Bases Empty 95.1 mph 95.4 mph 88.4 mph 81.3 mph 86.7 mph
Runners on Base 95.2 mph 94.9 mph 88.0 mph 81.5 mph 87.2 mph

Strasburg’s average velocity with runners on base is 91.5 mph, compared to 91.0 mph with the bases empty, so he’s actually throwing the ball harder when there’s runners on base. That can’t be the problem. He’s also not walking a significant amount more batters when there are runners on base, so it’s not like he’s sacrificing control for increased speed.

Without any numbers to provide a reason, it appears Strasburg’s struggles when striking out batters with runners on base are either based purely in luck or are completely mental. This is not necessarily a good thing, as we have no idea if or when he will sort it out. With his skill, Strasburg has the potential to be one of the best in the game. He just needs to get out of his own head, and maybe get just a little bit luckier.


Why is Bronson Arroyo Still Throwing a Changeup?

I respect the change-up. As a pitcher myself, I know how difficult it is to throw a good one (thus I don’t). It’s not the most glamorous pitch in baseball, but certainly an effective one if executed correctly. Plus, what constitutes a good off-speed offering reads like a laundry list of mechanical and ball path attributes that have to be repeated over and over again. Proper grip on the baseball. Delivery and arm speed must be identical to the fastball. Velocity needs to be lower than the fastball. The ball should move (ideally both horizontally and vertically) and spotted in a good location. And lastly, there’s the intangible pitching IQ of understanding when to throw it.

The Diamondbacks Bronson Arroyo and his change-up seem to be missing a majority of these qualities… but for some reason he continues to throw the darned thing. 16% of the time in 2013, in fact, and already almost 18% of the time this season. I’m baffled.

Now, of course I can’t know what’s going on in his head (although if someone can point me to an all-encompassing Pitching IQ metric I would be more than happy to apply it). And I also can’t measure his arm velocity at release. So I can’t quantify all of his deficiencies. But there is, fortunately, hard numerical and visual data showing he’s lacking the necessary skills to throw a change-up well.

Let’s look at Arroyo compared to pitchers who threw more than 200 change-ups between 2011 and 2013:

Movement:

Since change-ups (especially the circle change) tend to move down and to the right for right-handed pitchers versus down and to the left for southpaws, absolute value of x-Mov and z-Mov is used to standardize axis movement for both.

2011-2013 Abs(x-Mov) Abs(z-Mov)
League Average 7.17 4.30
Arroyo 6.00 3.60

I’ll give him a C- for movement. F’s are left for the likes of a Samuel Deduno, who posted a whopping 0.3″ of lateral and 1.6″ vertical (ignoring the natural pull of gravity) movement in 2013.

Velocity:

Again, keep in mind this does not include all pitchers, just ones who have thrown 200 or more change-ups between 2011 and 2013.

2011-2013 vFA (pfx) vCH (pfx)
League Average 90.9 82.9
Arroyo 86.6 78.2

When batters are already sitting on a below average fastball, it’s fair to say it won’t take much of an adjustment to catch up to the change. Below average may even be an understatement. There are only 12 guys in this data set of 275 with a lower average vFA. Jamie Moyer is one of them.

D+.

Location:

There are very few pitchers that can have success locating the change-up for called strikes.  Fernando Rodney being the freak off-speed guru who fools batters looking with a career 46.2 Swing%, 48.8 Zone% and 1.51 Val/C on the change. Typically the best change hurlers induce swings. And those swings either result in bad contact or a flat out whiff. But location of the pitch is still overwhelmingly crucial to achieve either.

I’ll use 2013 poor contact master Hyun-Jin Ryu and Braves injured whiff king Kris Medlen for illustration.

Ryu, with his 56.2 Swing% and 70.9 Contact% is looking to get bat on ball with the change. Ending 2013 with a .187 BABIP, the pitch worked beautifully to induce dribbling grounders (54.7 GB%) to an already above average Dodgers defense (3.1 UZR/150). How did he do it? Pin-perfect location (courtesy of Brooks Baseball).

 photo 74025e6d-0ca0-4068-802d-d2575977591e_zps07ccd3d1.png

Arroyo also induces hitters to get the bat on the ball with the change… at a whopping 85.5 Contact% rate. But is he getting poor contact with the pitch? I somehow don’t think .600+ SLG and 23 HR  over the past three full seasons would constitute bad contact. Let’s compare his zone chart with that of Ryu.

 photo 53238386-6da5-4f8b-9c21-44707dbd34a3_zpsc37ace95.png

 

Not quite, Bronson.

“But what about whiffs?” you ask. With a 6.8 career SwStr%, batters aren’t swinging and missing Arroyo’s meatballs either.

Let’s look at Medlen who owns a 27.5 career SwStr% on the pitch for comparison.
 photo 312d97a3-59b7-474d-8d46-43e4196b2988_zps9c5924cd.png

Pretty, no?

I’ll give Arroyo a D- for location. At least he’s not hanging them up and in on lefties.

So overall grade: barely passing.

I really don’t know what to say at this point. I’m miffed. Confounded. And who is the culprit to blame in the grand mystery of why he continues to throw this sub-par pitch? Batters have already gone deep on it twice in 2014. Is it the catchers? Do we point the finger at Devin Mesoraco, Ryan Hanigan, and now Miguel Montero for keeping blind faith and confidence? Are these guys cursed with chronic short-term memory loss? Or do we blame Arroyo for stubbornly going out there outing after outing and continuing to shove that ball in the back of his palm and firing away? If that’s the case, I get it. I’m a pitcher. I’ve stood there on the mound and though, “This next one will be better, guys. I swear!”

So, please, Bronson. In the end, there is really nothing good that has come from you throwing the thing so often. I like you. I really do. I will forever be indebted to you for giving my beloved 2004 Red Sox their first World Series since “tarnation” was a common curse word. But please. Enough change-ups already.


Battle of the Ks: K/9, K/BB and K%

The great debate has been raging for years: which strikeout-related metric is a better predictor of actual pitching success? Some would say there is no right or wrong answer — that each metric has it’s own unique merit and value. That one must look at certain strikeout-related metrics in combination with others. Unfortunately, as tragic as it may seem, statistical evidence begs to differ. Statistics tell us there is in fact a right answer, and it’s a whopper.

Let’s start with K/9. Looking at all 2013 pitchers with 80+ innings, the correlation (R2) between strikeouts per 9 and ERA is a solid  .1081. This correlation has been consistent, plus or minus a few hundredths, for the past five years. So nothing exciting or anomalous can be found in looking at other seasons. Yu Darvish leads the category with Tony Cingrani, Max Scherzer, Anibal Sanchez, and A.J. Burnett rounding out the top five. Additionally, eight of the top ten K/9 leaders ended up with sub 3.10 ERAs. So a decent indicator all-around.

 photo 53a65e17-24d6-482d-b2de-766753f09051_zps2940fbe7.png

K/BB get’s a bit more interesting. We see a jump in linear correlation to .1671 — more than a 50% increase over K/9. Clayton Kershaw, Cliff Lee, and Adam Wainwright  all leap into the top ten of this metric, with Hisashi Iwakuma climbing into the top fifteen — four elite hurlers in 2013 left out of the K/9 leaderboard.

 photo 98225caf-a307-44c3-850b-d610a9444d32_zps70ee67d9.png

But the real gem is K%. It shows double the correlation versus K/9. Plus, the top fifteen in this category ended the year with sub 3.30 ERA — whereas Scott Kazmir (4.04) and Josh Johnson (6.20) smeared the good name of the K/9 leaderboard; with Kevin Slowey (4.11) and Dan Haren (4.67) unpleasantly loitering on the K/BB board.

The reason K% is so powerful is that it simplifies how effective a pitcher is at simply striking out each batter he faces. When BABIP gets involved — as it does for K/9 (high BABIP pitchers are rewarded on K/9 since the number of outs remains the same even if they’re giving up, say, 10+ hits per game) — the value of each strikeout is severely reduced.

 photo 17feabf1-8665-45c5-af39-48d69923e54a_zpsf45972cf.png

 

To recap:

2013 R2 (correlation to ERA)
K/9 .1081
K/BB .1671
K% .2089

So should we end the debate completely? No. But if you asked me to put money on Tim Lincecum, a career 25.8 K% pitcher with no decline in the stat over the past 2 years, over Tyler Chatwood, a career 13.0 K% who had a breakout year in 2013 with his freakish 76.3% LOB, I would bet on Lincecum every doggone time.


The R.A. Dickey Effect – 2013 Edition

It is widely talked about by announcers and baseball fans alike, that knuckleball pitchers can throw hitters off their game and leave them in funks for days. Some managers even sit certain players to avoid this effect. I decided to analyze to determine if there really is an effect and what its value is. R.A. Dickey is the main knuckleballer in the game today, and he is a special breed with the extra velocity he has.

Most people that try to analyze this Dickey effect tend to group all the pitchers that follow in to one grouping with one ERA and compare to the total ERA of the bullpen or rotation. This is a simplistic and non-descriptive way of analyzing the effect and does not look at the how often the pitchers are pitching not after Dickey.

Dickey's Dancing Knuckleball
Dickey’s Dancing Knuckleball (@DShep25)

I decided to determine if there truly is an effect on pitchers’ statistics (ERA, WHIP, K%, BB%, HR%, and FIP) who follow Dickey in relief and the starters of the next game against the same team. I went through every game that Dickey has pitched and recorded the stats (IP, TBF, H, ER, BB, K) of each reliever individually and the stats of the next starting pitcher, if the next game was against the same team. I did this for each season. I then took the pitchers’ stats for the whole year and subtracted their stats from their following Dickey stats to have their stats when they did not follow Dickey. I summed the stats for following Dickey and weighted each pitcher based on the batters he faced over the total batters faced after Dickey. I then calculated the rate stats from the total. This weight was then applied to the not after Dickey stats. So for example if Janssen faced 19.11% of batters after Dickey, it was adjusted so that he also faced 19.11% of the batters not after Dickey. This gives an effective way of comparing the statistics and an accurate relationship can be determined. The not after Dickey stats were then summed and the rate stats were calculated as well. The two rate stats after Dickey and not after Dickey were compared using this formula (afterDickeySTAT-notafterDickeySTAT)/notafterDickeySTAT. This tells me how much better or worse relievers or starters did when following Dickey in the form of a percentage.

I then added the stats after Dickey for starters and relievers from all four years and the stats not after Dickey and I applied the same technique of weighting the sample so that if Niese’12 faced 10.9% of all starter batters faced following a Dickey start against the same team, it was adjusted so that he faced 10.9% of the batters faced by starters not after Dickey (only the starters that pitched after Dickey that season). The same technique was used from the year to year technique and a total % for each stat was calculated.

The most important stat to look at is FIP. This gives a more accurate value of the effect. Also make note of the BABIP and ERA, and you can decide for yourself if the BABIP is just luck, or actually better/worse contact. Normally I would regress the results based on BABIP and HR/FB, but FIP does not include BABIP and I do not have the fly ball numbers.

The size of the sample was also included, aD means after Dickey and naD is not after Dickey. Here are the results for starters following Dickey against the same team.

Dickey Starters

It can be concluded that starters after Dickey see an improvement across the board. Like I said, it is probably better to use FIP rather than ERA. Starters see an approximate 18.9% decrease in their FIP when they follow Dickey over the past 4 years. So assuming 130 IP are pitched after Dickey by a league average set of pitchers (~4.00 FIP), this would decrease their FIP to around 3.25. 130 IP was selected assuming ⅔ of starter innings (200) against the same team. Over 130 IP this would be a 10.8 run difference or around 1.1 WAR! This is amazingly significant and appears to be coming mainly from a reduction in HR%. If we regress the HR% down to -10% (seems more than fair), this would reduce the FIP reduction down to around 7%. A 7% reduction would reduce a 4.00 FIP down to 3.72, and save 4.0 runs or 0.4 WAR.

Here are the numbers for relievers following Dickey in the same game.

Dickey Bullpen

Relievers see a more consistent improvement in the FIP components (K, BB, HR) between each other (11.4, 8.1, 4.9). FIP was reduced 10.3%. Assuming 65 IP (in between 2012 and 2013) innings after Dickey of an average bullpen (or slightly above average, since Dickey will likely have setup men and closers after him) with a 3.75 FIP, FIP would get reduced to 3.36 and save 3 runs or 0.3 WAR.

Combining the un-regressed results, by having pitchers pitch after him, Dickey would contribute around 1.4 WAR over a full season. If you assume the effect is just 10% reduction in FIP for both groups, this number comes down to around 0.9 WAR, which is not crazy to think at all based off the results. I can say with great confidence, that if Dickey pitches over 200 innings again next year, he will contribute above 1.0 WAR just from baffling hitters for the next guys. If we take the un-regressed 1.4 WAR and add it to his 2013 WAR (2.0) we get 3.4 WAR, if we add in his defence (7 DRS), we get 4.1 WAR. Even though we all were disappointed with Dickey’s season, with the effect he provides and his defence, he is still all-star calibre.

Just for fun, lets apply this to his 2012. He had 4.5 WAR in 2012, add on the 1.4 and his 6 DRS we get 6.5 WAR, wow! Using his RA9 WAR (6.2) instead (commonly used for knucklers instead of fWAR) we get 7.6 WAR! That’s Miguel Cabrera value! We can’t include his DRS when using RA9 WAR though, as it should already be incorporated.

This effect may even be applied further, relievers may (and likely do) get a boost the following day as well as starters. Assuming it is the same boost, that’s around another 2.5 runs or 0.25 WAR. Maybe the second day after Dickey also sees a boost? (A lot smaller sample size since Dickey would have to pitch first game of series). We could assume the effect is cut in half the next day, and that’d still be another 2 runs (90 IP of starters and relievers). So under these assumptions, Dickey could effectively have a 1.8 WAR after effect over a full season! This WAR is not easy to place, however, and cannot just be added onto the teams WAR, it is hidden among all the other pitchers’ WARs (just like catcher framing).

You may be disappointed with Dickey’s 2013, but he is still well worth his money. He is projected for 2.8 WAR next year by Steamer, and adding on the 1.4 WAR Dickey Effect and his defence, he could be projected to really have a true underlying value of almost 5 WAR. That is well worth the $12.5M he will earn in 2014.

For more of my articles, head over to Breaking Blue where we give a sabermetric view on the Blue Jays, and MLB. Follow on twitter @BreakingBlueMLB and follow me directly @CCBreakingBlue.


TIPS, A New ERA Estimator

FIP, xFIP, SIERA are all very good ERA estimators, and their predictability is well documented. It is well known that SIERA is the best ERA estimator over samples that occur from season to season, followed very close by xFIP, with FIP lagging behind. FIP is best at showing actual performance though, because is uses all real events (K, BB, HR). Skill is commonly best attributed to either xFIP or SIERA. ERA is also well known to be the worst metric at predicting future performance, unless the sample size is very large <500IP with the pitcher remaining in the same or a very similar pitching environment.

FIP, xFIP, and SIERA are supposed to be Defense Independent Metrics, and they are. Well, they are independent of field defense, but there is one small error in the claim of defense independent. K’s and BB’s are not completely independent of defense. Catcher pitch framing plays a role in K’s and BB’s. Catchers can be good or bad at changing balls into strikes and this affects K’s and BB’s. Umpire randomness and umpire bias also play a role in K’s and BB’s. It is unknown how much of getting umpires to call more strikes is a skill for a pitcher or not. Some pitchers are consistent at getting more strike calls (Buehrle, Janssen) or less strike calls (Dickey, Delabar), but for most pitchers it is very random (especially in small sample sizes). For example Jason Grilli was in the top 5% in 2013 but was in bottom 10% in 2012.

I wanted to come up with another ERA estimator that eliminates catcher framing, umpire randomness and bias, and eliminates defense. I took the sample of pitchers who have pitched at least 200IP since 2008 (N=410) and analyze how different statistics that meet this criteria affect ERA-. I used ERA- since it takes out park factors and adjusts for the changes in the league from year to year. I looked at the plate discipline pitchf/x numbers (O-Swing, Z-Swing, O-Contact, Z-Contact, Swing, Contact, Zone, SwStr), the six different results based off plate discipline (zone or o-zone, swing or looking, contact or miss for ZSC%, ZSM%, ZL%, OSC%, OSM%, OL%), and batted ball profiles (GB%, LD%, FB%, IFFB%). *Please note that all plate discipline data is PitchF/X data, not the the other plate discipline on FanGraphs, this is important as the values differ*

The stats with very little to absolutely no correlation (R^2<0.01) were: Z-Swing%, Zone%, OSC%, ZSC%, ZL% (was a bit surprised as this would/should be looking strike%), GB%, and FB%. These guys are obviously a no-no to include in my estimator.

The stats with little correlation (R^2<0.1) were: Swing%, LD%, and IFFB%. I shouldn’t use these either.

O-Contact% (0.17), Z-Contact%, (.302), Contact% (.319), OSM% (0.206), and ZSM% (.248) are all obviously directly related to SwStr%. SwStr% had the highest correlation (.345) out of any of these stats. There is obviously no need to include all of the sub stats when I can just use SwStr%. SwStr% will be used in my metric.

OL% (0.105) is an obvious component of O-Swing% (0.192). O-Swing had the second highest correlation of the metrics (other than the components of SwStr%). I will use it as well. The theory behind using O-Swing% is that when the batter doesn’t swing it should almost always be a ball (which is bad), but when the batter swings, there are a two outcomes, a swing and miss (which is a for sure strike) or contact. Intuitively, you could say that contact on pitches outside the zone is not as harmful to pitchers as pitches inside the zone, as the batter should get worse contact. This is partially supported in the lower R^2 for O-Contact% to Z-Contact%. It is more harmful for a pitcher to have a batter make contact on a pitch in the zone, than a pitch out of the zone. This is why O-Swing is important and I will use it.

Using just SwStr% and O-Swing%, I came up with a formula to estimate (with the help of Excel) ERA-. I ran this formula through different samples and different tests, but it just didn’t come up with the results I was looking for. The standard deviation was way too small compared to the other estimators, and the root mean square error was just not good enough for predicting future ERA-.

I did not expect/want this estimator to be more predictive than xFIP or SIERA. This is because xFIP and SIERA have more environmental impacts in them that remain fairly constant. K% is always a better predictor of future K% than any xK% that you can come up with. Same with BB% Why? Probably because the environment of catcher framing, and umpire bias remain somewhat constant. Also (just speculation) pitchers who have good control can throw a pitch well out of the zone when they are ahead in the count, just to try and get the batter to swing or to “set-up” a pitch. They would get minus points for this from O-Swing, depending on how far the pitch is off the plate, but it may not affect their K% or BB% if they come back and still strike out the batter.

So I didn’t expect my statistic to be more predictive, but the standard deviation coupled with not that great of RMSE (was still better than ERA and FIP with a min of 40IP), caused me to be unhappy with my stat.

I then started to think about if there were any stats that were only dependent on the reaction between batter an pitcher that are skill based that FanGraphs does not have readily available? I started thinking about foul balls and wondered if foul ball rates were skill based and if they were related to ERA-. I then calculated the number of foul balls that each pitcher had induced. To find this I subtracted BIP (balls in play or FB+GB+LD+BU+IFFB) from contacts (Contact%*Swing%*Pitches). This gave me the number of fouls. I then calculated the rates of fouls/pitch and foul/contacts and compared these to ERA-. Foul/Contact or what I’m calling Foul%, had an R^2 of .239. That’s 2nd to only SwStr%. This got me excited, but I needed to know if Foul% is skill based and see what else it correlates with.

This article from 2008 gave me some insight into Foul%. Foul% correlates well to K% (obviously) and to BB% (negative relationship), since a foul is a strike. Foul% had some correlation to SwStr%, this is good as it means pitchers who are good at getting whiffs are also usually good at getting fouls. Foul% also had some correlation to FB% and GB%. The more fouls you give up, the more fly balls you give up (and less GB). This doesn’t matter however, as GB% and FB% had no correlation to ERA-. Foul% is also fairly repeatable year to year as evidenced in the article, so it is a skill. I will come up with a new estimator that includes Foul% as well.

I decided to use O-Looking% instead of O-Swing%, just to get a value that has a positive relationship to ERA (more O-looking means higher ERA), because SwStr% and O-Swing are negatively related. O-Looking is just the opposite of O-Swing and is calculated as (1 – O-Swing%).

The formula that Excel and I came up with is this: (I am calling the metric TIPS, for True Independent Pitching Skill)

TIPS = 6.5*O-Looking(PitchF/x)% – 9.5*SwStr% – 5.25*Foul% + C

C is a constant that changes from year to year to adjust to the ERA scale (to make an average TIPS = average ERA). For 2013 this constant was 2.68.

I converted this to TIPS- to better analyze the statistic. FIP, xFIP, and SIERA were also converted to FIP-, xFIP-, and SIERA-. I took all pitchers’ seasons from 2008-2013 to analyze. The sample varied in IP from 0.1 IP to 253 IP. I found the following season’s ERA- for each pitcher if they pitched more than 20 IP the next year and eliminated any huge outliers. Here were the results with no min IP. RMSE is root mean square error (smaller is better), AVG is the average difference (smaller is better), R^2 is self explanatory (larger is better), and SD is the standard deviation.

N=2316 ERA- FIP- xFIP- SIERA- TIPS-
RMSE 77.005 51.647 43.650 43.453 40.767
AVG 43.941 34.444 30.956 30.835 30.153
R^2 0.021 0.045 0.068 0.147 0.169
SD 69.581 38.654 24.689 24.669 15.751

Wow TIPS- beats everyone! But why? Most likely because I have included small samples and TIPS- is based off per pitch, as opposed to per batter (SIERA) or per inning (xFIP and FIP). There are far more pitches than AB or IP so TIPS will stabilize very fast. Let’s eliminate small sample sizes and look again.

Min 40 IP
N=1619 ERA- FIP- xFIP- SIERA- TIPS-
RMS 40.641 36.214 34.962 35.634 35.287
AVG 29.998 26.770 25.660 25.835 26.115
R^2 0.063 0.105 0.120 0.131 0.101
SD 26.980 19.811 15.075 17.316 13.843

 

Min 100 IP
N=654 ERA- FIP- xFIP- SIERA- TIPS-
RMSE 32.270 29.949 29.082 28.848 29.298
AVGE 24.294 22.283 21.482 21.351 22.038
R^2 0.080 0.118 0.143 0.145 0.095
SD 20.580 16.025 12.286 12.630 10.985

Now, TIPS is beaten out by xFIP and SIERA, but beats ERA and and is close to FIP (wins in RMSE, loses in R^2). This is what I expected, as I explained earlier K% and BB% are always better at predicting future K% and BB% and they are included in SIERA and xFIP. SIERA and xFIP take more concrete events (K, BB, GB) than TIPS. I didn’t want to beat these estimators, but instead wanted a estimator that is independent of everything except for pitcher-batter reaction.

TIPS won when there was no IP limit, so it obviously is the best to use in smaller sample sizes, but when is it better than xFIP and SIERA, and where does it start falling behind? I plotted the RMSE for my entire sample at each IP. Theoretically these should be an inverse relationship. After 150 IP it gets a bit iffy, as most of my sample is less than 100 IP. I’m more interested in IP under 100 anyhow.

Orange is TIPS, Blue is ERA, Red is FIP, Green is xFIP, and Purple is SIERA. If you can’t see xFIP, it’s because it is directly underneath SIERA (they are almost identical). This is roughly what the graph should look like to 100 IP:

Looking at the graph, at what IPs is TIPS better than predicting future ERA than xFIP and SIERA? It appears to be from 0 IP to around 70 IP.

Here is the graph for 1/RMSE (higher R^2). Higher number is better. This is the most accurate graph as the relationship should be inverse.

The 70-80 IP mark is clear here as well.

I’m not suggesting my estimator is better than xFIP or SIERA, it isn’t in samples over 75 IP, but I think it is, and can be, a very powerful tool. Most bullpen pitchers stay under 75 IP in a season. This means that my unnamed estimator would be very useful for bullpen arms in predicting future ERA. I also believe and feel that my estimator is a very good indicator of the raw skill of a pitcher. It would probably be even more predictive if we had robo-umps that eliminated umpire bias and randomness and pitch framing.

2013 TIPS Leaders with 100+IP

Name ERA FIP xFIP SIERA TIPS
Cole Hamels 3.6 3.26 3.44 3.48 3.02
Matt Harvey 2.27 2 2.63 2.71 3.09
Anibal Sanchez 2.57 2.39 2.91 3.1 3.23
Yu Darvish 2.83 3.28 2.84 2.83 3.23
Homer Bailey 3.49 3.31 3.34 3.39 3.26
Clayton Kershaw 1.83 2.39 2.88 3.06 3.32
Francisco Liriano 3.02 2.92 3.12 3.5 3.34
Max Scherzer 2.9 2.74 3.16 2.98 3.36
Felix Hernandez 3.04 2.61 2.66 2.84 3.37
Jose Fernandez 2.19 2.73 3.08 3.22 3.42

 

And Leaders from 40IP to 100IP

Name ERA FIP xFIP SIERA TIPS
Koji Uehara 1.09 1.61 2.08 1.36 1.87
Aroldis Chapman 2.54 2.47 2.07 1.73 2.03
Greg Holland 1.21 1.36 1.68 1.5 2.29
Jason Grilli 2.7 1.97 2.21 1.79 2.36
Trevor Rosenthal 2.63 1.91 2.34 1.93 2.42
Ernesto Frieri 3.8 3.72 3.49 2.7 2.45
Paco Rodriguez 2.32 3.08 2.92 2.65 2.50
Kenley Jansen 1.88 1.99 2.06 1.62 2.50
Glen Perkins 2.3 2.49 2.61 2.19 2.54
Edward Mujica 2.78 3.71 3.53 3.25 2.54