Archive for xstats

Can One Month of Statcast Data Be Used To Evaluate Hitters?

You have probably done this if you know about Statcast, “xStats,” and Baseball Savant. You pull up the xStats list, sort by under- or over-performers, and use it to draw broad and sweeping conclusions about your fantasy teams. Which of your fantasy players are poised for quick resurgence, or which of your opponents’ players are prime trade targets? Which guys should you be selling high on before the bottom drops out?

But in the same way that you can’t really sort the FanGraphs leaderboards by ERA minus FIP and just magically find pitching diamonds in the rough (homer rates complicate things…), this is maybe not the best way to be applying our vast wealth of fancy Statcast-based metrics. I’ve personally found that early-season Statcast data is difficult to trust, so I decided to dive in and see what exactly we can learn from one month of xStats.

It turns out there may be something useful here — the method I arrived at after this work would have advised you to buy-in on José Ramírez after his rough start to 2019! But we’ll get to that.

I will dive into gritty details below, but first to quickly outline, here are the major questions I’m setting out to answer (and what I ended up finding):

Read the rest of this entry »


The White Sox Might Have Found A No. 2 Starter For Nothing

The White Sox’ rotation this year can charitably be described as “rocky”. They began the year projected to have the worst rotation in the majors by WAR and thus far they’ve ranked 28th, between the Jeter-decimated Marlins and the aging Rangers. That’s not terribly surprising considering they’ve given out the most walks by far at 4.61 BB/9; besides them, only the Cubs’ rotation is over 4 at 4.21. The White Sox’ rotation also has the lowest strikeout rate in the majors this year at 6.20 K/9. The only thing preventing them from having the worst FIP of any team’s starters is middle-of-pack home run prevention, but their home field is a launching pad come summer.

As I stated before, they weren’t expected to have a good roster of starters, but being a rebuilding club filled with young and therefore volatile players, there was at least theoretically the chance that they made the jump to competence and beyond earlier than expected and surprise people like the Braves have this year. That obviously has not happened, but back in February, when everything is possible, Rian Watt took a look at the surprisingly large error bars in the projections for Chicago’s starters. The backstories of their projected starters agreed with what those large error bars said about a wide range of outcomes.

Lucas Giolito, a former No. 1 global prospect traded to the Sox last year from the Nationals, looked very sharp in spring training, having apparently rediscovered the massive 12-6 curve and some of the fastball velocity that had made him such a vaunted prospect and pairing it with newly found command and an improving, fading changeup. Reynaldo Lopez, fellow right-hander and former top-100 prospect who came over from the Nationals, had disappointing strikeout numbers despite big stuff, between a fastball that averaged 95 MPH, above-average curve and average slider and change– perhaps an improvement in sequencing or location would tap into the strikeouts he clearly had the talent to produce. Carson Fulmer, former No. 7; overall draft pick, has a lively arsenal in which everything moves in unpredictable ways that hitters dislike, albeit unpredictable to him too; perhaps he could make a mechanical adjustment and find the control and therefore success he had in college. Carlos Rodon, former No. 3 overall pick, was out with minor shoulder surgery (bursitis) until June but can flash complete dominance with his overpowering fastball/slider combo from the left side. Everyone knows about the world-class talent of Michael Kopech, who is currently stuck vaporizing poor saps in Triple-A (12.13 K/9!) until he limits his walks to acceptable levels. Bringing up the rear were Miguel Gonzalez, Hector Santiago, and James Shields, three veterans for whom the reasonable hopes were “eat innings better than cannon fodder”.

This article is not about any of the eight pitchers above, or their struggles with control (Giolito, Fulmer), relative successes (Shields), or weirdness (Lopez, who is having some success despite still not getting many strikeouts). Instead, it’s… Dylan Covey?

Yes, the Dylan Covey who ran both an ERA and FIP over seven last year in seventy innings as a rookie, good for -1.1 WAR. Pitching like, well, cannon fodder is not exactly an auspicious start to one’s major league career. Brief background of Covey: He was considered an elite high school arm, the riskiest category of draft picks, thought of high enough to be selected fourteenth overall in 2010 by Milwaukee– one pick after the White Sox selected a certain stick-figure lefty at a little-known Florida college whom Covey out-dueled earlier this June. During his pre-signing medicals, though, Covey was diagnosed with Type 1 diabetes, and he decided not to sign in order to learn how to deal with the disease before the stresses of pro ball. He chose to attend San Diego State and three years later was selected in the fourth round by Oakland.

After another three years of middling results hampered by injuries, Oakland left him off the 40-man roster despite an encouraging AFL and Chicago pounced in the Rule V draft. It was a bit of an unusual choice in that Covey was quite raw, almost akin to the Padres’ Rule V hijacking of prospects straight from A-ball, because Covey had thrown all of six starts at his highest level (Double-A). After hearing that, it probably makes a lot more sense why A) he got rocked the way he did last year and B) there was and is still hope for him. Although he was 25, the rawness showed, but the White Sox were entirely alright with absorbing the losses, as they would only help them pick higher in the 2018 Draft anyways (Nick Madrigal says hello).

Ironically, when he was drafted fourteenth overall in 2010, he was considered as safe as any high school arm could possibly be, on the basis of a low to mid-nineties sinker, above-average curve, ideal workhorse frame (currently listed at 6-2/195), and remarkably clean mechanics for his age. Ground balls, control, good health, and a reasonable number of strikeouts sounds like the perfect profile of a high-floor starter prospect. Of course, it didn’t work out that way in 2010, nor did he really come around while with Oakland. Thus, one might reasonably conclude, this article is being written because he appears to be finally delivering on his talent in his second year with the White Sox.

And so he has. Of course, the disclaimer of “small-sample size” applies here, as Covey has seven starts, and 35.1 innings total in those starts this year, but still, those 35.1 innings have been a complete reversal from his performance in 2017. He’s gotten a shot only because two rotation spots needed filling before Kopech was ready (i.e. past his Super Two deadline). First, Gonzalez went down with a shoulder injury in mid-April; that spot was filled by Santiago sliding from the bullpen into the rotation as he was signed to do. By mid-May, Fulmer’s wildness became too much to bear, and he was sent down to Triple-A to work on that, and Covey was called up to Chicago to get his second shot in the bigs. He’s taken that chance and run with it.

Thus far this year, Covey is the proud owner of a 2.29 ERA, 2.17 FIP, 3.31 xFIP, and 3.48 SIERA, good for a 1.3 fWAR (!) that currently leads all White Sox pitchers. No, I don’t think Covey is suddenly the third-best pitcher in baseball, and yes, that SIERA is a over a run higher than the FIP, and that’s because Covey has yet to give up a home run. That SIERA is still really good, though: among starters this year with at least 30 IP, the highest bar Covey clears, that would be good for 29th, slotting between Blake Snell and Alex Wood. Other pitcher evaluation metrics mostly agree: Baseball Savant’s xwOBA-against judges him at .293, 21st-best among starters. Baseball Prospectus’ DRA, how ever, does not like what he’s done, as his DRA this year is 5.38. There have been 4 unearned runs against him this year, so BBRef’s RA/9 dings him for that but still evaluates him well at 3.31 (Note: two of those unearned runs scored as inherited runners off a reliever). I cannot say why DRA hates him, but when a black-box statistic is in complete disagreement with literally every other ERA estimator, I have to ignore it.

Of course, the instinct of any saber-savvy fan is dismiss this as a fluke, small sample, etc. Anything can happen in small samples– once upon a time, Philip Humber threw a perfect game! That’s what I said, so when I trawled through Covey’s peripherals just to make sure this was a fluke, I kept expecting to find something or another that screamed regression. If there is a statistical red flag for harsh regression beyond his steadfast refusal to give up a home run, it remains as elusive to me as the average Bigfoot. His K% is a bit above average at 22.2% (starters’ average this year is 21.7%), his walk rate is a little better than average at 7.4% (avg is 8.2%), for a just above average K-BB% of 14.8% (avg of 13.6%). His LOB% is a bit low at 71.1% (avg 73.0%), and his BABIP-against is maybe a touch unlucky at .333 (avg .288). His WHIP is a smidge worse than average at 1.30 (avg 1.28). There is, in sum, absolutely nothing out of the ordinary there; by those measures he looks like a league average or slightly above starter. Which isn’t bad, as it suggests that his floor is that of a perfectly cromulent major-league starter, which is already a great outcome for a Rule V pick and vast improvement over last year.

Where Covey starts getting real interesting is when you start looking at the ways in which he might be suppressing home runs. I already told you that Covey’s primary pitch as a high schooler was a heavy sinker, and he’s gone back to his roots with it this year. In 2017, he threw fastballs about 60% of the time, splitting usage about evenly between his sinker and a four-seam. This year, he’s throwing even more fastballs, up to 68.3%, but he’s ditched the four-seam almost entirely; those are nearly exclusively sinkers he’s thrown. The point of a sinker is to get ground balls, and boy oh boy has his sinker done so.

Put simply, Covey’s been a ground ball machine. Among all starters with at least 30 IP this year, he’s tops in ground ball rate at 61.0%. The sinker has done most of that work; when batters put it in play, they beat it into the ground 68.1% of the time, 8th among starters. As one would expect, he’s also not allowed many fly balls; his FB% is a tiny 23.5%, seventh-lowest among his peers. Also unsurprisingly, he’s got the fourth-highest GB/FB, at 2.56, of starters. If his FIP is low because he’s not allowed a home run, well, it’s at least in part because it’s rather difficult to get a home run out of a grounder. When examined more closely, the metrics on his sinker back up its excellent results.

First of all, he’s added some velocity to it. This year his sinker is averaging 94.4 MPH, compared to last year’s 92.9 MPH. The addition of 1.5 to 2 MPH this year versus last is found in all his other pitches, too. Throwing harder across the board: always a good sign! It’s more than just respectably hard. Although Statcast classifies it as a 2-seamer, the pitch has the 29th-lowest average spin rate among either sinkers or 2-seamers this year.

While that and the velocity of the pitch (26th-fastest in the same mix of starters’ 2-seams & sinkers) are both good-not-great numbers, the combination of the two is actually pretty unusual– fastball velocity and spin rate usually have a positive correlation. Less spin is good in this case; the spin is mostly backspin, and the less backspin on a sinker, the more it sinks and (probably) the better it is. Of the 25 starters that throw their 2-seamers/sinkers harder than Covey does, only two– Erick Fedde and Fernando Romero, both rookies with small sample sizes themselves, also have lower spin rates. Stephen Strasburg and Sal Romano also throw harder and barely missed the spin rate cutoff. For comparison, the 2018 preview on Fedde’s FG page describes his sinker as “potentially premium”, Romano and Romero both have their fastballs graded by the FG prospect experts as 70s (plus-plus), and Strasburg rarely throws his 2-seamer.

In short, his sinker is elite for the sum of its parts. It’s generated an exactly league-average 6.8% whiff rate, which doesn’t sound special, but when it’s put in play, hitters can’t help but beat it into the ground. Its grounder/ball in play rate is an incredible 68.1%, 4th among starters and 10th among all pitchers this year. As would be expected, hitters haven’t done too well against it, with a xwOBA against of 0.324, checking in at 13th of all starters’ sinkers/2-seamers.

The three guys ahead of him on the starter list– Trevor Cahill, J.A. Happ, and Marcus Stroman— are interesting for comps, too. None strike out a ton of guys– all have career K/9s under eight– and none walk too many either, like Covey. Unsurprisingly, Stroman and Cahill, sinker/slider righties like Covey, are No. 2 and  No. 3 in starter GB% after Covey. Cahill’s having his best year yet in the A’s rotation, having upped his strikeouts to almost 9 K/9, cut his walks to 2 BB/9, and limiting home runs enough that ERA & ERA estimators are all around 3. Stroman, though he’s been hurt and not pitched well this year, has a track record of four years of being a solid No. 2 starter, especially according to SIERA.

Covey’s secondary pitches– slider (15.6% usage), curve (8.2%) and split-finger changeup (8.7%)– are all about average or better. The slider’s whiff rate is 13.5%, not spectacular but solidly above the league-average slider whiff of 9.0%. It’s not been murdered when it gets hit, either; Statcast’s xwOBA against the pitch is a pitiful .209, good for 16th among starters’ sliders. The change is an effective swing-and-miss pitch too, also with an above-average whiff rate at 15.6%. Hitters haven’t hit the change well either, with a xwOBA against of just .220, 16th among starters’ changeups. The curve hasn’t generated many swings-and-misses (just 2 out of 44 thrown, 4.5%) but hasn’t killed him at an xwOBA of .273, about middle of the pack for starters.

Baseball Savant sure doesn’t think that Covey’s just been extremely lucky in home-run suppression, but just to be sure, I went to go see what xStats.org thought of him. It thinks he should have given up 1.5 homers so far. Ignoring for a moment the fact that one cannot in fact hit half a home run, although a ground rule double seems close to it, that works out to a deserved rate of 0.382 HR/9. Which, in case you’re wondering, would still be good for fourth-lowest HR/9 of starters— Covey of course currently has the lowest of all at 0. Not perfect, then, but damn close to it. The other names in the top 10 lowest HR/9 are unsurprisingly for the most part really good to great pitchers: Arrieta, Nola, Severino, Bauer, Chatwood (???), deGrom, Buehler, Cueto, and Carlos Martinez, in ascending (towards lowest) order.

So that’s Dylan Covey in 2018: a pitcher with an excellent bread-and-butter sinker, two very good secondaries, and a passable fourth pitch. He’s not walking many, striking out close to a batter per inning, getting ground balls like they’re going out of fashion, and bucking the home run trend. I’m particularly reminded of Stroman in overall profile, but Covey has the advantages of size, a bit of youth, a home field with dirt instead of turf (grounders come off turf faster, meaning more hits), and a considerably younger and rangier infield behind him. He’s also got Don Cooper and Herm Schnieder on his coaching staff, which makes it less likely that he’ll be derailed by either mechanical or health issues. I for one didn’t see this coming, but the White Sox’ patience has already been rewarded with an unexpected breakout by Matt Davidson, so why couldn’t they have found another post-prospect gem? It’s at least interesting to note that Dallas Keuchel and Jake Arrieta, probably the best examples of guys who became great pitchers out of more or less nowhere after given time to reinvent themselves on rebuilding squads, are both in the top 20 in ground ball rate for starters– the category, of course, wherein Covey currently reigns supreme. I don’t really know what more to say. Small sample size notwithstanding, how about Dylan Covey, No. 2 starter?

Notes on process: with a small sample size of just seven starts at time of writing, the minimum cutoffs I employed to compare Covey to other pitchers were usually the minimum that he himself cleared– 30 IP with his 35.1 IP, 10 PA for his xwOBA against his curveball that has 13 PAs, etc. As he gets more starts, the exact numbers and rankings will of course change; the rankings are there not to be exact but rather to give some context for the raw numbers, most of which are obscure enough that the average reader likely cannot evaluate how “good” it is. Everyone knows a 2.29 ERA & 2.16 FIP are great, but I doubt many readers can instantly discern how good, say, a xwOBA of .220 against a certain pitcher’s changeup is. I also made the decision to evaluate almost exclusively against other starters’ 2018 years, as the baseball is again different this year and relievers are increasingly a different, turbo-powered breed of pitcher that cannot fairly be compared to starters.


Introducing xFantasy: Translating Hitters’ xStats to Fantasy

2016 has been a garbage year. At least, that’s what everyone seems to be talking about right now as the year draws to a close. But here in the baseball world, it’s been a banner year for many reasons, not the least of which is the new era of analysis that has arrived thanks to publicly available Statcast data. I, and I’m sure every other FG reader, have enjoyed following the quality Statcast analysis being developed in these electronic pages, particularly Andrew Perpetua’s “xStats”. In fact, I’m going to go ahead and stake the claim that I may have ‘coined’ (or at least influenced the creation of) the term xStats in the comments section of Andrew’s first xBABIP post. Inspired by the work of Perpetua, along with Alex Chamberlain (BIS-based xBABIP and xISO), and frequent leaguemate and Trevor-Story-lover Andrew Dominijanni (statcast xISO), I’ve decided to spend the offseason digging into xStats a bit deeper.

Perpetua has developed a great set of data using his binning strategy, most recently explained and updated this week, producing xBABIP, xBACON, and xOBA numbers based on Statcast’s exit velocity/launch angle data, along with the resulting ‘expected’ versions of the typical slash-line stats, xAVG/xOBP/xSLG. Throughout the year, I followed these stats fairly closely, often using ‘xStats’ to influence my fantasy baseball decisions. Given the opaque nature of translating a slash-line to actual fantasy stats, I generally went to the spreadsheet with the simple question “over- or under-performing?”, but that was about as far as I got. I found myself coming to probably-wrong conclusions such as “hey, maybe Sandy Leon isn’t actually that bad.” I was frustrated at my inability to turn a seemingly useful tool into actionable numbers for fantasy purposes.

This post serves as a starting point for that translation process. Way back in 2011, Jeff Zimmerman explained a basic approach for projecting R and RBI using only AVG, BB%, and HR% as inputs. I’ll similarly start here by coming up with simple models that translate rate stats (AVG, OBP, ISO) into fantasy-relevant ones, and then finally sub in the ‘x’ versions of those stats to come up with an ‘xFantasy’ line. I’ll stress that these are meant to be simple — I train the models based on all players that reached at least 300 PA in 2016, and I introduce a few team-related factors and shortcuts to improve fits, but I’m not looking to create a new Steamer or ZiPS here, just easy translations.

Home Runs

Starting with the surprisingly easy model, HR per PA is modeled well by ISO alone, with an R2 of .902 (excuse my simpleton’s application of statistics here; if you’re hoping for RMSE, p-values, etc., this will be a very disappointing post for you).

HR/PA = 0.2814*ISO – 0.01553

Runs and Runs Batted In

R and RBI per PA are interesting given their strong dependence on lineup position. To de-convolute that a bit, I’ve combined R+RBI into a single category (we can always separate them later). ‘R+RBI’ could be modeled using SLG alone, with an R2 of .758, but we can do better by separating SLG into AVG and ISO, and including terms for ‘team R+RBI total’ (player R/RBI totals are influenced by the team’s overall run production) and ‘average batting order position.’ Tanner Bell’s preseason post from this year explains and tabulates the influence of team offense and lineup position on R+RBI production. After doing some work to combine and normalize the data from Tanner’s tables, you can see the dependence of R+RBI/PA on lineup position can be roughly modeled as quadratic:

Average batting order position doesn’t appear to be easily accessible within the FanGraphs leaderboards, but thanks to the new ‘splits leaderboards’, it is possible to calculate with some elbow grease. Integrating all these factors to modify the original SLG model, R+RBI/PA is modeled by ‘SLGmod’ with an R2 of .807.

R+RBI/PA = 0.3292*SLGmod – 0.04751
SLGmod = AVG + 1.800*ISO + 2.061e-4*TeamR+RBI – 2.023e-3*ABO2 + 1.227e-2*ABO
                    TeamR+RBI = season total R+RBI for player’s team
                    ABO = average position of player in batting order

I mentioned that R+RBI could be separated later. Rather than demand the model predict the breakdown of R vs. RBI for each player, and introduce more sources of variation, I’m taking a shortcut here. The model calculates a value of x(R+RBI), and that is decomposed into R and RBI according to the actual proportion of R vs. RBI accumulated by the player in 2016. For instance, Mike Trout had 123 R and 100 RBI (223 R+RBI), and the model predicts 214.3 R+RBI, so we’ll give him (123/223)*214.3 = 118.2 R, and (100/223)*214.3 = 96.1 RBI.

Stolen Bases

SB per PA is a strange beast, a stat that’s much more dependent upon the whims and opportunities of the player and team than it is on the physical speed of the player. It can be tough to model given the large number of players that never run, or very rarely run. Much like SLG and R+RBI, I found that the SPD metric alone predicts SB/PA well, with an R2 of .662 when using a third-order polynomial fit. Is SPD cheating a bit? Maybe. For the uninitiated, it uses SB%, SB attempt frequency, triples percentage, and runs-scored percentage as inputs. You can see how SB/PA would fall directly out of that calculation, especially given the fact that teams tend to only turn runners loose on the basepaths if they are above a certain SB%. In any case, I’ll continue by modifying SPD to improve the fit, though the contribution of xStats to SB/PA will be much smaller than for the other stats.

Two rate stats serve to improve the fit, and they make intuitive sense: OBP, as players need to be on base in order to steal bases, and ISO, as players that hit for too much power tend not to spend as much time standing on first base, trying to steal second. I’ll again include a team factor, ‘team SB/PA,’ to quantify teams’ (or managers’) willingness to send runners, as well as ‘average batting order position,’ as players near the middle of the order tend not to steal as often. In this case I may have failed my initial criteria of a simple model, but it’s nevertheless a nice fit. Integrating it all into ‘SPDmod’, we can model SB/PA with an R2 of .834.

SB/PA = 0.2200*SPDmod3 – 0.3524*SPDmod2 + 0.2132*SPDmod – .04170
SPDmod = SPD/10 + 0.8206*OBP – 0.4670*ISO + 9.180*TeamSB – 9.192e-4*ABO2
                    TeamSB = average steals per plate appearance for player’s team

Average

Does batting average need its own section? I’m just going to use xAVG.

xFantasy

Now that I’ve reinvented the wheel and created a sort-of-okay way to calculate a 5×5 line based on rate stats, it’s a simple matter of substituting in the Perpetua xStats versions of AVG, OBP, and ISO to arrive at an ‘xFantasy’ line. I’ve also done a quick calculation of 2016 $ values using my normal z-score method, along with x$ values to allow easy comparison (no positional adjustments to either of them, though). The full sheet with 429 players’ 2016 xFantasy stats is found here, and I’ll include below the top-10 and bottom-10 players* whose lines improved/declined most when using xStats:

As one might hope, the top of the list is populated by several of the players that were identified as xStats’ undervalued darlings in 2016, like Mauer and Morales. In Belt, we might be seeing a place where park factors could improve xStats, though the disparity between his 17 HR and 29 xHR is still hard to ignore. Meanwhile, at the bottom of the list, it seems likely that the xSB model fails to adequately predict the SB totals for MLB’s most prolific runners, with Villar, Hamilton, and Nunez all getting hammered in the xSB category. But, it’s also possible that this is a knock-on effect from speedy players getting an unfair shake in xOBP. With Blackmon, it’s certainly possible that this is the other end of the park-factor spectrum, with his 20 xHR flagging way behind the 29 HR he put up.

Finally, one might ask how we solve the ‘Gary Sanchez problem,’ and it’d be quite useful to see what xStats project for players that only played partial seasons, to get an idea of what they ‘should’ have done over a full complement of PAs. Much like the ‘Steamer600’ projections hosted here at FanGraphs, I’ve calculated xFantasy600 values, where each player’s xFantasy line is normalized to 600 plate appearances. Or in other words, in this case, we’re evaluating players on a per-PA basis. Below, we have the top 20 players by xFantasy600 (x$600) in 2016:

Some new names rise to the top here, with Trea Turner, Gary Sanchez, and Trevor Story checking in as the third- (!!!), eighth- (!!), and 16th- (!) best players by xStats in 2016. On the one hand, they all appear to have over-performed in 2016 (check their wOBA vs. xOBA scores), but even regressing back to xStats in 2017 would comfortably land them among the best players in fantasy. The rest of this list is generally a who’s who of the best players in baseball, outside of Rickie Weeks, who was apparently highly effective as a platoon player last year. It’s fun to see that Big Papi went out on top, as the king of xFantasy. Miggy comes in at a very close No. 2, and I’ve seen him kicking around as a second-rounder on some early 2017 rankings – he might be the biggest bargain in drafts this year if that holds up. Overall, I’m very satisfied with this list’s ability to peg the best fantasy players, outside of the potential issue of underrating SBs.

Next time

The next step in this process is to evaluate xStats and xFantasy as a predictive tool. Throughout 2016, I pondered the fact that xStats might tell you more about “what happened” rather than “what will happen.” However, it’s hard to resist the allure of using them to project forward in-season, as they should stabilize faster than their standard statistical counterparts. One thing I have theorized is that xStats might be most helpful in evaluating ‘new swing’ guys, ‘new pitch’ guys, or new call-ups, as we wouldn’t expect traditional projection systems to capture these sorts of things. Craig Edwards has actually released an exceedingly timely look at “Did Exit Velocity Predict Second-Half Slumps, Rebounds?” I’ve now started work on the next chapter of the xFantasy story, comparing first-half and second-half numbers for 2015/2016 (the ‘Statcast era’) using traditional stats, xStats, and Steamer projections (h/t to Andrew Perpetua for updating his sheet to include first/second-half xStats splits).

This first look at xFantasy was a fun exploration of rudimentary projections and xStats. Hopefully others find it interesting; hit me up in the comments and let me know anything you might have noticed, or if you have any suggestions.