Archive for Home Runs

The White Sox Might Have Found A No. 2 Starter For Nothing

The White Sox’ rotation this year can charitably be described as “rocky”. They began the year projected to have the worst rotation in the majors by WAR and thus far they’ve ranked 28th, between the Jeter-decimated Marlins and the aging Rangers. That’s not terribly surprising considering they’ve given out the most walks by far at 4.61 BB/9; besides them, only the Cubs’ rotation is over 4 at 4.21. The White Sox’ rotation also has the lowest strikeout rate in the majors this year at 6.20 K/9. The only thing preventing them from having the worst FIP of any team’s starters is middle-of-pack home run prevention, but their home field is a launching pad come summer.

As I stated before, they weren’t expected to have a good roster of starters, but being a rebuilding club filled with young and therefore volatile players, there was at least theoretically the chance that they made the jump to competence and beyond earlier than expected and surprise people like the Braves have this year. That obviously has not happened, but back in February, when everything is possible, Rian Watt took a look at the surprisingly large error bars in the projections for Chicago’s starters. The backstories of their projected starters agreed with what those large error bars said about a wide range of outcomes.

Lucas Giolito, a former No. 1 global prospect traded to the Sox last year from the Nationals, looked very sharp in spring training, having apparently rediscovered the massive 12-6 curve and some of the fastball velocity that had made him such a vaunted prospect and pairing it with newly found command and an improving, fading changeup. Reynaldo Lopez, fellow right-hander and former top-100 prospect who came over from the Nationals, had disappointing strikeout numbers despite big stuff, between a fastball that averaged 95 MPH, above-average curve and average slider and change– perhaps an improvement in sequencing or location would tap into the strikeouts he clearly had the talent to produce. Carson Fulmer, former No. 7; overall draft pick, has a lively arsenal in which everything moves in unpredictable ways that hitters dislike, albeit unpredictable to him too; perhaps he could make a mechanical adjustment and find the control and therefore success he had in college. Carlos Rodon, former No. 3 overall pick, was out with minor shoulder surgery (bursitis) until June but can flash complete dominance with his overpowering fastball/slider combo from the left side. Everyone knows about the world-class talent of Michael Kopech, who is currently stuck vaporizing poor saps in Triple-A (12.13 K/9!) until he limits his walks to acceptable levels. Bringing up the rear were Miguel Gonzalez, Hector Santiago, and James Shields, three veterans for whom the reasonable hopes were “eat innings better than cannon fodder”.

This article is not about any of the eight pitchers above, or their struggles with control (Giolito, Fulmer), relative successes (Shields), or weirdness (Lopez, who is having some success despite still not getting many strikeouts). Instead, it’s… Dylan Covey?

Yes, the Dylan Covey who ran both an ERA and FIP over seven last year in seventy innings as a rookie, good for -1.1 WAR. Pitching like, well, cannon fodder is not exactly an auspicious start to one’s major league career. Brief background of Covey: He was considered an elite high school arm, the riskiest category of draft picks, thought of high enough to be selected fourteenth overall in 2010 by Milwaukee– one pick after the White Sox selected a certain stick-figure lefty at a little-known Florida college whom Covey out-dueled earlier this June. During his pre-signing medicals, though, Covey was diagnosed with Type 1 diabetes, and he decided not to sign in order to learn how to deal with the disease before the stresses of pro ball. He chose to attend San Diego State and three years later was selected in the fourth round by Oakland.

After another three years of middling results hampered by injuries, Oakland left him off the 40-man roster despite an encouraging AFL and Chicago pounced in the Rule V draft. It was a bit of an unusual choice in that Covey was quite raw, almost akin to the Padres’ Rule V hijacking of prospects straight from A-ball, because Covey had thrown all of six starts at his highest level (Double-A). After hearing that, it probably makes a lot more sense why A) he got rocked the way he did last year and B) there was and is still hope for him. Although he was 25, the rawness showed, but the White Sox were entirely alright with absorbing the losses, as they would only help them pick higher in the 2018 Draft anyways (Nick Madrigal says hello).

Ironically, when he was drafted fourteenth overall in 2010, he was considered as safe as any high school arm could possibly be, on the basis of a low to mid-nineties sinker, above-average curve, ideal workhorse frame (currently listed at 6-2/195), and remarkably clean mechanics for his age. Ground balls, control, good health, and a reasonable number of strikeouts sounds like the perfect profile of a high-floor starter prospect. Of course, it didn’t work out that way in 2010, nor did he really come around while with Oakland. Thus, one might reasonably conclude, this article is being written because he appears to be finally delivering on his talent in his second year with the White Sox.

And so he has. Of course, the disclaimer of “small-sample size” applies here, as Covey has seven starts, and 35.1 innings total in those starts this year, but still, those 35.1 innings have been a complete reversal from his performance in 2017. He’s gotten a shot only because two rotation spots needed filling before Kopech was ready (i.e. past his Super Two deadline). First, Gonzalez went down with a shoulder injury in mid-April; that spot was filled by Santiago sliding from the bullpen into the rotation as he was signed to do. By mid-May, Fulmer’s wildness became too much to bear, and he was sent down to Triple-A to work on that, and Covey was called up to Chicago to get his second shot in the bigs. He’s taken that chance and run with it.

Thus far this year, Covey is the proud owner of a 2.29 ERA, 2.17 FIP, 3.31 xFIP, and 3.48 SIERA, good for a 1.3 fWAR (!) that currently leads all White Sox pitchers. No, I don’t think Covey is suddenly the third-best pitcher in baseball, and yes, that SIERA is a over a run higher than the FIP, and that’s because Covey has yet to give up a home run. That SIERA is still really good, though: among starters this year with at least 30 IP, the highest bar Covey clears, that would be good for 29th, slotting between Blake Snell and Alex Wood. Other pitcher evaluation metrics mostly agree: Baseball Savant’s xwOBA-against judges him at .293, 21st-best among starters. Baseball Prospectus’ DRA, how ever, does not like what he’s done, as his DRA this year is 5.38. There have been 4 unearned runs against him this year, so BBRef’s RA/9 dings him for that but still evaluates him well at 3.31 (Note: two of those unearned runs scored as inherited runners off a reliever). I cannot say why DRA hates him, but when a black-box statistic is in complete disagreement with literally every other ERA estimator, I have to ignore it.

Of course, the instinct of any saber-savvy fan is dismiss this as a fluke, small sample, etc. Anything can happen in small samples– once upon a time, Philip Humber threw a perfect game! That’s what I said, so when I trawled through Covey’s peripherals just to make sure this was a fluke, I kept expecting to find something or another that screamed regression. If there is a statistical red flag for harsh regression beyond his steadfast refusal to give up a home run, it remains as elusive to me as the average Bigfoot. His K% is a bit above average at 22.2% (starters’ average this year is 21.7%), his walk rate is a little better than average at 7.4% (avg is 8.2%), for a just above average K-BB% of 14.8% (avg of 13.6%). His LOB% is a bit low at 71.1% (avg 73.0%), and his BABIP-against is maybe a touch unlucky at .333 (avg .288). His WHIP is a smidge worse than average at 1.30 (avg 1.28). There is, in sum, absolutely nothing out of the ordinary there; by those measures he looks like a league average or slightly above starter. Which isn’t bad, as it suggests that his floor is that of a perfectly cromulent major-league starter, which is already a great outcome for a Rule V pick and vast improvement over last year.

Where Covey starts getting real interesting is when you start looking at the ways in which he might be suppressing home runs. I already told you that Covey’s primary pitch as a high schooler was a heavy sinker, and he’s gone back to his roots with it this year. In 2017, he threw fastballs about 60% of the time, splitting usage about evenly between his sinker and a four-seam. This year, he’s throwing even more fastballs, up to 68.3%, but he’s ditched the four-seam almost entirely; those are nearly exclusively sinkers he’s thrown. The point of a sinker is to get ground balls, and boy oh boy has his sinker done so.

Put simply, Covey’s been a ground ball machine. Among all starters with at least 30 IP this year, he’s tops in ground ball rate at 61.0%. The sinker has done most of that work; when batters put it in play, they beat it into the ground 68.1% of the time, 8th among starters. As one would expect, he’s also not allowed many fly balls; his FB% is a tiny 23.5%, seventh-lowest among his peers. Also unsurprisingly, he’s got the fourth-highest GB/FB, at 2.56, of starters. If his FIP is low because he’s not allowed a home run, well, it’s at least in part because it’s rather difficult to get a home run out of a grounder. When examined more closely, the metrics on his sinker back up its excellent results.

First of all, he’s added some velocity to it. This year his sinker is averaging 94.4 MPH, compared to last year’s 92.9 MPH. The addition of 1.5 to 2 MPH this year versus last is found in all his other pitches, too. Throwing harder across the board: always a good sign! It’s more than just respectably hard. Although Statcast classifies it as a 2-seamer, the pitch has the 29th-lowest average spin rate among either sinkers or 2-seamers this year.

While that and the velocity of the pitch (26th-fastest in the same mix of starters’ 2-seams & sinkers) are both good-not-great numbers, the combination of the two is actually pretty unusual– fastball velocity and spin rate usually have a positive correlation. Less spin is good in this case; the spin is mostly backspin, and the less backspin on a sinker, the more it sinks and (probably) the better it is. Of the 25 starters that throw their 2-seamers/sinkers harder than Covey does, only two– Erick Fedde and Fernando Romero, both rookies with small sample sizes themselves, also have lower spin rates. Stephen Strasburg and Sal Romano also throw harder and barely missed the spin rate cutoff. For comparison, the 2018 preview on Fedde’s FG page describes his sinker as “potentially premium”, Romano and Romero both have their fastballs graded by the FG prospect experts as 70s (plus-plus), and Strasburg rarely throws his 2-seamer.

In short, his sinker is elite for the sum of its parts. It’s generated an exactly league-average 6.8% whiff rate, which doesn’t sound special, but when it’s put in play, hitters can’t help but beat it into the ground. Its grounder/ball in play rate is an incredible 68.1%, 4th among starters and 10th among all pitchers this year. As would be expected, hitters haven’t done too well against it, with a xwOBA against of 0.324, checking in at 13th of all starters’ sinkers/2-seamers.

The three guys ahead of him on the starter list– Trevor Cahill, J.A. Happ, and Marcus Stroman— are interesting for comps, too. None strike out a ton of guys– all have career K/9s under eight– and none walk too many either, like Covey. Unsurprisingly, Stroman and Cahill, sinker/slider righties like Covey, are No. 2 and  No. 3 in starter GB% after Covey. Cahill’s having his best year yet in the A’s rotation, having upped his strikeouts to almost 9 K/9, cut his walks to 2 BB/9, and limiting home runs enough that ERA & ERA estimators are all around 3. Stroman, though he’s been hurt and not pitched well this year, has a track record of four years of being a solid No. 2 starter, especially according to SIERA.

Covey’s secondary pitches– slider (15.6% usage), curve (8.2%) and split-finger changeup (8.7%)– are all about average or better. The slider’s whiff rate is 13.5%, not spectacular but solidly above the league-average slider whiff of 9.0%. It’s not been murdered when it gets hit, either; Statcast’s xwOBA against the pitch is a pitiful .209, good for 16th among starters’ sliders. The change is an effective swing-and-miss pitch too, also with an above-average whiff rate at 15.6%. Hitters haven’t hit the change well either, with a xwOBA against of just .220, 16th among starters’ changeups. The curve hasn’t generated many swings-and-misses (just 2 out of 44 thrown, 4.5%) but hasn’t killed him at an xwOBA of .273, about middle of the pack for starters.

Baseball Savant sure doesn’t think that Covey’s just been extremely lucky in home-run suppression, but just to be sure, I went to go see what xStats.org thought of him. It thinks he should have given up 1.5 homers so far. Ignoring for a moment the fact that one cannot in fact hit half a home run, although a ground rule double seems close to it, that works out to a deserved rate of 0.382 HR/9. Which, in case you’re wondering, would still be good for fourth-lowest HR/9 of starters— Covey of course currently has the lowest of all at 0. Not perfect, then, but damn close to it. The other names in the top 10 lowest HR/9 are unsurprisingly for the most part really good to great pitchers: Arrieta, Nola, Severino, Bauer, Chatwood (???), deGrom, Buehler, Cueto, and Carlos Martinez, in ascending (towards lowest) order.

So that’s Dylan Covey in 2018: a pitcher with an excellent bread-and-butter sinker, two very good secondaries, and a passable fourth pitch. He’s not walking many, striking out close to a batter per inning, getting ground balls like they’re going out of fashion, and bucking the home run trend. I’m particularly reminded of Stroman in overall profile, but Covey has the advantages of size, a bit of youth, a home field with dirt instead of turf (grounders come off turf faster, meaning more hits), and a considerably younger and rangier infield behind him. He’s also got Don Cooper and Herm Schnieder on his coaching staff, which makes it less likely that he’ll be derailed by either mechanical or health issues. I for one didn’t see this coming, but the White Sox’ patience has already been rewarded with an unexpected breakout by Matt Davidson, so why couldn’t they have found another post-prospect gem? It’s at least interesting to note that Dallas Keuchel and Jake Arrieta, probably the best examples of guys who became great pitchers out of more or less nowhere after given time to reinvent themselves on rebuilding squads, are both in the top 20 in ground ball rate for starters– the category, of course, wherein Covey currently reigns supreme. I don’t really know what more to say. Small sample size notwithstanding, how about Dylan Covey, No. 2 starter?

Notes on process: with a small sample size of just seven starts at time of writing, the minimum cutoffs I employed to compare Covey to other pitchers were usually the minimum that he himself cleared– 30 IP with his 35.1 IP, 10 PA for his xwOBA against his curveball that has 13 PAs, etc. As he gets more starts, the exact numbers and rankings will of course change; the rankings are there not to be exact but rather to give some context for the raw numbers, most of which are obscure enough that the average reader likely cannot evaluate how “good” it is. Everyone knows a 2.29 ERA & 2.16 FIP are great, but I doubt many readers can instantly discern how good, say, a xwOBA of .220 against a certain pitcher’s changeup is. I also made the decision to evaluate almost exclusively against other starters’ 2018 years, as the baseball is again different this year and relievers are increasingly a different, turbo-powered breed of pitcher that cannot fairly be compared to starters.


Home Runs and Temperature: Can We Test a Simple Physical Relationship With Historical Data?

Unlike most home-run-related articles written this year, this one has nothing to do with the recent home run surge, juiced balls, or the fly-ball revolution. Instead, this one’s about the influence of temperature on home-run rates.

Now, if you’re thinking here comes another readily disproven theory about home runs and global warming (a la Tim McCarver in 2012), don’t worry – that’s not where I’m going with this. Alan Nathan nicely settled the issue by demonstrating that temperature can’t nearly account for the large changes in home-run rates throughout MLB history in his 2012 Baseball Prospectus piece.

In this article, I want to revisit Nathan’s conclusion because it presents a potentially testable hypothesis given a large enough data set. If you haven’t read his article or thought about the relationship between temperature and home runs, it comes down to simple physics. Warmer air is less dense. The drag force on a moving baseball is proportional to air density. Therefore (all else being equal), a well-hit ball headed for the stands will experience less drag in warmer air and thus have a greater chance of clearing the fence. Nathan took HitTracker and HITf/x data for all 2009 and 2010 home runs and, using a model, estimated how far they would have gone if the air temperature were 72.7°F rather than the actual game-time temperature. From the difference between estimated 72.7°F distances and actual distances, Nathan found a linear relationship between game-time temperature and distance. (No surprise, given that there’s a linear dependence of drag on air density and a linear dependence of air density on temperature.) Based on his model, he suggests that a warming of 1°F leads to a 0.6% increase in home runs.

This should in principle be a testable hypothesis based on historical data: that the sensitivity of home runs per game to game-time temperature is roughly 0.6% per °F. The issue, of course, is that the temperature dependence of home-run rates is a tiny signal drowned out by much bigger controls on home-run production [e.g. changes in batting approach, pitching approach, PED usage, juiced balls (maybe?), field dimensions, park elevation, etc.]. To try to actually find this hypothesized temperature sensitivity we’ll need to (1) look at a massive number of realizations (i.e. we need a really long record), and (2) control for as many of these variables as possible. With that in mind, here’s the best approach I could come up with.

I used data (from Retrosheet) to find game-time temperature and home runs per game for every game played from 1952 to 2016. I excluded games for which game-time temperature was unavailable (not a big issue after 1995 but there are some big gaps before) and games played in domed stadiums where the temperature was constant (e.g. every game played at the Astrodome was listed as 72°F). I was left with 72,594 games, which I hoped was a big enough sample size. I then performed two exercises with the data, one qualitatively and one quantitatively informative. Let’s start with the qualitative one.

In this exercise, I crudely controlled for park effects by converting the whole data set from raw game-time temperatures (T) and home runs per game (HR) to what I’ll call T* and HR*, differences from the long-term median T and HR values at each ball park over the whole record. Formally, for any game, T* and HR* are defined such that T* = T – Tmed,park and HR* = HR – HRmed,park, where Tmed,park and HRmed,park are median temperature and HR/game, respectively, at a given ballpark over the whole data set. A positive value of HR* for a given game means that more home runs were hit than in a typical ball game at that ballpark. A positive value for T* means that it was warmer than usual for that particular game than on average at that ballpark. Next, I defined “warm” games as those for which T*>0 and “cold” games as those for which T*<0. I then generated three probability distributions of HR* for: 1) all games, 2) warm games and 3) cold games. Here’s what those look like:

The tiny shifts of the warm-game distribution toward more home runs and cold-game distribution toward fewer home runs suggests that the influence of temperature on home runs is indeed detectable. It’s encouraging, but only useful in a qualitative sense. That is, we can’t test for Nathan’s 0.6% HR increase per °F based on this exercise. So, I tried a second, more quantitative approach.

The idea behind this second exercise was to look at the sensitivity of home runs per game to game-time temperature over a single season at a single ballpark, then repeat this for every season (since 1952) at every ballpark and average all the regression coefficients (sensitivities). My thinking was that by only looking at one season at a time, significant changes in the game were unlikely to unfold (i.e. it’s possible but doubtful that there could be a sudden mid-season shift in PED usage, hitting approach, etc.) but changes in temperature would be large (from cold April night games to warm July and August matinees). In other words, this seemed like the best way to isolate the signal of interest (temperature) from all other major variables affecting home run production.

Let’s call a single season of games at a single ballpark a “ballpark-season.” I included only ballpark-seasons for which there were at least 30 games with both temperature and home run data, leading to a total of 930 ballpark-seasons. Here’s what the regression coefficients for these ballpark-seasons look like, with units of % change in HR (per game) per °F:

A few things are worth noting right away. First, there’s quite a bit of scatter, but 75.1% of these 930 values are positive, suggesting that in the vast majority of ballpark-seasons, higher home-run rates were associated with warmer game-time temperatures as expected. Second, unlike a time series of HR/game over the past 65 years, there’s no trend in these regression coefficients over time. That’s reasonably good evidence that we’ve controlled for major changes in the game at least to some extent, since the (linear) temperature dependence of home-run production should not have changed over time even though temperature itself has gradually increased (in the U.S.) by 1-2 °F since the early ‘50s. (Third, and not particularly important here, I’m not sure why so few game-time temperatures were recorded in the mid ‘80s Retrosheet data.)

Now, with these 930 realizations, we can calculate the mean sensitivity of HR/game to temperature, resulting in 0.76% per °F. [Note that the scatter is large and the distribution doesn’t look very Gaussian (see below), but more Dirac-delta like (1 std dev ~ 1.66%, but middle 33% clustered within ~0.4% of mean)].

Nonetheless, the mean value is remarkably similar to Alan Nathan’s 0.6% per °F.

Although the data are pretty noisy, the fact that the mean is consistent with Nathan’s physical model-based result is somewhat satisfying. Now, just for fun, let’s crudely estimate how much of the league-wide trend in home runs can be explained by temperature. We’ll assume that the temperature change across all MLB ballparks uniformly follows the mean U.S. temperature change from 1952-2016 using NOAA data. In the top panel below, I’ve plotted total MLB-wide home runs per complete season (30 teams, 162 games) season by upscaling totals from 154-game seasons (before 1961 in the AL, 1962 in the NL), strike-shortened seasons, and years with fewer than 30 teams accordingly. In blue is the expected MLB-wide HR total if the only influence on home runs is temperature and assuming the true sensitivity to be 0.6% per °F. No surprise, the temperature effect pales in comparison to everything else. Shown in the bottom plot is the estimated difference due to temperature alone in MLB-wide season home run totals from the 1952 value of 3,079 (again, after scaling to account for differences in number of games and teams). You can think of this plot as telling you how many of the total home runs hit in a season wouldn’t have made it over the fence if air temperatures at remained constant at 1952 levels.

While these anomalies comprise a tiny fraction of the thousands of home runs hit per year, one could make that case (with considerably uncertainty admitted) that as many as 59 of these extra temperature-driven home runs were hit in 2016 (or about two per team!).


How to Make Yourself Interesting

Allow me to start this post off with a couple of charts without any context about the player we are talking about.

 

Let’s talk about this player for a second. His name is not of consequence, yet. This player has fluctuated from being an above-average producer of runs and slightly-below-average producer of runs for close to 10 years now. This means he’s been around a long time, so his profile as a hitter is solidified; he has a reputation. Something funny has happened in 2016 and 2017 as evidenced by the LARGE upward line. That’s good! Can you guess who this player is? No? Come on, one guess. Okay, fine. It’s Mark Reynolds! Yes, that Mark Reynolds!

Mark Reynolds once hit 44 home runs. Do you remember that? When I said, he had a reputation, I meant to say that he’s well-known for the three true outcomes: walks, strikeouts, and home runs. Not much else. He’s a first baseman, which means his defensive value is minimal at best. So basically, his value is his offense. He’s signed for $1.5 million this year and is currently a top-10 first baseman in the MLB by fWAR. He’s top-8 by wRC+, and top-4 by wOBA. He’s already exceeded the value of his contract. The obvious caveat here: it’s May 9th. The other obvious caveat is he plays for the Rockies now, which means he gets to play 81 games (give or take) at Coors Field.

I don’t know if he can sustain this. I too see the name Mark Reynolds and think, 30% K rate, with a decent amount of power. The thing is, he’s not striking out in 30% of his plate appearances. He’s not even striking out in 25% of his plate appearances. You want to know how often he’s striking out? After today’s day game with the Cubs, he’s striking out only 21.1% of the time. That’s, dare I say, below league average (League wide K% currently is 21.5%). I’m going to throw some more numbers together to try and articulate an idea: Mark Reynolds is up to something.

This doesn’t seem to be a one-year fluke. Reynolds is a slightly different player than he was two years ago. His K% has been on the decline since 2015, when it was 28%. Last year it was 25.4% and obviously now it’s 21.1%. So let’s go to his plate discipline to see what’s changed.

Looking at his O-Swing% and PITCHf/x O-Swing%, there isn’t a huge difference. They both hover in and around his rate of 26-27%, though PITCHf/x has him at 29.5%. The real difference is in his Z-Swing%, where he has decreased his percentage over the last two years. In 2015, it was around his career norm of 70% by Baseball Info Solutions and 67% by PITCHf/x. The last two years: 69.4% and 66.2% respectively by Baseball Info Solutions, 66.2% and 64.9% respectively by PITCHf/x. He seems to be pickier in the zone overall and there is a tangible result.

His Z-Contact% career average as calculated by PITCHf/x and Baseball Info Solutions is 74.3% and 74%, respectively. In 2015, he made contact with pitches in the zone 80% of the time by both systems. Last year? 81.9% by Baseball Info Solutions and 84.6% by PITCHf/x. This year? 85.4% and 84%. He’s making more contact overall for the last two years, as it’s been in the 70% range rather than the 60% range. His SwStr% has been decreasing too! It’s been below 13% the last two years, where his career average is 15.7%. This is a different Mark Reynolds.

Maybe Reynolds is trying to take more pitches in the zone so he can focus in on his best pitch. The power is there — his ISO is .339, with 12 home runs thus far. Probably not sustainable, but 30 home runs can be reached even with a return to the average.

About that park factor, though. He really hasn’t hit much differently at Coors versus away from Coors.

The same amount of hits, admittedly more home runs, same amount of strikeouts, same amount of walks. Slightly odd thing — he has a reverse platoon split. Let’s chalk that up to small sample size. One more chart that I feel is important:

This chart befuddles me. He’s hitting fewer fly balls than league average (opposite league trend as touched on at FG main page), more ground balls than league average, and slightly more line drives than league average. Something funny is happening here. So, here’s the thing. His HR/FB is 44%. Aaron Judge is at 46.4% and no one expects him to sustain it. League average is 12.8% and Reynolds’ career high is 26%. His career average is 19.4%, which he hasn’t reached since 2011.
It all comes back to small samples, but even if he comes crashing back down, there’s still proof he’s trying to make a change. He’s making more contact and we know contact is a good thing, and this has been happening for more than just 30 games. If he sustains a fraction of this pace, he becomes trade bait at the deadline, or he stays part of a contender, and he may even get a pay raise in free agency. Mark Reynolds has made himself interesting.

Prospect Watch: 5 Future All-Stars No One Is Talking About

I chose to stick with hitters in this article, because pitching prospects are extremely difficult to predict, and I think the pitchers who do get the hype are typically deserving. However, I do see a trend of some unnoticed hitting prospects turning out great careers in the majors. Let’s get right to it.

1. Travis Demeritte – 2B – ATL

In 2016, Demeritte went from the Rangers’ to the Braves’ system and spent the entire year in high-A ball, where he dominated at the plate. A 2B with power like Cano, good speed and the ability to get on base is such a rarity.

In my opinion, Demeritte has the highest chance of being a perennial All-Star out of these five prospects. The middle infield in Atlanta has an extremely bright future. I’m predicting that Demeritte will make his splash in 2018, and make his first ASG appearance by 2020 (age 25). Let’s look at his numbers from a season ago:

 

Name Age G AB PA H 2B 3B HR BB SO SB CS BB% K% OPS ISO wOBA wRC+
Travis Demeritte 21 145 547 635 145 33 13 32 78 200 20 4 12.3% 31.5% 0.905 0.283 0.393 139


Let’s compare these to the four All-Star 2B in 2016 and Brian Dozier.

Name G AB PA H 2B 3B HR BB SO SB CS BB% K% OPS ISO wOBA wRC+
Jose Altuve 161 640 717 216 42 5 24 60 70 30 10 8.4% 9.8% 0.928 0.194 0.391 150
Robinson Cano 161 655 715 195 33 2 39 47 100 0 1 6.6% 14.0% 0.882 0.235 0.37 138
Brian Dozier 155 615 691 165 35 5 42 61 138 18 2 8.8% 20.0% 0.886 0.278 0.37 132
Dustin Pedroia 154 633 698 201 36 1 15 61 73 7 4 8.7% 10.5% 0.825 0.131 0.358 120
Ian Kinsler 153 618 679 178 29 4 28 45 115 14 6 6.6% 16.9% 0.831 0.196 0.356 123


Some things to keep in mind as we compare these players: Demeritte was playing in A+ ball, but he did play an average of 12 less games than these major-leaguers. As you can see, it’s basically a two-man race (other than Dozier’s 42 HRs) between Altuve and Demeritte here. While we cannot expect these A+ ball numbers to translate directly against ML pitching, Demeritte definitely deserves more attention in top-prospect lists. While he’s not quite as speedy as Altuve, he has more power, and he walks at a far higher rate. The one glaring weakness is the K numbers for Demeritte. However, some of the top players in the league K at very high rates. As long as the OPS stays high, it doesn’t really matter how a guy makes outs anymore.

I should note that 2016 was a breakout year for Demeritte; in years past he didn’t quite live up to his potential, and also served an 80-game PED suspension. These could be the main reasons why he hasn’t garnered much attention yet. He still has to prove himself to most. However, I’m sold. I’d pencil him in for the majority of the 2020s’ ASGs right now.

 

2. Ramon Laureano – OF – HOU

Laureano has all the tools: he can play any OF spot well, he has speed and pop, and he gets on base. Houston’s farm has taken a bit of a hit due to some trades in the last two years, but that’s because they knew they had guys like Laureano who don’t have super high trade value, but have a chance to be great ML players like the guys they traded. Let’s look at Laureano’s 2016 numbers.

Name Age G AB PA H 2B 3B HR BB SO SB CS BB% K% OPS ISO wOBA wRC+
Ramon Laureano 21 128 461 555 146 32 9 15 73 128 48 15 13.2% 23.1% 0.943 0.206 0.418 159


The numbers speak for themselves. This is the making of a star; where is the hype? I know it’s not a huge sample size, and we don’t have much to go off from the previous year either, but in A+ and AA last year he put up those phenomenal numbers you see above.

If those aren’t All-Star numbers, then I don’t know what are. Laureano’s ability to play all three OF spots will keep him in the lineup everyday and help his chances of making it to the ASG. When he does get the call-up, if his numbers stay relatively close to this, there’s no way he doesn’t make three to four All-Star Games. As of now, he’s more of a speed threat, but as he develops, the speed/power combo will even out and he will be an Andrew McCutchen-type player. Keep tabs on this guy.

 

3. Christin Stewart – OF – DET

While researching Stewart, I couldn’t find an article more recent than September of 2015. There’s no one talking about him…why? As we know, Detroit is aging and looking to deal top players. So, I’m assuming we will be seeing a lot of opportunities for young guys to step up and prove themselves. Detroit’s system isn’t super deep, but that could change anytime if they do decide to move some key pieces. Regardless, I see Stewart as the prospect to watch moving forward; he has the tools to be an All-Star. Let’s check out his numbers from 2016.

Name Age G AB PA H 2B 3B HR BB SO SB CS BB% K% OPS ISO wOBA wRC+
Christin Stewart 22 147 514 622 132 29 2 31 93 154 4 2 15.0% 24.8% 0.883 0.245 0.407 156


The power is impressive, and by this chart he looks even a bit better than the two previous guys I mentioned. However, with the K numbers pretty high up there, and not a whole lot of speed, Stewart is a player that could fall into slumps. Often times, adjusting to the majors can be challenging, and some top prospects never quite figure it out. While Stewart’s MiLB numbers are pretty insane, his slump potential makes him a pretty risky pick here. However, I do believe that if he does indeed figure it out, he will make it to a few ASG and serve as an everyday player in this league for a decade. HRs and BBs get it done. Keep an eye on Stewart.

 

4. Jason Martin – OF – HOU

Another Houston OF prospect…another future All-Star? I think so. The future is certainly bright over at Minute Maid Park: Altuve is a cornerstone, Correa is a centerpiece, Springer is a baller, and they have prospects for days. If they can just figure out how to pitch, they could be a WS contender for the next eight years.

Why Martin, though? Let’s check out his 2016 numbers from high-A ball.

Name Age G AB PA H 2B 3B HR BB SO SB CS BB% K% OPS ISO wOBA wRC+
Jason Martin 20 121 431 502 114 25 7 23 63 112 22 12 12.5% 22.3% 0.874 0.251 0.382 131


Impressive, to say the least. At just 20 years old, he pumped out 23 homers in 121 games. He walks every eight at-bats, and he also grabbed 22 bags on the season. The ability to walk and run (lol) will typically keep guys out of major slumps. While Martin is not a highly-touted prospect at this point, I think he will be a household name by 2022. I expect him to get the call-up in 2019 and play a significant role during a pennant race that year. In 2020, he will burst onto the scene and prove his worth to this franchise.

With Houston’s current build, this might be a guy we see dealt if they are trying to add talent at the deadline this year. That doesn’t change my prediction, however. I see Martin suiting up for the ASG a few times throughout his career. Stay posted.

 

5. Tom Murphy – C – COL

You can’t keep putting Yadier Molina in there every year. And with Buster Posey most likely making that change to 1B full-time within three years, Jonathan Lucroy getting dealt to the AL, Kyle Schwarber playing OF, etc, pathways for guys like Tommy Murphy open up. Making the All-Star Game as a C is not saying as much as other positions, in my opinion. A decent hot streak in the first half will inflate your hitting numbers. For example, Derek Norris in 2014. It may seem like he was the best catcher in the league at the halfway point, but, as usual, it evened out by season’s end.

With that being said, Murphy has proven he has pop, and playing in Colorado is a huge advantage for him. While I don’t think he will be a Hall-of-Fame catcher, I do think he’s flying under the radar right now and will probably open some eyes in 2017. I’d say he makes two appearances in the ASG before 2022. However, once he gets up near 30 and he’s no longer playing in Colorado, I think he will have trouble keeping a job.

I have him on the list, first of all, because he meets the criteria, and also because I think people should pay attention to him, and lastly because he’s ML-ready, unlike the rest of these guys. Trevor Story didn’t have a whole lot of hype; most people didn’t expect him to make the team out of spring, but with the Jose Reyes situation, the kid got a shot and as we all know, he ran with it. I’m not saying Murphy will make a cannonball-esque splash like Story, but I think he will turn some heads and maybe even get some ASG votes this year. Anything can happen, especially in Colorado. Keep tabs on him.

Honorable Mentions

Dylan Cozens – OF – PHI

There’s not a lot of buzz surrounding Cozens, which is surprising to me, because usually when we see 40 HR in 134 games, we really perk up. In his age-22 season, he played all 134 games at the AA level for the Phillies affiliate, Reading Fightin’ Phils, a place where most Phillies prospects prosper. The reason why Cozens doesn’t quite make the cut here is because of the words, “future All-Star.” He is one of those lefties that mash in the right ballpark and against RHP, but usually career platoon hitters, even if they are highly effective, don’t make the ASG.

Rhys Hoskins – 1B – PHI

Hoskins is another AA player in the Phillies system. He probably has a little bit more of a well-rounded hitting ability than does Cozens, but he’s a 1B, and that’s an overloaded position. You have to be incredible to crack that ASG squad, and I just don’t think Hoskins will ever be quite at that level. I do believe he will pan out to be an everyday guy for a good amount of time in this league. He has really good power and he gets on base, two things that will keep you in the lineup more often than not.

Bobby Bradley – 1B – CLE

Bradley is another guy I would keep an eye on; I’m just not sold on him yet. He has a a lot of raw power, but a really high K rate in the low levels of the minors. Also, he’s a 1B, so once again, really hard to make the ASG at that position.


The Homer Numbers of a Hypothetically-Healthy Giancarlo Stanton

Giancarlo Stanton has missed significant playing time since his MLB debut in 2010 and has never played more than 150 games of a 162-game season (145 and 123 games being his next two highest totals). In spite of his injury-shortened seasons, Stanton has still been among the league home-run leaders in 2011, 2012, and 2014 (his 150, 123, and 145-game seasons, respectively).

Giancarlo Stanton Since Debut (June 2010)
Season Games PA HR HR MLB Rank Injury Report
2010 100 396 22 T-55 ——
2011 150 601 34 9 Hamstring issues limited time
2012 123 501 37 7 15-day DL: Arthroscopic knee surgery
2013 116 504 24 T-31 15-day DL: Strained right hamstring
  2014* 145 638 37 2 Season-ending facial fracture
2015 74 318 27 T-25 15-day DL: Season-ending hamate (hand) fracture
2016 119 470 27 48 15-day DL: Strained left groin
*=finished 2nd in NL MVP race (Clayton Kershaw)

Career-wise, Stanton has amassed a total of 208 home runs, good enough for 16th-most of any player through their age-26 season and among the likes of Miguel Cabrera and Jose Canseco.

HR-leaders through Age-26 season
Rank Player HR
1 Alex Rodriguez 298
2 Jimmie Foxx 266
3 Eddie Matthews 253
4 Albert Pujols 250
5 Mickey Mantle 249
6 Mel Ott 242
7 Frank Robinson 241
8 Ken Griffey, Jr. 238
9 Orlando Cepeda 222
10 Andruw Jones 221
11 Hank Aaron 219
12 Juan Gonzalez 214
13 Johnny Bench 212
14 Miguel Cabrera 209
14 Jose Canseco 209
16 Giancarlo Stanton 208

Given Stanton’s injury-plagued career, his career home-run numbers are a lower bound on what he may have accomplished had he played full, injury-free seasons following his debut. To quantify how Stanton’s injuries have suppressed Stanton’s career power numbers thus far, I extrapolated the home-run totals of Stanton’s injury-shortened seasons into full-season hypothetical home-run totals (hHR) using the formula below:

hHR = FLOOR(HR/G * 162)

The formula simply assumes that Stanton maintains his HR/G rate through a whole 162-game season and then conservatively rounds down. We can now compare home-run totals between the real Giancarlo Stanton and our hypothetical Giancarlo Stanton. I excluded his 2010 debut from the extrapolation.

Real Giancarlo Stanton vs. Hypothetical Giancarlo Stanton
Season Games HR HR MLB Rank hGames hHR hHR MLB Rank
2010 100 22 T-55 100 22 T-55
2011 150 34 9 162 36 8
2012 123 37 7 162 48 1
2013 116 24 T-31 162 33 T-9
2014 145 37 2 162 41 1
2015 74 27 T-25 162 59 1
2016 119 27 48 162 36 T-16

The real Stanton never led the MLB in home runs, but our hypothetical Stanton climbs into the MLB lead in three of his hypothetical seasons (2012, 2014, and 2015).

Career-wise, our hypothetical Stanton would have hit 275 total home runs. This hypothetical Stanton adds 67 home runs to his real total, jumping from 16th to second place on the Age-26 leaderboard, only 23 home runs behind the far-away leader, Alex Rodriguez.

HR-leaders through Age-26 season
Rank Player HR
1 Alex Rodriguez 298
2 Giancarlo Stanton (hypothetical) 275
3 Jimmie Foxx 266
4 Eddie Matthews 253
5 Albert Pujols 250
6 Mickey Mantle 249
7 Mel Ott 242
8 Frank Robinson 241
9 Ken Griffey, Jr. 238
10 Orlando Cepeda 222
11 Andruw Jones 221
12 Hank Aaron 219
13 Juan Gonzalez 214
14 Johnny Bench 212
15 Miguel Cabrera 209
16 Jose Canseco 209
17 Giancarlo Stanton (real) 208

Of note, using the same formula to calculate Stanton’s career strikeout totals predicts a whopping 1271 strikeouts for our hypothetical Stanton. His 977 strikeout “real” total through age 26 (second-highest) balloons and surpasses Justin Upton’s age-26-leading 1026 for a clear command of first place.

In reality, Stanton is a three-time All-Star, a Silver Slugger (2014), and a Home Run Derby champion (2016), and he historically ranks among the best in home-run totals for his age, all while facing injury issues in all of his first six full big-league seasons. Our hypothetically-healthy Giancarlo Stanton greatly improves his career numbers and garners himself a few MLB home-run crowns, giving a glimpse into how much larger his career numbers could be today had his first six full seasons been injury-free. As Stanton’s career progresses, it will be interesting to see where his home-run totals end up, and, unfortunately, how much greater they could have been.

Credit to Baseball-Reference for all publicly available data.


xHR%: Questing for a Formula (Part 5)

This is the long-delayed fifth part in the xHR series. If you really want to read the first four parts, they can be located here, here, here, and here.

More than a month late, the highly anticipated follow-up to the first iteration of xHR has arrived. Once more, that increasingly trivial metric will grace the page of FanGraphs, wallowing in the mostly prestigious Community Research section (on the other hand, this section is most definitely the best section on the World Wide Web for experimental metrics and amateur analyses).

Unless the reader has an impeccable memory for breezily scanned, frivolous articles, he or she likely needs a reminder as to what xHR% is and aims to be. xHR% is a metric that describes at what rate a player should have hit runs over a given season. From this, expected home runs, a more understandable counting statistic, can be found by multiplying plate appearances by xHR%. It cannot be emphasized enough that the metric is not predictive; it only aims to describe. Without further ado, the formula is here:

I know that’s a lot to look at, and it isn’t exactly self-evident what all of the variables mean. As such, an explication of each part is necessary and provided below. (For logical rather than chronological purposes, the Kn variable will be analyzed last.)

AeHRD – One of the biggest differences between this formula and the last one is that this one does not use home run distance. This iteration uses expected distance, rendering it a combination of simple math, sabermetric theory, and physics. As such, expected home run distance strips out one of the biggest factors in luck — the weather.

Expected home run distance is found by utilizing a method taken from Newtonian Mechanics to calculate how far objects go. By using ESPN’s HitTracker website, I was able to obtain launch angles and velocities for nearly every home run hit in 2015. From this, I was able to resolve velocity into its respective parts, velocity in the x-direction (Vx) and velocity in the y-direction (Vy). After that, I calculated the amount of time the ball would be in the air with the formula vf=vi+gt, where vf is final velocity (0 m/s), vi is initial velocity (Vy), and g is simply the gravitational acceleration constant. Finally, I multiplied Vx by time in order to get the total expected distance.

I repeated that process for every home run hit by a given player in order to find his average expected home run distance. By doing this, I was able to strip out all weather-related components.

AeHRDH – Utilizing the same process as above, I found the average expected home run distance for every stadium. This is the player’s home stadium’s average home run distance, regardless of team.

AeHRDL – The same as above, but done for every home run hit in the majors last season.

When put together in the numerator and the denominator, the above variables serve as a “distance constant” of sorts that will at most adjust the resulting expected home runs by plus or minus two. Occasionally, the impact is negligible because the average expected distance is very close to that of the player’s home stadium and the league. Averaging the mean expected home run distance of the league and of the home stadium allows the metric to paint a more accurate picture of where the player hit his home runs and whether or not they should have left the park. Nevertheless, it’s important to note that this formula still fails to account for fly balls that fell just short of the wall due to the wind and other factors, meaning that there are still expected home runs unaccounted for.

FB% – If you remember correctly, or took the time to briefly review the previous posts, then you will recall that in the prior iteration of the formula there was a section very similar to this one. The only differences are that the weights on each year of data have changed (those are still somewhat arbitrary, however, but I am working on getting them to more precisely reflect holdover talent from past years) and the primary statistic used.

Previously, HR/PA was used, but it had to be abandoned because the results were too closely correlated with reality. This time, I looked at how similarly descriptive formulas were quantified. Oftentimes, those metrics did not use the target expected metric in their formulas. Rather, they utilized other metrics that correlated moderately well or strongly with their expected metric. In this case, I decided to use FB% because it’s a relatively stable metric (especially in comparison with HR/FB), and it has a strong correlation with HR% (about .6).

As a clarification, the subscript Y3, Y2, and Y1 indicate the years away from the season being examined, where Y1 is really Y0 because it’s zero years away. So just to be clear, Y1 is the in-season data from the year being examined. In the data to be examined, for example, Y1 is 2015, Y2 is 2014, and Y3 is 2013.

Kn – As you can well imagine, FB% numbers are always far greater than HR% numbers*, resulting in some truly ridiculous results if a constant isn’t applied that relates HR% to FB%. For instance, without a constant to modify the results, Jose Bautista would have been expected to hit 304 home runs last season. That’s a lot of home runs. Just two and a half seasons of playing at that level and he’d have the home run record in the bag. Luckily, I’m not stupid enough to think that that’s actually possible, and so I initially related FB% and xHR% with a constant, called KCon.

Unfortunately, KCon didn’t work as well as I’d hoped because it skewed expected home run results way up for terrible home run hitters and way down for the best home run hitters. By skewed, I mean bad by more than six home runs. And so I, in my infinite (and infantile) amateur mathematical wisdom, made it into a seven part piecewise** function. By this, I mean that there’s a different constant for each piece of the formula, defined by HR% at somewhat arbitrary, though round points. For clarity, here they are:

K1 = HR%<1

K2 = 1≤HR%<2

K3 = 2≤HR%<3

K4 = 3≤HR%<4

K5 = 4≤HR%<5

K6 = 5≤HR%<6

K7 = 6<HR%

It works quite well. I am very excited about the current iteration of xHR%, its implications, and all it has to offer. Of course, it is not finished, but I think I’m getting closer. Please comment if you have any questions, an error to point out, or anything of that nature. There will be a results piece published soon on the 2015 season, so keep an eye out.

*It wouldn’t be surprising if Ben Revere became the first player to have a HR% equal to FB% (both at 0%, naturally).

**It is neither continuous nor differentiable.


xHR%: Questing for a Formula (Part 2)

Part 2 of a series of posts regarding a new statistic, xHR%, and its obvious resultant, xHR, this article will examine formula 1. The primer, Part 1, was published March 4.

As a reminder, I have conceptualized a new statistic, xHR%, from which xHR (expected home runs) can be derived. Furthermore, xHR% is a descriptive statistic, meaning that it calculates what should have happened in a given season rather than what will happen or what actually happened. In searching for the best formula possible, I came up with three different variations, all pictured below with explanations.

HRD – Average Home Run Distance. The given player’s HRD is calculated with ESPN’s home run tracker.

AHRDH – Average Home Run Distance Home. Using only Y1 data, this is the average distance of all home runs hit at the player’s home stadium.

AHRDL – Average Home Run Distance League. Using only Y1 data, this is the average distance of all home runs hit in both the National League and the American League.

Y3HR – The amount of home runs hit by the player in the oldest of the three years in the sample. Y2HR and Y1HR follow the same idea. In cases where there isn’t available major league data, then regressed minor league numbers will be used. If that data doesn’t exist either, then I will be very irritated and proceed to use translated scouting grades.

PA – Plate appearances

(Apologies for my rather long-winded reminder, but if you really forgot everything from Part 1, then you should really invest in some Vitamin E supplements and/or reread the first post.)

The focus formula of this post is the first one, which also happens to be the one I think will work the least well because it relies too heavily on prior seasons to provide an accurate and precise estimate of what should have happened in a given season.

In the second piece of the formula, with only fifty percent of the results from the season being studied taken into account, it likely fails to take into account the fact that breakouts occur with regularity. As a result, it probably predicts stagnation rather than progress.

Methodology

Luckily for myself and the readers, the process was an incredibly simple one. Pulling data from FanGraphs player pages, ESPN’s Home Run Tracker, and various Google searches, I compiled a data set from which to proceed. From FanGraphs, I collected all information for Part Two of the formula, including plate appearances and home runs. Unfortunately, because a few of the players from the sample were rookies or had fewer than three years of major league experience, I had to use regressed minor league numbers. In some cases, where that data wasn’t applicable, I dug through old scouting reports to find translatable game power numbers based off of scouting grades (and used a denominator of 600 plate appearances).

Then, from ESPN’s amazingly in-depth Home Run Tracker website, I obtained all relevant data for player home run distance, average home run distance for the player at home, and league average home run distance. Due to my limited time, I only used players that qualified for the batting title during the 2015 season, yielding an iffy sample of only 130 players. Additionally, before anyone complains, please realize that the purpose of my research at this point is only to obtain the most viable formula and refine it from there.

Results

Using Microsoft Excel, I calculated the resultant xHR% and xHR. Some key data points:

League Average HR% (actual):  3.03%

Average xHR%:  2.85%

Average Home Runs: 18.7

Expected Home Runs: 17.7

Please note that there is a significant amount of survivorship bias in this data. That is, because all of these players played enough to qualify for the batting title, they are likely significantly better than replacement level, which is why the percentages and home runs seem so high.

Clearly, the numbers match up fairly well, with this version of the formula expecting that the league should have hit home runs at a .18% lower clip, and one fewer per player, which amounts to a significant difference. Over the course of a 600 plate appearance season, the difference between them is still only a little more than one home run, an acceptable distance.

Correlation between xHR% and HR%: 0.960506092

R² for above: 0.922571953

HR% Standard Deviation: 1.5769373

xHR% Standard Deviation: 1.3883746

Correlation between xHR and HR: 0.966224253

R² for above: 0.933589307

HR Standard Deviation:  10.43771886

xHR Standard Deviation: 9.201355342

While xHR% using this formula apparently explains about 92% of the variance, correlation may not be the best method of determining whether or not the formula works adequately. This holds at least for between xHR% and HR%, because there’s only a minuscule difference between their numbers (but one that matters), meaning it’s not a particularly explanatory method and that it may not have the descriptive power I’m looking for. Nevertheless, it is important to note that the correlation is not a product of random sampling, as p<.005. Unsurprisingly, the standard deviation for xHR% is smaller than that of HR% (nearly insignificantly so), indicating that the data is clumped together close to the mean as a result of using this formula, a potentially good thing (in terms of regression).

A better indicator of the success of the formula is the correlation between xHR and HR, a relatively high value of ≈.97. Here, presumably because the separation between home runs and expected home runs is greater, the formula ostensibly explains approximately 94% of the variance in outcomes and resultant data. However, in this case, the standard deviation for actual home runs is about 10.4, while for xHR it’s about 9.2, suggesting that, after being multiplied out by plate appearances, xHR is spaced nearly as evenly as HR. Ergo, it likely serves as a decent predictor of actual home runs.

Players of Interest

Mr. Bryce Harper – It’s likely there isn’t a better candidate for regression according to this formula than Bryce Harper, who the formula says have hit only 32 home runs as opposed to his actual total of 42. While he did lead his league in “Just Enough” home runs with 15, he’s also always been known for having prodigious power (or at least a potential for it). Furthermore, Mr. Harper dramatically changed his peripherals last season to ones more conducive to power. Suggesting this are the facts that he increased his pull percentage from 38.9% to 45.4%, his hard hit percentage from 32% to 40%, and his fly ball percentage from 34.6% to 39.3%. On their own, all of the previous statistics lend credence to the idea that Harper changed his profile to a more home-run-drive one, but when taken together they significantly suggest that. His season was no fluke, and the formula certainly failed him here because it weighted prior seasons far too heavily.

Mr. Brian Dozier – No surprises here. Mr. Dozier has certainly been trending upward for a long time, and in a model that heavily weights prior performance such as this one, upticks in performance are punished. Nevertheless, the data vaguely supports the idea that Dozier should have hit 24 home runs instead of 28. While he did significantly increase his pull percentage to an incredibly high 60% from 53%, he did play in a stadium where it’s of an average difficult to hit pull home runs as a right-handed hitter. Moreover, 10 of his 28 home runs were rated as “Just Enough” home runs, in addition to his average home-run distance being 12 feet below average (admittedly not a huge number, nor a perfect way of measuring power). If I were a betting man, I’d expect him to hit 4-6 fewer home runs this coming season.

Keep watch for Part 3 in the coming days, which will detail the results of the other formulas. Something to watch for in this series is the issue that the results of the formula correspond too closely to what actually happened, which would render it useless as a formula.

Note that because I have never formally taken a statistics course, I am prone to errors in my conclusions. Please point out any such errors and make suggestions as you see fit.


xHR%: Questing for a Formula (Part 1)

One of the most important developments in statistics — and its subordinate field, sabermetrics — is the usage of multiyear data to produce an expected outcome in a given year. It’s an old concept, one that’s been around for centuries, but it likely originated in sabermetrics circles with Bill James. In Win Shares (arguably the birth of WAR), the sabermetric response to Principia Mathematica, he details a procedure of finding park factors wherein the calculator uses a weighted average of several years of data in conjunction with league averages to find park factors for a certain ballpark.

Methods such as Mr. James’s allow the amateur sabermetrician (and even the mighty professional statistician) to determine what ought to have happened over a specific time period. Essentially, a descriptive statistic. The best example of a descriptive statistic for the unlearned reader is xFIP, which basically describes what a pitcher’s fielding-independent average runs allowed would have been if the pitcher had a league-average home runs per fly ball rate.

Several statistics fluctuate greatly from year to year and are thus considered unstable. Examples include BABIP, HR/FB% for pitchers, and line-drive percentage. HR/FB% in particular is very fluid because all sorts of variables go into whether a ball leaves the park or not. For instance, on a particularly windy day, an otherwise certain dinger might end up in the glove of an expectant center fielder on the warning track instead of in the beer glass of your paunchy friend in the cheap seats. Rendered down, xFIP takes the uncontrollable out of a pitcher’s runs-allowed average.

With this, and an excellent article about xLOB% from The Hardball Times, in mind, I started developing my own statistic a few days ago. xHR%, as I dubbed it, attempts to find an expected home-run percentage, and from there one can easily find expected home runs (xHR) by multiplying xHR% by plate appearances, a more understandable idea to the casual baseball fan. In order to calculate this, I wrote several different (albeit very similar) formulas:

More likely than not, your eyes glazed over in that section, so I will explain.

HRD – Average Home Run Distance. The given player’s HRD is calculated with ESPN’s Home Run Tracker.

AHRDH – Average Home Run Distance Home. Using only Y1 data, this is the average distance of all home runs hit at the player’s home stadium.

AHRDL – Average Home Run Distance League. Using only Y1 data, this is the average distance of all home runs hit in both the National League and the American League.

Y3HR – The amount of home runs hit by the player in the oldest of the three years in the sample. Y2HR and Y1HR follow the same idea. In cases where there isn’t available major-league data, then regressed minor-league numbers will be used. If that data doesn’t exist either, then I will be very irritated and proceed to use translated scouting grades.

PA – Plate appearances

(For the uninitiated, HR% is HR/PA)

Essentially, what I have created is a formula that describes home-run percentage. First off, I used (.5)(AHRDH) + (.5)(AHRDL) in the denominator of the first part because a player spends half his time at home and half on the road. If I were so inclined, I could factor in every single stadium that gets visited, weight the average of them, and make that the denominator, but that’s just doing way too much work for a negligible (but likely more accurate) effect. Besides, writing that out in a formula would be a disaster because then there essentially couldn’t be a formula. Furthermore, having half of the denominator come from the player’s home stadium factors in whether or not the stadium is a home-run suppressor or inducer, which helps paint a more accurate picture of the player.

Dividing the player’s average HRD by(.5)(AHRDH) + (.5)(AHRDL) allows the calculator to get a good idea of whether or not the player was “lucky” in his home runs. If his average home-run distance is less than the average of the league and his home stadium, then it follows that he is a below-average home-run hitter and his home-run totals ought to be lesser.

Since the values in the numerator and the denominator will invariably end up close in value to each other, I decided that this part of the formula could be used as the coefficient (as opposed to just throwing it out) because it will change the end number only slightly. Moreover, the xCo (as I call it) acts as a rough substitute for batted-ball distance and park dimensions in order to factor those into the formula.

The second part, the meat of the formula, uses a weighted average of multiple years of home-run-percentage data to help determine what should have been the home-run percentage in year one (the year being studied). Basically, it helps to throw out any extreme outlier seasons and regress them back a little bit to prior performance without stripping out everything that happened in that season (notice that in every formula the biggest weight is given to the season studied).

At this juncture, I cannot say for certain how much weight ought to be given to prior seasons. Obviously, a player can have a meaningful and lasting breakout season, with continued success for the rest of his career, making it inaccurate to heavily weight irrelevant data from a season two years ago. On the other hand, a player can have a false breakout, making it better to include more data from previous seasons. Undoubtedly that will be the subject of future posts. At present, the formula is a developmental one that will no doubt experience heavy changes in the future.

For the interested reader, some prior iterations of the formula are below:

As a reminder, with some small addenda, here is the explanation for each variable:

HRDY3 – Average Home Run Distance Year Three (year three being the oldest of the three years in the sample). HRD is calculated with ESPN’s home run tracker. HRDY2 and HRDY1 follow the same idea.

AHRDH – Average Home Run Distance Home. Using only Y1 data, this is the average distance of all home runs hit at the player’s home stadium by any player.

AHRDL – Average Home Run Distance League. Using only Y1 data, this is the average distance of all home runs hit in both the National League and the American League.

Y3HR – The amount of home runs hit by the player in the oldest of the three years in the sample. Y2HR and Y1HR follow the same idea. n cases where there isn’t available major league data, then regressed minor league numbers will be used. If that data doesn’t exist either, then I will be very irritated and proceed to use translated scouting grades.

PA – Plate appearances

(You should be initiated at this point, so figure out HR% for yourself.)

The reason these formulas were thrown out was that the xCo relied too heavily on seasons past to provide an accurate estimate. When I briefly tested this one on a few players, it delivered incredibly scattered results. Furthermore, there wouldn’t be any data available for rookies to use these iterations on because there’s no such thing as a minor-league or high-school home-run tracker (and if there were I probably wouldn’t trust it). The first formulas described are overall more elegant and more accurate.

Stay tuned for Part 2, when results will be delivered instead of postulations.


Peter O’Brien’s Raw Power: Estimating Batted-Ball Velocities in the Minor Leagues

On May 20th Peter O’Brien hit a massive home run to straight away center clearing the 32 foot tall batter’s eye at Arm & Hammer Park more the 400 feet from home plate.  O’Brien is currently 1 home run behind Joey Gallo, in what looks to be an exciting competition for the minor league home run title.  O’Brien isn’t as highly touted a prospect as Gallo, but he still has some of the most impressive power in the minor leagues.  Reggie Jackson saw O’Brien’s home run and said it was one of hardest hit balls in the minor leagues that he had ever seen (and Reggie knows a thing or two about tape measure home runs).

How hard was that ball actually hit?  It is impossible to figure out exactly how hard and how far the ball was hit from the available information.  You can however use basic physics to make a reasonable estimation.

Below I explain the assumptions and thought process I used to get to an estimate of how hard the ball was hit.  If that does not interest you, then just skip to the end to find out what it takes to impress Reggie Jackson. But, if you’re curios or skeptical stick around.

OBSERVATIONS

I started off by watching the video to see what information I could gather (O’Brien’s at bat starts at the 37 second mark in the video).

TIME OF FLIGHT From the crack of the bat, to the ball leaving the park – it appears to take 5 seconds. If you watched the video, you can tell this is not a perfect measurement since the camera doesn’t track the ball very closely. If you think you have a better estimation, let me know and I’ll rework the numbers.  

LOCATION LEAVING THE PARK  The ball was hit to straight away center. From the park dimensions we know when it left the park it was 407 feet from home plate and at least 32 feet in the air to clear the batter’s eye.

ASSUMPTIONS

COEFFICIENTS OF DRAG (Cd) – The Cd determines how much a ball will slow down as it moves through the air. I chose 0.35 for the Cd because it is right in the middle of the most frequently inferred Cd values for the home runs that Allan Nathan was looking at in this paper.In looking at the Cds of baseballs, Allan Nathan showed there is reason to believe that there is some significant (meaning greater than what can be explained by random measurement error) variation in Cd from one baseball to another.

ORIGIN OF BALL I assume the ball was 3.5 feet off the ground and 2 feet in front of home plate when it was hit.  These are the standard parameters in Dr. Nathan’s trajectory calculator. But what if the location is off by a foot? The effects of the origin on the trajectory are translational. One foot up, one foot higher. One foot down, one foot lower. The other observations and assumptions are more significant in determining the trajectory of the home run.

Using these assumptions and the trajectory calculator, I was able to determine the minimum speed and backspin a ball would need in order to clear the 32 foot batter’s eye 5 seconds after being hit at different launch angles.  The table below shows the vertical launch angle (in degrees), the back spin (in RMPs) and the speed of the balled ball (in MPH).

Vertical launch angle Back spin Speed off Bat
19 14121 101
21 6817 101.9
23 4155 102.75
25 2779 103.69
27 1940 104.7
29 1375 105.89
30 1156 106.5
32 805 107.88
34 536 109.4
36 322 111.1
38 149 112.99
40 4 115.1

The graph shows a more visual representation of the trajectories in the table above (with the batter’s eye added in for reference).

http://i1025.photobucket.com/albums/y314/GWR87/OBrienhomerun_zpsb1507cf4.png

Looking at the graph you will notice that all of these balls would be scraping the top of the batter’s eye.  This makes sense because the table shows the minimum velocities and back spins needed for the ball to exactly clear the batter’s eye.

What is the slowest O’Brien could have hit the ball?

If you were in a rush, looking at the table you would think the slowest O’Brien could have hit the ball would be 101 MPH at 19o. But, not so fast! The amount of backspin required for the ball to travel at that trajectory is humanly impossible.

What is a reasonable backspin?

I am highly skeptical of backspin values greater than 4,000 rpm based on the Baseball Prospectus article by Alan Nathan “How Far Did That Fly Ball Travel?.” The backspin on home runs Nathan examined ranged from 500 to 3,500 rpm, with most falling in around 2,000. The first 3 entries in the table have backspins of over 4,000 and can be eliminated as possibilities. If the ball with the 19o launch angle only had 3,500 rpm of back spin it would have hit the batter’s eye less than 11 feet off the ground instead of clearing it.  Maybe you’re skeptical that I eliminated the 3rd entry because it’s close to the 4,000 rpm cut off.  Think about it this way, if a player was able to hit a ball with over 4,000 rpm of back spin, they would have to be hitting at a much higher launch angle than 23o (Higher launch angles generate greater spin while lower launch angles generate less spin).

The high launch angle trajectories with very little back spin (like the bottom three in the table) are also not very likely.  A ball hit with a 40o launch angle would almost certainly have more than 4 rpm of back spin.  If the ball hit with the 40o launch angle had 1,000 rmp of back spin (instead of 4) it would have been 70 feet off the ground, easily clearing the 32 foot batter’s eye.

Accounting for reasonable back spin, the slowest O’Brien could have hit the ball is 103.69 MPH at 25o with 2,779rpm of backspin.

So what do all these observations and assumptions get us?

We can say that the ball was likely hit 103.69 MPH or harder, with a launch angle of 25o or greater.  103.69 MPH launch velocity is not that impressive, it is essentially the league average launch velocity for a home run.  Distance wise, how impressive of a home runs was it? Unobstructed the ball would have landed at least 440 feet from home plate (assuming the 25o scenario).  The ball probably went further than 440 because it did not scrape the batter’s eye. So, how rare is a 440+ foot home run? Last year during the regular season there were 160 home runs that went 440 feet or further, there were a total of 4661 home runs that season, meaning only 3.4% of all home runs were hit at least that far.

For those of you who wanted to just skip to the end. My educated guess is that the ball went at least 440 feet and left the bat at at least 103.69 MPH.

If you like this, you can read other articles on my blog GWRamblings, or follow me on twitter  @GWRambling

None of this would have been possible without Alan Nathan’s great work on the physics of baseball.  I used his trajectory calculator to do this, and I referenced his articles frequently to make sure I wasn’t way making stupid assumptions. The information on major league home run distance is based off of hittrackeronline.com


What’s Behind A-Rod’s Power Outage: The Sequel

When Yankeeist last looked at Alex Rodriguez’s declining power numbers, I (and several others) came to the rather obvious conclusion that his paltry 8.3% HR/FB rate would soon escalate. Alex’s five home runs since that point in time have indeed bumped his rate up, but it’s only sitting at 12.1%, still well below his 23.3% career percentage.

A-Rod has eight home runs on the season to date; the lowest number through 58 games of his career and only the second time he has accumulated less than 10 this deep into a season — in 1997, he had nine through 58 games with the Mariners. So what’s going on with Alex?

The below table shows historical batted ball numbers for A-Rod, his year-to-date home run totals (in this case, through the first 58 games of each season), and his season home run totals (all data c/o Fangraphs and B-Ref):

Despite five big flies, Alex’s fly ball percentage is down from when I last looked at the numbers on May 10. Accordingly, his line drive percentage is also down, to 17.9% (though this is barely off his career rate) and his ground ball percentage is up, to 46.2% (pretty well above his 42% career rate).

As you can see, Alex has never had a Fly Ball % this low in a full season for as long as Fangraphs has recorded this data, which partially explains why his HR/FB rate has only risen by 3.8 points — he’s just not hitting as many fly balls as he usually does. Assuming his Fly Ball % normalizes to his career rate, we should see a corresponding uptick in the HR/FB percentage.

Here are the different pitch types A-Rod has seen:

Pitchers are obviously aware that A-Rod isn’t hurting the baseball as much as he usually does, as they are challenging him with more fastballs than ever before. Correspondingly he’s seeing less of every other pitch type since May 10, with the exception of a slight increase in changeups and split-fingered fastballs. Looks like the book on ‘Rod remains challenging him with the heater, which means he’s going to have to make some adjustments to his approach, as there’s no reason Alex shouldn’t be able to adequately handle a steady diet of fastballs.

And here are his swing percentages:

Since I last conducted this analysis, Alex is swinging at even more pitches out of the zone (25.8%) but making less contact with them (60.1%), and also swinging at more pitches in the zone (65.7%) and making less contact with those as well (91.8%, down from a crazy high of 97.3%). His overall contact percentage is 81.4%, still a good deal higher than his career rate of 75.5%.

It would appear Alex’s biggest problem is that he’s trying to make too many things happen with the bat right now — swinging at pitches out of the zone has contributed to an above-average (for Alex) contact rate, which is resulting in more balls being pounded into the ground than lofted into the air (hence the career-low Fly Ball %).

Alex has also eschewed his trademark patience during the past month. He had 19 walks through 31 games, but has only walked seven times since then over his last 27 games. His OBP has dropped from .381 on May 10 to .360.

While A-Rod still has time to improve his numbers, and ZIPS ROS projection has him hitting a robust .284/.378/.512, .392 wOBA and 18 home runs the rest of the way, that would still only get A-Rod to a full season line of .285/.371/.499 with a .381 wOBA and 26 bombs, which would mark his lowest SLG, home run total and wOBA since 1997.

Basically, A-Rod needs to stop swinging at bad pitches, take a few more walks and show pitchers he can still punish the fastball if we’re going to see significant improvements in the Fly Ball % and HR/FB rate and get his numbers anywhere near his career line of .304/.389/.573. I realize that’s a rather obvious conclusion that probably didn’t require a comprehensive statistical analysis, but it’s nice to see that the numbers support it.

Larry Koestler eats, drinks, sleeps and breathes the Yankees at his blog, Yankeeist.