Archive for Projection

Replacing Replacement Value in Fantasy Auctions

With the baseball season rapidly approaching and recent posts by FanGraphs authors converting projected statistics into auction values, I thought I would share my approach towards valuation I have used in a long-standing A.L. league with 12 teams, 23 player rosters selected through auction (C, C, 1B, 3B, CI, 2B, SS, MI, 5 OF, 1 DH), a $260 budget, a 17-player reserve snake draft and the ability to keep up to 15 players from one year to the next, an attribute that inflates the value of the remaining pool and can further distort disparate talent across positions and categories.

We have traditionally used a 4×4 format, and while I have persuaded my co-owners to switch to a 5×5 for the coming year, what follows is my process for a 4×4 league.

There was a distant time when I was a whiz at math but my utter lack of a work ethic for advanced math collided with university-level calculus and I crumbled as surely as a weak-kneed lefty facing Randy Johnson. So my understanding of some key statistical processes is compromised. And by some I mean most.

But what I lack in math I hope I make up in approach:

(1) For categories over multiple years in this league, teams finish in a standard bell-shaped curve, with two or three teams well ahead, two or three well behind and six to eight clumped more closely together.

(2) In a 12-team league, a third-place finish in a category bets you 10 points. Across eight categories, averaging a third-place finish gets you 80 points, which is enough points to win out league between 80% and 90% of the time.

(3) Given both (1) and (2), my goal is to finish in third in every category, because doing do will far more often than not win my league, and because that target is a comfortable space above the pack in the middle, creating a margin for error within which I can still secure a win.

(4) I calculate what totals I need for each category to finish third based upon the specific history of our league, giving greater weight to more recent and relevant trends.

(5) I calculate the totals needed to finish dead middle in the pack for each category, again based upon the specific history of our league, giving greater weight to more recent and relevant trends.

(6) The difference between the third-place totals and the median totals become my spread, in a sense, the yardstick against which I then measure all projected player performance.

(7) I don’t weight pitchers and hitters evenly because my league does not – the marketplace of my league places significantly less value on pitchers, spending between $70 and $100 on them, and I adjust values to account for that. Perhaps that is also justified by either greater volatility or more injuries for pitchers. In any case, I divide the total value for hitters by 14 and for pitchers by 9 to come up with the average value for hitters or pitchers.

(8) I calculate what each of 14 hitters and 9 pitchers would need to contribute per player for each category for both the top and the bottom of the spread.

(9) For each category, I divide the median production per player by the difference in the gap to find the incremental value of each unit of production.

(10) For each player and for each category, I start with the median value of median production for all four categories, than add or subtract the incremental value depending upon if their projected production is above or below the median.

(11) I do the same for keepers to calculate inflation value, then list both the value and inflated value next to each player, broken down by position, so I can track both availability and the ebb and flow of inflation in real time.

(12) Finally, my league is mostly inelastic except for dumping trades. That means it is not easy to trade surplus categories for deficit categories. So I create a running tally of my projected production, starting with my keepers and adding players I gain in the auction with the goal or at least reaching each of the target levels needed for projected third-places finished in each category.

(13) I don’t adjust assigned value based on the position played but of course I consider position as I bid in order to reach my targets in an inelastic league. I may deliberately pay somewhat more than inflation cost for a good player if the likely alternatives is paying over inflation value for a poor player and being left with more money to spend then there is talent to spend it on. I do so knowing my keepers will produce to much surplus value that I can win simply getting players close to inflation value.

At least in my league, my projected values, adjusted for inflation, are pretty close to the mark notwithstanding the outliers that will come in any marketplace, both for individual players and for more systemic biases (my league overpays for closers, for example). I don’t win every year, but when I fall short, it is not because my valuations were off but because of too many failures in projecting specific players.

Is there a statistical basis for tossing replacement value as a baseline for creating auction values or statistical benefit to instead using league-specific gaps between middling and winning teams? Frankly, I don’t know, however intuitive my system seems to me. But I’d welcome feedback on my approach, statistical arguments for and against it, and whether it warrants further exploration.


Application of my fWAR League Adjustment Method

This article is a follow-up to my previous one in which I will work through some examples. You should try to get an intuition on it. If the concept seems too complicated I have to apologize for not explaining myself well because I sincerely think this is very straightforward and no voodoo and could help improve fWAR even further… which is mindboggling if you think about it. It could improve projection systems as well as the correlation of WAR and actual wins while also handling players changing from the AL to the NL or vice versa more elegantly.

I will simply follow my steps 1-4 from my previous article to figure out the proper league adjustment and continue with some WAR calculations. I will use the 2014 season as my guinea pig.

While playing around with it I also stumbled upon a wRC+ adjustment that has to be done because of a) the independence of both leagues and b) the differing league strengths. I will tackle this issue in my next article.

All right, here are steps 1) –4).

1) I need to figure out the wOBA values, R/PA, FIP, R/W, cFIP for each league individually. These can normally be found here. I will not list every single wOBA value here because that doesn’t add much to the explanation and saves me some time.

AL (2014):

wOBA: .312

R/PA: .110

FIP: 3.82

R/W: 9.25

cFIP: 3.16

 

NL (2014):

wOBA: .308

R/PA: .105

FIP: 3.66

R/W: 8.97

cFIP: 3.10

 

The exact values for all of MLB found on the Guts! page is conveniently exactly the arithmetic mean of my AL and NL values.

2) All right, we now move on to step 2 which is to figure out the interleague record. I suggested that a 3 year rolling regressed average could be a possibility with years N-1, N and N+1 as inputs. I cannot see into the future, for that reason I will simply use the 2012-2014 interleague record based on pythagenpat. This comes out to a .539 W% for the AL. Conveniently, the actual W% is exactly the same. For demonstration purposes let’s just do a farmer’s regression and call that a “true talent” .530 W%.

3) This is the seemingly tricky part but once you got your head around it is is very easy to grasp. As a reminder: the three necessary “true” replacement levels needed for all WAR calculations are .294 in general for teams – this is where the fixed 1,000 WAR each year comes from – the .380 replacement level for starting pitchers and the .470 for relievers.

Imagine an NL team that is a .500 team within the NL. This team plays a .500 AL team within the AL. That needs to be stressed. Those teams are NOT of equal strength, even if both have a .500 record. Why, you ask? Because if they were, we would not see an advantage for the AL in interleague play. We would see a balanced .500 interleague record. That is not our reality and we can confidently conclude that the NL is the weaker league as of today.

Following this line of thought, what happens if two replacement teams out of each league play each other? Well, this means a .294 NL team plays a .294 AL team. What would the outcome be? A .530 winning percentage in favor of the AL. This comes straight out of the interleague record.

How much better than a .294 W% would this NL team have to be in order to win exactly half of its games against this .294 AL team? This is where the odds ratio comes into play and it spits out a .320 winning percentage. That means if a .320 NL team faces a .294 AL team in an environment, in which the AL wins 53% of all interleague games, we would finally expect parity. A .500 interleague record. This .320 is our new “artificial” replacement level for the NL in 2014.

On the other hand we have to ask the question: How much worse than a .294 can an AL team be when facing a .294 NL team and still win half of its games? Odds ratio says a .270 AL team would still win 50% of all games against a .294 NL team in a context where the AL wins 53% of all interleague games. This .270 is our new “artificial” replacement level for the AL in 2014.

4) Remember that our “regressed” interleague record suggests the AL to be the stronger league, thus worthy of receiving more share of the WAR-pie. Now it is time to figure out how much more they deserve.

We figured out a .270 “artificial” replacement level for the AL. Therefore, we can distribute (.500-.270)*15*162 = 559 WAR towards the AL. This is split up 57/43 between position players and pitchers.

In the National League we found a .320 “artificial” replacement level. Therefore, we can distribute (.500-.320)*15*162 = 437 WAR towards the NL. Same 57/43 split.

Now 559+437 = 996, which is not equal to 1,000. This is because of the odds ratio being non-linear the closer it gets to the extremes but I might be totally mistaken here. This usually is where Tangotiger appears out of the dark and helps out with fancy math or steps in when the math gets hurt. I don’t really see it as a problem.

We could either distribute the remaining 4 WAR 50/50 between both leagues or adjust the replacement levels slightly to arrive at exactly 1,000 WAR. Both would change individual WAR figures only on an atomic level.

I want to point out that this kind of inconsistency is very common in the implementations of WAR. rWAR and fWAR both have some adjustment runs to match inconsistencies like that. This doesn’t even make a difference on a player level. It would not even change a team’s WAR figure by 1/10 I guess.

WAR calculations

After you have come this far you are probably interested in how much certain player’s WAR figure might change. Again, I won’t list every step necessary but only the actual results. If you ask yourself how I have done it, you should take a look here, here and here. If that doesn’t help out, just comment with your question and I will walk you through.

My example will be Mike Trout. I will show the differences of some of the more important and interesting stats as (OLD/NEW). Forgive me for not being a formatting wizard.

NOTE: For sake of better comparison I will present the “new” run values with an exchange rate of 9.117 R/W (currently used). Otherwise 1 run wouldn’t have the same meaning since in my WAR calculations 1 win equals 9.25 runs.( See step 1  ) This makes this an apples to apples comparison.

Trout:

wOBA: (.403 /.402)

wRC+* : (167 / 170)

WAR**: (7.8 / 8.0)

batting: (52.1 / 54.0)

UBR: (3.0 / 3.0) unchanged

wSB: (1.8 / 1.7)

Fld: (-9.8 / -9.8) unchanged

Pos: (1.4 / 1.4) unchanged

Lg: (2.9 / 2.9)

Rep***: (19.9 / 19.9 )

 

 

*  I use a slightly different wRC+ calculation here. My league adjustment method would also improve the accuracy of wRC+ as a comparison tool between the two leagues. I will write another article dealing with the modified wRC+ calculation, as well as the wRAA and replacement runs modifications to improve the accuracy of fWAR.

** Fielding runs, UBR and positional adjustment were not changed. These three will never change, the league adjustment however will undoubtedly change, as well as wSB, although the changes would be tiny. It involves complete league stats, i.e. every single player’s stats.

*** The value of replacement runs will never be affected in my league adjustments even though I use different replacement levels for my calculations. Replacement runs will always be based on the .294 baseline. I hope this makes sense to you. If not I point out to the upcoming article of mine.

Outlook

In my next article I will lay out the modifications that have to be applied to wRAA, wRC+, batting runs and the replacement runs. I will show why my modifications make wRC+ more accurate in comparing both leagues and explain why this new league adjustment influences position player WAR more than pitcher WAR. Because right now, the fWAR-process for pitchers leans heavily, not entirely though, towards the independency treatment of both leagues – a cornerstone of my league adjustments.

Also look forward to a table of the players with the biggest and the smallest increase in WAR and the corresponding losses. In both the AL and NL there are players who gain or lose more than others. This has to do with the different run environments is my best educated guess so far. In the NL – the lower scoring league – extra-base hits become slightly more valuable. So does base-stealing. Opposite for the AL. So look forward to my next piece, fellows!


Ottoneu Tools: Advanced Standings Part Two

In early May I introduced Ottoneu players to the Advanced Standings Dashboard, a tool that allows team owners to decipher the early season standings in an effort to better gauge where their team might be headed as the 2014 season comes together. You can download that tool here (http://goo.gl/pbXI5), but now that we’ve just entered July, the traditional halfway point of the baseball season, it’s time to take a deeper look at a few ways this tool can be used to effectively to manage your team into contention in the second half.

Since the tool can be updated easily with just a couple of copy/paste actions, I use this tool almost daily in my own FGPoints Ottoneu league.  But for fun, let’s walk through a few features as they apply to the FanGraphs Staff League, with a special focus on Eno Sarris’ team, “It’s A Perm“.

Eno enters July as a 3rd place team, nearly 400 points out of 1st place, and 150 out of 2nd.  In general, with at least seven teams over the 8,000 point mark, this league looks competitive at a glance.  But with the recent pickup of Ryan Braun, Eno clearly has his sights set on a title (https://twitter.com/enosarris/status/483016142831644672), so let’s break down the standings using the tool to see if Eno has the momentum to win it all in the 2nd half.

The first tab of the tool is simply the statistical breakdown of the Ottoneu standings into some common sabermetric calculations.  While we can easily see Eno leads the league offensively at 5.44 P/G, the underlying statistics also support it, showing he maintains an (slight) advantage in OPS, OPS+, wOBA, Runs Created, and Total Bases.  What may be more interesting is that Eno has more points scored from his offense than any other team in the league.  In fact, just over 58% of his points have come from his hitters (tab 3, ‘Projected Finish”). With roughly 55% of league scoring in Ottoneu coming from offense, Eno is clearly banking on this approach of shoring up the side of the ledger that carries the most weight.  The acquisition of Braun will only help.

So It’s A Perm is built on bats, but what about the pitching? Unfortunately, this is a weak spot, as Eno’s FIP, WHIP, and BB/9 are all higher than the two teams he’s chasing.  I’m sure he knows this instinctively as his 5.03 P/IP is below the league average of 5.13 P/IP (and further below the top 7 teams of 5.19 P/IP), but the dashboard makes it quicker and easier to point out these pitching deficiencies.  One possible area of improvement: the bullpen.  Without looking at his roster, I can tell you pretty quickly he’s probably pretty frustrated with his bullpen, which has been almost 42% less effective (“PEN” = Saves + Holds/IP) than the league leader, A Little Out of Context.  Shoring up a bullpen is often easier and cheaper than finding an ace SP mid season, so does Eno speculate on the eventual Sergio Romo replacement? Does he approach John Heyman’s Last Sirloin about shedding some of his bullpen pieces in a plea to “deal from strength”?

Once you’ve taken the time to digest some of the traditional sabermetric outputs in the Dashboard, your eyes will naturally gravitate toward the end of the first tab into the “League Projections” section, which is where the real power of the tool comes alive.  The key takeaways here are the “Otto” score and the “Pace” columns.  The Otto score can be better explained here by Chad Young (http://goo.gl/KK4Xy), while the “Pace” attempts to project the season-ending point totals for each team based up a range of factors, including current P/G and P/IP values, remaining IP and GP, and league averages in these areas.  In many leagues these are the columns that can better identify contenders from pretenders, but for the FanGraphs Staff league we see more evidence that the actual standings are, for the most part, very accurate, as Eno is also projected to end the season with the 3rd most points (18,042, or about 400 points out of 1st place).

There are a few interesting things to note here, however. First, John Heyman’s Last Sirloin actually has the third highest Otto score (13.66), but is still projected for 4th place, most likely due to his slower pace in IP (1,416 projected).  If this team can pick up the IP pace in the 2nd half with similar quality IP (5.54 P/IP), this team could make up ground quickly.  This team is clearly riding a league-best bullpen and trying to maximize its RP innings as much as possible.

Second, Ground Rule Double Helmet, despite sitting in 4th place with a strong 9,000 points, has had to overtax a very week pitching staff (4.74 P/IP) just to get there (1,577 IP projected).  The tool sees as much and projects this team to end the season in 5th place, but unless the pitching staff sees a significant improvement in the 2nd half, I’d expect this team to possibly fall even further as the season shakes out.

And that’s just the first tab…Once you get familiar with the tool, you’ll actually find the third tab, “Projected Finish” to be the most useful summary of some of these features described above, as it will give you a daily update of the projected champion for the league.  With Eno just 400 points out of both the actual and projected season-ending standings, this league is just too close to call on July 1st, but there are at least four clear contenders here, and It’s A Perm is one of them.  Will Ryan Braun help the cause? Just for fun, let’s say Braun increases Eno’s offense by just 2.00% (from 5.44 to 5.55).  Well, that could be all it takes, as that small increase moves the needle for It’s A Perm enough to overtake Johan Santa Claus by 100 points in the projected season-ending standings, and less than 200 points out from 1st place.  Of course, that’s if everything else stays the same, and, as in life, the only thing constant in baseball is change.  This will be a fun league to watch as the summer heats up, so enjoy the tool and use it where possible to get that 2% edge.


Ranking Batters in Fantasy Leagues with Alternate Stats

Draft prep: Framing the problem

So you’re preparing for your fantasy draft. You’re caught up on FanGraphs, checked for recent injuries at Rotoworld, maybe skimmed a few headlines from your other top 11 baseball news sites. Maybe you’ve even downloaded the FanGraphs positional rankings, and are planning to keep the file open during the draft as a reality check against the pre-set rankings of the site your league uses.

But really, what do the guys at FanGraphs know? Sure, they know a lot about baseball, and statistics, and this year’s projections, and a handful of underlying stats that tend to predict future performance. But what they don’t know is whether your league uses OBP instead of AVG, or OPS, SLG, or batters’ strikeouts, or maybe holds and FIP and pitcher fielding percentage. If this is your situation, then I feel your pain. My fantasy league uses eight statistics for batters and pitchers, three each beyond the usual five. (In case you’re curious, the mysterious six are: Batter hits, K’s, & OPS; Pitcher holds, losses & complete games).

These differences matter. If your league uses OBP, Joey Votto turns from a fantasy player who’s solid in four categories (including average, where his impact is limited because he walks all the time) to a guy with a truly elite skill. Maybe it’s easy for you to account for the relative value of a Joey Votto, but how well can you project the 25th through 35th outfielders? Some might be much better or worse in your league. If you have batter strikeouts, as in my league, how do you value Mark Trumbo and his home run power against the elite contact skills of Norichika Aoki?

Generating your own rankings

One answer, and the one I opted for, is to generate rankings based on your own league’s stats. Now, this may sound a bit too work-intensive and time-consuming for most of you (especially those of you with relatively normal priorities), but in reality it wasn’t as time-consuming as I expected.*

First of all, there’s no need to reinvent the wheel. There are lots of projection systems out there that are available to the public, and some of them are quite good. I decided I would simply download all the projections listed on FanGraphs, and average them out. And then, after thinking for a little while about the costs and benefits of that approach, I decided I wouldn’t do that at all, and instead would use the results of just one projection system. But which one should I use? Luckily, that’s yet another bit of analysis we don’t need to bother with, because the Interwebs are full of crazy mathematicians who love baseball and have nothing better to do. After searching for a few articles that evaluate projection systems, like this one and this meta-one, I decided that the forecasts I trusted most (and were easiest to obtain) were Steamer for batters and FanGraphs fans for pitchers. (The high accuracy of the latter shocked me at first, but then I realized that fans assimilate the results of all the projection systems into their own player projections, departing from them only as dictated by common sense, inside scoop, and hope.)

Operationalizing the Solution

Here’s where it gets tricky. What advanced data manipulation packages and techniques are best for downloading reams of data from the FanGraphs site into your spreadsheet? Certainly there was no need for me to copy and paste the data 50 players at a time like someone living the dark ages, was there? No, of course not. And I probably never really did that.

Instead – bear with me if you’re not technically inclined – I hit the gray “Export Data” button to the upper right of my chosen projection page. This involved a lot of loading the correct page, hovering my mouse over the text, and clicking, but in the end it was worth all the work, because 5 minutes of sweat, plus a beer, had finally paid off in spreadsheets full of data.

*If you’re not interested in these details, the fun stuff is posted in a couple of tables towards the end. (I like writing, so this is likely to go on for a while.)

Z-scoring your data points

Z-scoring batter projections is easy. The problem lies in determining what set of players to use in order to calculate means and standard deviations.

This is an important question, at least to the extent that any question in fantasy baseball is important. For example, if you must use every hitter in the league, including the guys projected for 8 at-bats, you create the illusion that lots of players bat .220 or score only 4 runs, as opposed to your league’s reality in which .270 with 70 runs is pretty ordinary. For a little math fun, I compared the results generated using means and deviations 500 players deep (the equivalent of a 25-team league that rosters 20 position players) versus one with more reasonable assumptions. It caused huge increases in variance in runs and rbi’s, so a guy who drove in and scored 100 compared no better to the mean either way (~2+ standard deviations), but smaller increases in the variance in SB’s, HR’s, and OPS, which, together with the lower means, meanings this system overvalues guys who produce in these categories. Martin Prado and Torii Hunter were made sad, whereas Billy Hamilton was elevated to a demigod (or at least a top-40 hitter).

So how do you generate values that represent your player pool?

One method – and a very reasonable one – is to use the final statistics compiled by your league the previous year. With this data, it’s easy to generate per-slot averages based on last year’s performance, and to compare projected performance against it. But I did not choose this method. A more savvy number-cruncher might say that projection systems, while designed to be as accurate as possible for each player, may be systematically biased on the whole, and therefore determining the value of this year’s projections based on last year’s actual statistics is tantamount to comparing apples and oranges.

I was more worried about lazy owners. Any league can have a couple of careless owners who are in it just for fun (the gall!), or who keep BJ Upton when he can’t even see the Mendoza line, because of that one time his cousin shook BJ’s hand at a Jay-Z concert. I know of what I speak. If your goal is to win your league, you want to base your evaluation on the best players available, rather than the happenstance of which Atlanta outfielders spent the whole year on someone’s roster.

I generated means using very precise data, plus a random stab in the dark. First, I looked up the exact number of players at each position in my league from the previous year. Then I mostly ignored this data. Although it’s true that player values vary greatly between leagues depending on how many players start, and how many are rostered, this is the sort of thing you can keep track of during the draft. Don’t draft another first baseman if you already have three of them and no shortstop, and don’t draft a first baseman just because he’s ranked ahead of a shortstop if there are another seven first basemen ranked close behind.

My league rostered only 123 regulars last year. Not a deep league. I used a lot more than 123 in my calculations in an effort to lower the means a bit, to account for the existence of catchers and second basemen. I then haphazardly created sort variables so I could bring the best 150 to 180 players to the fore, with the goal of getting a fair representation of the quality of players in my league. I tried various formulas like [(HR+1) * R * RBI * (SB +1) * AVG * OPS] (adding 1’s so as not to exclude players projected for 0 HR’s or SB’s ) and PA * wOBA. Virtually every one of them produced a good representation of the best hitters projected for regular playing time. In the end, the best way to evaluate the sort is to look at the list and see if the guys near the cutoff are fringe players who are familiar from last year’s waiver wire.

Calculating projected player values

Once you determine which players you want to include, Excel is happy to instantaneously calculate averages and standard deviations for each stat. Once you have these values, you can re-include the entire player pool, or as much of it as you wish, and the formula for each player in each category is simply (his projected value – the average projected value)/standard deviation.

The next challenge is to generate ranks from the Z-scores. The simplest way is simply to add them together (being sure to subtract ones where lower scores are better, such as pitcher walks or batter strikeouts). But here, I discovered another issue. A potential superstar who might not have a full-time job could end up ranked about the same or below a mediocre player who was guaranteed to start. If I wanted my draft rankings to make sense at a glance when I have just 90 seconds to pick a player while eating a sandwich, I needed to distinguish accumulators from guys with potential.

Ranking performance and potential

It matters whether a player is an okay guaranteed performer or a unpredictable potential star. If I find myself with no second basemen in the 22nd round, I might want to take the best guy who’s pretty much guaranteed 140 days in the starting lineup, like an Anthony Rendon or a Howie Kendrick. If my roster’s pretty much set, I might prefer a hitter who has a better chance to bust out and hit 45 home runs, like Chris Carter (unless I’m in my league, in which his 80% strikeout rate falls 37 standard deviations below the mean).

What I decided to do was generate two rankings for each batter, one based on projected totals, and one based on projections per plate appearance. Luckily, Steamer has already done the work for us by projecting everyone in both ways. For instance, Everth Cabrera is projected as the 479th-best player by wOBA, with 74 runs and 45 stolen bases. At the other extreme, Colorado’s Kris Parker is projected to be the 50th-best hitter in the league, just ahead of Dustin Pedroia, with a .279 batting average and .465 slugging percentage, despite getting only one plate appearance, and not getting a hit.

At this point, there are 2 sets of columns for each batter: 1 set of columns for his Steamer projections for each relevant stat, and 1 for the associated Z-scores. To this, I added 2 more sets of columns: 1 for per plate-appearance projections for each stat, and 1 for those associated Z-scores. (Dividing hits into plate appearances rather than at-bats feels unnatural, but that’s what you need to do if your league counts total hits.) Calculating per-PA quality is then easy, as you can just add the Z-scores (or subtract for negative statistics). But once you have projected rate statistics in your per-PA rankings, it becomes apparent that it doesn’t make sense to include the exact same values in your projected accumulated totals.

To handle this, I weighted the Z-scores for the rate stats. I multiplied the Z-score for AVG by projected AB’s/average projected AB’s, and you can do the same for OBP, using PA’s. My league uses OPS, a value generated by adding two fractions with different denominators (aka OBP & SLG), so to weight those Z-scores I multiplied them by projected (AB’s + PA’s)/average projected (AB’s + PA’s). I then added these weighted Z-scores to the other Z-scores for projected totals. The result of adding these weights is that a player who is one standard deviation above average in both AVG and OPS, and who has an average number of AB’s and PA’s, would get +2 from these categories in the variable used to rank projected totals. By the same lights, the aforementioned Kyle Parker’s AVG and OPS would essentially get no weighting at all, and have no effect at all on his projected totals, just as in real life his performance is not expected to have any effect at all on the rate stats of your team.

The Fun Stuff

And that’s about it. Once you have Z-scores, it’s very easy to rank players, to change the formulas to rank them by different systems, or to sort players by certain categories to see who stands out the most.

Two common variations on the traditional 5 stats are to include OBP instead of AVG, or to play in a points league. (For a points league, just change the Z-score weighting to reflect the point system). Here are the top players in these alternate systems using this evaluation method (I threw my own league in too, just for kicks):

Rank Trad 5 OBP 5 Points Crazy 8s
1 Miguel Cabrera Miguel Cabrera Miguel Cabrera Miguel Cabrera
2 Mike Trout Mike Trout Mike Trout Mike Trout
3 Carlos Gonzalez Carlos Gonzalez Joey Votto Carlos Gonzalez
4 Yasiel Puig Paul Goldschmidt Paul Goldschmidt Andrew McCutchen
5 Paul Goldschmidt Jose Bautista Andrew McCutchen Troy Tulowitzki
6 Andrew McCutchen Prince Fielder Prince Fielder Adrian Beltre
7 Troy Tulowitzki Andrew McCutchen Carlos Gonzalez Prince Fielder
8 Ryan Braun Edwin Encarnacion Troy Tulowitzki Yasiel Puig
9 Prince Fielder Jose Abreu Giancarlo Stanton Paul Goldschmidt
10 Jose Abreu Yasiel Puig Jose Bautista Edwin Encarnacion
11 Chris Davis Giancarlo Stanton Yasiel Puig Albert Pujols
12 Edwin Encarnacion Chris Davis Edwin Encarnacion Ryan Braun
13 Jose Bautista Troy Tulowitzki Ryan Braun Robinson Cano
14 Adrian Beltre Ryan Braun Chris Davis Adrian Gonzalez
15 Giancarlo Stanton Joey Votto Shin-Soo Choo Jacoby Ellsbury
16 Albert Pujols Shin-Soo Choo Jose Abreu Buster Posey
17 Jacoby Ellsbury Albert Pujols David Ortiz Jose Bautista
18 Wilin Rosario David Ortiz Adrian Gonzalez Joey Votto
19 David Ortiz Adrian Beltre Adrian Beltre Jose Abreu
20 Adam Jones Evan Longoria Albert Pujols Eric Hosmer
21 Joey Votto Bryce Harper Anthony Rizzo Billy Butler
22 Carlos Beltran Jacoby Ellsbury Robinson Cano David Ortiz
23 Shin-Soo Choo Anthony Rizzo Evan Longoria Carlos Beltran
24 Adrian Gonzalez Carlos Beltran Buster Posey Chris Davis
25 Robinson Cano David Wright David Wright Anthony Rizzo
26 Bryce Harper Matt Holliday Matt Holliday Giancarlo Stanton
27 Anthony Rizzo Adrian Gonzalez Billy Butler Shin-Soo Choo
28 Evan Longoria Robinson Cano Joe Mauer Adam Jones
29 Eric Hosmer Jason Heyward Freddie Freeman Jose Reyes
30 Michael Cuddyer Adam Jones Carlos Beltran Allen Craig
31 Carlos Gomez Billy Butler Bryce Harper Matt Holliday
32 David Wright Freddie Freeman Allen Craig Norichika Aoki
33 Matt Holliday Carlos Gomez Eric Hosmer Pablo Sandoval
34 Billy Butler Eric Hosmer Pablo Sandoval David Wright
35 Buster Posey Justin Upton Michael Cuddyer Dustin Pedroia
36 Alex Rios Wilin Rosario Jacoby Ellsbury Michael Cuddyer
37 Matt Kemp Buster Posey Alex Gordon Wilin Rosario
38 Hanley Ramirez Matt Kemp Jason Heyward Joe Mauer
39 Freddie Freeman Michael Cuddyer Carlos Santana Martin Prado
40 Jose Reyes Jay Bruce Justin Upton Bryce Harper

(Note: I evaluated points leagues the same way as the other leagues, generating both a points total and a points/PA score for each player. I scaled the two values to give them approximately equal weight, and ranked players by the mean of the two.)

I expected Joey Votto to be a stud in OBP leagues, but in reality Joey Bats benefits more. Jason Heyward too. Meanwhile, CarGo is top 3 in every other system, but falls to the bottom half of the first round in a points league. In my own crazy league, Norichika Aoki projects as a contact-hitting top-40 stud, while Mark Trumbo’s contact deficiencies show up in strikeouts and hits, as well as AVG, and he drops to 82nd.

I also thought it would be cool to see which players project to be affected most under different scoring systems. Here are the players with the largest variation in ranks between systems (weighted to prefer higher-ranked and therefore more interesting players):

Player Trad 5 OBP 5 Points
Billy Hamilton 42 45 166
Joey Votto 21 15 3
Carlos Santana 101 46 39
Carlos Gonzalez 3 3 7
Carlos Gomez 31 33 69
Yasiel Puig 4 10 11
Alex Rios 36 60 90
Jose Bautista 13 5 10
Adam Jones 20 30 46
Rajai Davis 102 115 208
Joe Mauer 67 57 28
Wilin Rosario 18 36 43
Leonys Martin 58 72 121
Jacoby Ellsbury 17 22 36
Ben Zobrist 93 68 45
Starling Marte 45 67 92
Troy Tulowitzki 7 13 8
Matt Carpenter 125 119 62
Jose Abreu 10 9 16
Martin Prado 88 105 53
Josh Willingham 121 71 73
Jean Segura 51 81 96
Jonathan Villar 139 132 220
Pablo Sandoval 52 63 34
Miguel Montero 197 155 110
Ryan Braun 8 14 13
Allen Craig 41 55 32
Yoenis Cespedes 46 47 72
Giancarlo Stanton 15 11 9
Mike Napoli 99 58 89
Mark Teixeira 71 42 59
Drew Stubbs 135 126 197
George Springer 206 184 293
Jason Heyward 48 29 38
Prince Fielder 9 6 6
Shin-Soo Choo 23 16 15
Nick Swisher 107 79 68
Adam Dunn 239 151 230
Coco Crisp 56 51 78
Alfonso Soriano 90 93 133

Billy Hamilton projects to be a one-category stud in any system that ranks stolen bases, but many people doubt whether he’ll be an especially good ballplayer in 2014, and the points system shares their skepticism. Carlos Santana will benefit enormously from any league using deeper measures than AVG, while Adam Dunn jumps from irrelevance to potential rosterability in OBP leagues only. A couple more notable players: Alex Rios is vastly more valuable in leagues with the standard five categories, and least valuable in points league, and Adam Jones follows a very similar, if somewhat less drastic, pattern.

And there you have it – the results of one approach to generating player values for leagues with alternative categories.


When is 27 Old?

What do Andrew McCutchen, Buster Posey, Jay Bruce, and 12 other players have in common? They will all be in their age 27 season for 2014 and so we should expect that as a group their wOBAs will decline by 3 points on average. That may not sound like a lot, but it is the start of what will likely be the slow, steady offensive decline phase of their careers. Some will defy those odds, but which ones, and what might be signs of imminent decline?

To begin to answer these questions, I examined how hitters’ aging trends would be affected if certain skills did not decline with age. For example, how would a hitter age if his BABIP was consistently league average? I used wOBA as the measure of performance. The two age profiles in Figure 1 show this story. The solid line shows how players typically age, peaking at age 26, and declining steadily from there (see below for technical details on how this figure was constructed). The dashed line draws out a hypothetical age curve that assumes players’ BABIP is constant over time. Because players under age 30 tend to have above average BABIP, the dashed line is below the solid line for these young players. Conversely, older players would benefit from having their actual BABIP replaced by an average BABIP. The overall consequence of adjusting for BABIP is an age profile that is flatter—meaning that the effect of aging on wOBA is reduced.

Figure 1. Change in wOBA Age Profile when Adjusting for BABIP

Figure 1. Change in wOBA Age Profile when Adjusting for BABIP

The effect is reduced, but not eliminated. What other skills decline with age and how important are they to the decline in wOBA with age? Although swinging strike percentage, K rate, BB/K, and fly ball percentage all have important relationships with wOBA, these factors have little impact on the aging of wOBA. The wOBA age profile adjusted for these factors in Figure 2 is just slightly flatter than the unadjusted profile. This is because swinging strike percentage and K rate typically peak before age 26 and show little decline with age. To the extent that the adjusted curve in Figure 2 is flatter, trends with age in BB/K are most responsible.

Figure 2. Change in wOBA Age Profile when Adjusting for Swinging Strike Percentage, K%, BB/K, and Fly Ball Percentage

Figure 2. Change in wOBA Age Profile when Adjusting for Swinging Strike Percentage, K%, BB/K, and Fly Ball Percentage

Figure 1 demonstrated that BABIP plays an important role in the aging of wOBA. Adding both BABIP and HR/FB skills to the others from Figure 2 explains the entire decline in wOBA after age 26. Indeed, if a player’s BABIP and HR/FB skills (along with the others I’ve mentioned) remained average throughout his career, he would actually show continuous improvement through at least age 30. Figure 3 shows this result. The flatness of the adjusted line indicates that the full set of statistics used in the adjustment does a very good job of accounting for trends in wOBA with age.

Figure 3. Change in wOBA Age Profile when Adjusting for All of the Above and HR/FB

Figure 3. Change in wOBA Age Profile when Adjusting for All of the Above and HR/FB

So what does this mean for McCutchen and the others in our list? We may learn the most about how they will age based on observing trends in their BABIP and HR/FB rates from here on out. Doing so will be challenging because these are also among the least reliable measures of performance in a single season. Even so, a decline in these skills could indicate substantial performance losses to come. Additionally, players whose value derives from high walk or contact rates may age less precipitously than others.

A productive avenue for future analysis might be to assess whether there is a relationship between the amount of improvement in BABIP or HR/FB skills a player experiences before age 26 and the amount of decline in those skills after age 26. If so, then we might be able to better predict how a player’s offensive skills will age. However, we can learn a lot about averages, and our long-run projection for any particular player’s performance might improve, but it will always remain uncertain.

Technical Details

The age profiles adjust for what is often referred to as survivor bias—the fact that not all of the players in the sample at age 20 are also in the sample at age 35. To do this I used a technique commonly used by economists and others called fixed effects regression (see Jonah Rockoff’s work on changes in teachers’ performance with experience for one example). I run a regression that includes individual player fixed effects, ensuring that the relationship between wOBA and age is calculated using within-player variation in wOBA over time, rather than variation in performance across players. Consequently, the results are not affected by changes in the composition of MLB players by age. To calculate the adjusted profiles, I account for the other statistics as additional controls. Doing so means that when I adjust for BABIP, I am also adjusting for skills that are related to BABIP. Therefore my results do not depend on any particular model of how BABIP is related to performance. The above results are based on the 1,346 players who played in the majors for at least two seasons between 2002 and 2013. An alternative sample that is restricted to just the 304 players who played at least eight seasons during this period produces similar results, although the average wOBA levels are higher and the curves are less precisely estimated.

 

About the Author: Elias Walsh spends too much of his free time working with baseball data and trying to win his fantasy baseball league (or so his lovely wife informs him). His day job is a research economist at Mathematica Policy Research, where he conducts research to inform education policy decisions.


Ottoneu Tools: FGPoints

Below are two tools for Ottoneu FGPoints players to be used for the 2014 MLB season.  The first tool is a roster building tool that will provide 2013 statistics, including platoon splits, for offensive players.  Ottoneu players can use this tool to construct their team and prepare for 2014 auction drafts.

The second tool is a 2014 player projection tool that Ottoneu players (and commissioners) can use to estimate player projections for the 2014 season.  The tool incorporates Steamer, Oliver, and 3 Year Average stats for each player and then allows you to enter your own projections for the 2014 season.  Your own projections (will auto-populate FanGraph’s “Fans” projections as of 2.8.14…you can override these projections by entering your own) will load the team dashboard at the top of the tool and provide you with a summary of what you can expect from your Ottoneu team in 2014.

Roster Breakdown w/platoon splits:

http://bit.ly/1iCKkvl

2014 Team Projections Tool:

http://bit.ly/1eh614z


The Curious Case of Jason Castro

As we look for candidates to regress in 2014, a popular choice is Houston catcher Jason Castro for it seems the Astros backstop has two targets on his back: a high strikeout rate last year of 26.5% and a high BABIP of .351. Steamer and Oliver both project a steep drop in BABIP that will drag his batting average from a solid .276 to the .250s. As Brett Talley wrote, Castro screams regression.

Or does he?

Talley points to Castro’s strikeout rate that has been topped only 61 times in the past decade, and only four times the player matched or bettered a batting average of .276. But that measure may miss the mark. No one is suggesting Castro’s strikeout rate will worsen. When it comes to batting average, the critical question, then, is whether he can come close to maintaining a high BABIP.

On that question the evidence is more promising. In the last decade, only 38 of 1,509 batters have had an infield-fly rate lower than Castro’s 1.8%. Only 47 had a line-drive rate higher than Castro’s 25.2%. Taken together, those two select groups actually have 10 matches — players who managed both a lower infield-fly rate and higher line-drive rate. Here they are along with their BABIP, batting average and strikeout rate:

Player, year, BABIP, Avg., K-rate

Joe Mauer, 2013, .383, .324, 17.5%

Joey Votto, 2011, .349, .309, 12.9%

Howie Kendrick, 2011, .349, .297, 17.3%

Matt Carpenter, 2013, .359, .318, 13.7%

Michael Young, 2007, .366, .315, 15.5%

Joey Votto, 2013, .360, .305, 19%

Adam Kennedy, 2006, .313, .273, 14.3%

Bobby Abreu, 2006, .366, .297, 20.1%

Michael Young, 2011, .367, .338, 11.3%

Chris Johnson, 2012, .354, .281, 25%

 

What might we gather from this evidence?

(1) All but one of the players topped .276.

(2) The skills involved seem somewhat repeatable: Votto and Young each appear twice and as a group they generally in their careers combined a high LD rate, low IFFB rate and a high BABIP.

(3) We wouldn’t expect a player who whiffs a quarter of the time to have a batting average as high as someone who strikes out half as much while putting up similar LD and IFFB rates. Castro is unlikely to approach the median average of this group of .307.

(4) Castro doesn’t need to approach the median average to avoid significant regression. He is more likely to hit closer to last year’s mark than he is to hit in the .250s.


The R.A. Dickey Effect – 2013 Edition

It is widely talked about by announcers and baseball fans alike, that knuckleball pitchers can throw hitters off their game and leave them in funks for days. Some managers even sit certain players to avoid this effect. I decided to analyze to determine if there really is an effect and what its value is. R.A. Dickey is the main knuckleballer in the game today, and he is a special breed with the extra velocity he has.

Most people that try to analyze this Dickey effect tend to group all the pitchers that follow in to one grouping with one ERA and compare to the total ERA of the bullpen or rotation. This is a simplistic and non-descriptive way of analyzing the effect and does not look at the how often the pitchers are pitching not after Dickey.

Dickey's Dancing Knuckleball
Dickey’s Dancing Knuckleball (@DShep25)

I decided to determine if there truly is an effect on pitchers’ statistics (ERA, WHIP, K%, BB%, HR%, and FIP) who follow Dickey in relief and the starters of the next game against the same team. I went through every game that Dickey has pitched and recorded the stats (IP, TBF, H, ER, BB, K) of each reliever individually and the stats of the next starting pitcher, if the next game was against the same team. I did this for each season. I then took the pitchers’ stats for the whole year and subtracted their stats from their following Dickey stats to have their stats when they did not follow Dickey. I summed the stats for following Dickey and weighted each pitcher based on the batters he faced over the total batters faced after Dickey. I then calculated the rate stats from the total. This weight was then applied to the not after Dickey stats. So for example if Janssen faced 19.11% of batters after Dickey, it was adjusted so that he also faced 19.11% of the batters not after Dickey. This gives an effective way of comparing the statistics and an accurate relationship can be determined. The not after Dickey stats were then summed and the rate stats were calculated as well. The two rate stats after Dickey and not after Dickey were compared using this formula (afterDickeySTAT-notafterDickeySTAT)/notafterDickeySTAT. This tells me how much better or worse relievers or starters did when following Dickey in the form of a percentage.

I then added the stats after Dickey for starters and relievers from all four years and the stats not after Dickey and I applied the same technique of weighting the sample so that if Niese’12 faced 10.9% of all starter batters faced following a Dickey start against the same team, it was adjusted so that he faced 10.9% of the batters faced by starters not after Dickey (only the starters that pitched after Dickey that season). The same technique was used from the year to year technique and a total % for each stat was calculated.

The most important stat to look at is FIP. This gives a more accurate value of the effect. Also make note of the BABIP and ERA, and you can decide for yourself if the BABIP is just luck, or actually better/worse contact. Normally I would regress the results based on BABIP and HR/FB, but FIP does not include BABIP and I do not have the fly ball numbers.

The size of the sample was also included, aD means after Dickey and naD is not after Dickey. Here are the results for starters following Dickey against the same team.

Dickey Starters

It can be concluded that starters after Dickey see an improvement across the board. Like I said, it is probably better to use FIP rather than ERA. Starters see an approximate 18.9% decrease in their FIP when they follow Dickey over the past 4 years. So assuming 130 IP are pitched after Dickey by a league average set of pitchers (~4.00 FIP), this would decrease their FIP to around 3.25. 130 IP was selected assuming ⅔ of starter innings (200) against the same team. Over 130 IP this would be a 10.8 run difference or around 1.1 WAR! This is amazingly significant and appears to be coming mainly from a reduction in HR%. If we regress the HR% down to -10% (seems more than fair), this would reduce the FIP reduction down to around 7%. A 7% reduction would reduce a 4.00 FIP down to 3.72, and save 4.0 runs or 0.4 WAR.

Here are the numbers for relievers following Dickey in the same game.

Dickey Bullpen

Relievers see a more consistent improvement in the FIP components (K, BB, HR) between each other (11.4, 8.1, 4.9). FIP was reduced 10.3%. Assuming 65 IP (in between 2012 and 2013) innings after Dickey of an average bullpen (or slightly above average, since Dickey will likely have setup men and closers after him) with a 3.75 FIP, FIP would get reduced to 3.36 and save 3 runs or 0.3 WAR.

Combining the un-regressed results, by having pitchers pitch after him, Dickey would contribute around 1.4 WAR over a full season. If you assume the effect is just 10% reduction in FIP for both groups, this number comes down to around 0.9 WAR, which is not crazy to think at all based off the results. I can say with great confidence, that if Dickey pitches over 200 innings again next year, he will contribute above 1.0 WAR just from baffling hitters for the next guys. If we take the un-regressed 1.4 WAR and add it to his 2013 WAR (2.0) we get 3.4 WAR, if we add in his defence (7 DRS), we get 4.1 WAR. Even though we all were disappointed with Dickey’s season, with the effect he provides and his defence, he is still all-star calibre.

Just for fun, lets apply this to his 2012. He had 4.5 WAR in 2012, add on the 1.4 and his 6 DRS we get 6.5 WAR, wow! Using his RA9 WAR (6.2) instead (commonly used for knucklers instead of fWAR) we get 7.6 WAR! That’s Miguel Cabrera value! We can’t include his DRS when using RA9 WAR though, as it should already be incorporated.

This effect may even be applied further, relievers may (and likely do) get a boost the following day as well as starters. Assuming it is the same boost, that’s around another 2.5 runs or 0.25 WAR. Maybe the second day after Dickey also sees a boost? (A lot smaller sample size since Dickey would have to pitch first game of series). We could assume the effect is cut in half the next day, and that’d still be another 2 runs (90 IP of starters and relievers). So under these assumptions, Dickey could effectively have a 1.8 WAR after effect over a full season! This WAR is not easy to place, however, and cannot just be added onto the teams WAR, it is hidden among all the other pitchers’ WARs (just like catcher framing).

You may be disappointed with Dickey’s 2013, but he is still well worth his money. He is projected for 2.8 WAR next year by Steamer, and adding on the 1.4 WAR Dickey Effect and his defence, he could be projected to really have a true underlying value of almost 5 WAR. That is well worth the $12.5M he will earn in 2014.

For more of my articles, head over to Breaking Blue where we give a sabermetric view on the Blue Jays, and MLB. Follow on twitter @BreakingBlueMLB and follow me directly @CCBreakingBlue.


Evaluating 2012 Projections

Evaluating 2012 Projections

Hello loyal readers.  It’s time for the annual evaluation of last year’s player projections.  Last year saw Gore, Snapp, and Highly’s Aggpro forecasts win among hitter projections (http://www.fangraphs.com/community/comparing-2011-hitter-forecasts/) and Baseball Dope win among pitchers http://www.fangraphs.com/community/comparing-2011-pitcher-forecasts/ .  In general, projections computed using averages or weighted averages tended to perform best among hitters, while for pitchers, structural models computed using “deep” statistics (k/9, hr/fb%, etc.) did better.

2012 Summary

In 2012, there were 12 projections submitted for hitters and 12 for pitchers (11 submitted projections for both).  The evaluation only considers players where every projection system has a projection.

Read the rest of this entry »


Introducing BERA: Another ERA Estimator to Confuse You All

Coming up with BERA… like its [almost] namesake might say, it was 90% mental, and the other half was physical.  OK, maybe he’d say something more along the lines of “what the hell is this…” but that’s beside the point.    By BERA, I mean BABIP-estimating ERA (or something like that… maybe one of you can come up with something fancier).  It’s an ERA estimator that’s along the lines of SIERA, only it’s simpler, and—dare I say—better.

You know, I started out not knowing where I was going, so I was worried I might not get there.  As you may recall, I’ve been pondering pitcher BABIPs for a little while here (see article 1 and article 2), and whereas my focus thus far had been on explaining big-picture, long-term BABIP stuff in terms of batted ball data, one question that remained was how well this info could be used to predict future BABIPs.  After monkeying around with answering that question, though, I saw that SIERA’s BABIP component could be improved upon, so I set to work in coming up with BERA.  In doing so, I definitely piggybacked off of FIP and a little of what SIERA had already done.  You can observe a lot just by watching, you know.   I’m also a believer in “less is more” (except for when it comes to the size of my articles, obviously), so I tried to go for the best compromise of simplicity and accuracy that I could.

Read the rest of this entry »