Archive for Fantasy

2015 Fantasy Sleepers: Starting Pitching

The key to winning at fantasy baseball is finding players who will outperform their draft position.  This will be the first of a series of articles addressing undervalued and overvalued players that you should be targeting in your draft.

Read the rest of this entry »


Do Rookie Hitters Decline in the Second Half?

Do rookies perform worse after the All-Star break?

My claim over this statement is nonexistent, while the original thought of its occurrence was brought to my attention by Adam Aizer on the CBS Fantasy Baseball Podcast.

My judgment dissuaded, I thought that it would be worth the effort to look into the validity of the statement.

From the perspective of an offensive player, rookies infrequently make enough of an impact in the size of leagues (i.e. 10-team and 12-team leagues) that pedestrian Fantasy Baseball players occupy. For those sizes of leagues that the aforementioned owners participate in, a rookie hitter that is worth owning is either an elite prospect or a player that has preformed beyond their true talent level. As a result, the former is rare, while it would make sense for the latter to regress to their true talent level and is more common than the former. The idea that rookie hitters decline throughout the year is just a misevaluation of the player’s true talent level.

To put another way, it is the same logic that comes into play with a recent event: the Home Run Derby. Players that participate in the Home Run Derby are players that have exceptional first halves, which are often beyond their true talent level. These players often perform worse in the second half than they did in the first half, not because they participated in the monotonous and dated event that has become the Home Run Derby, but because, just like the rookies who perform worse in the second half of the season than the first, they have regressed toward their true talent level; when the rookies regress, they have just regressed to the point where they are not ownable.

The research looks at all player seasons between 1988 and 2013 where a batter was in their first season, had 250 plate appearances in the first half of the season, and had 250 plate appearances in the second half of the season.

Screen Shot 2014-07-20 at 8.48.48 PM

The rookie second half decline and the post Home Run Derby slump intuitively make sense, but intuition does not always bear truth. Through cognitive ease we rationalize that “Swinging that hard for that long throws off your timing”; “A rookie is too young to be able to make it through the long hot summer.”

Because most fantasy leagues are small, the only reason that the common rookie was on our teams to begin with is because they had to play beyond their ability in the first half of the season. The rookie who is on our team right now, unless he is a reputable prospect, is probably a safe bet to decline. But as a whole, we can see that there is no decline in rookie performance based on first half/second half splits.

Our desire to perceive a decline is just our desire to hold onto our ability as talent evaluators. We know that Yangervis Solarte is a great player, and the only reason he hasn’t been able to sustain his performance is because he is rookie that can’t play out the season: common baseball logic. In actuality, Solarte was not as good as some originally thought, and his true talent was never good enough to be on a 10 or 12 team league.

Summary:

Rookie hitters, as a generalization, are not good enough to play in 10 or 12 team leagues, and, as a generalization, those that do play in ten team leagues regress to their true talent level, which is not valuable enough to be ownable.

Devin Jordan is obsessed with statistical analysis, non-fiction literature, and electronic music. If you enjoyed reading him, follow him on Twitter @devinjjordan.


Ottoneu Tools: Advanced League Standings

Ottoneu Tools: Advanced Standings (Part One)

Whether you’re brand new to Ottoneu or a “seasoned” veteran in your fourth year, your league’s Standings page is likely to become your best (or worst) friend over the course of a given baseball season. However, if you often find yourself cheering or panicking based on just a few days’ worth of small but evolving linear weights data without the proper, broader context with which to make meaningful decisions about your team, you are not alone. Welcome to the Ottoneu “Advanced Standings” dashboard. The brainchild and early creation of Bill Porter (@wfporter1972), the Advanced Standings dashboard will provide you with the sabermetric performance data you want with the detail you need.

How It Works:

Before you get started you will need to download the current version of the Advanced Standings dashboard here (http://goo.gl/Tozhy4). Note: You will need the most up to date version of Excel to take full advantage of the dashboard features. Also, this Advanced Standings dashboard only works with FGPoints Ottoneu format leagues (for now).

While it may look overwhelming at first, the dashboard is designed to be easy to use. In fact, it’s designed to be updated quickly and often without requiring a lot of Excelmanship. With as little as two easy steps you will be able to see “inside” your league standings in a way not available on the website.

First, go to your league’s traditional STANDINGS page within Ottoneu. From the bottom right of the standings stats (begin just to the right of the last P/IP on the right hand bottom corner), highlight all standings data with your cursor (including team names). Do not export the standings to Excel. Also, do not highlight the headings bar that includes the column titles (AB, H, 2B, etc.). COPY this information and then go to the first tab (“Advanced Standings”) of the Excel dashboard. In cell A4 (1st team name in column), PASTE SPECIAL and select TEXT. When pasted, your league’s standings will populate this tab and you will have visibility of many advanced statistics tailored exactly to your league.

Second, to have more accurate league standings information, go to your league’s REPORT page within Ottoneu and at the bottom of the page highlight all the information (excluding the column headings) in the “Projected Games Played and Innings Pitched” section. Once selected, COPY this information (including team names), and PASTE SPECIAL – TEXT this data into Tab 2 (“Reports”) of the dashboard spreadsheet, in cell A2.

Yeah, it’s that easy.

In Part Two I will revisit some of the key features of the Advanced Standings Dashboard and how it can be best used to analyze your league  You can also learn how to go much deeper into these advanced standings from reading Bill’s recent post on this subject here (http://goo.gl/XkDzXV). Until then, enjoy playing with the dashboard tool. If you have questions or want access to some additional Ottoneu tools, feel free to DM me on Twitter @Fazeorange and I will send you a link to the Ottoneu Dropbox folder.

Enjoy


Ranking Batters in Fantasy Leagues with Alternate Stats

Draft prep: Framing the problem

So you’re preparing for your fantasy draft. You’re caught up on FanGraphs, checked for recent injuries at Rotoworld, maybe skimmed a few headlines from your other top 11 baseball news sites. Maybe you’ve even downloaded the FanGraphs positional rankings, and are planning to keep the file open during the draft as a reality check against the pre-set rankings of the site your league uses.

But really, what do the guys at FanGraphs know? Sure, they know a lot about baseball, and statistics, and this year’s projections, and a handful of underlying stats that tend to predict future performance. But what they don’t know is whether your league uses OBP instead of AVG, or OPS, SLG, or batters’ strikeouts, or maybe holds and FIP and pitcher fielding percentage. If this is your situation, then I feel your pain. My fantasy league uses eight statistics for batters and pitchers, three each beyond the usual five. (In case you’re curious, the mysterious six are: Batter hits, K’s, & OPS; Pitcher holds, losses & complete games).

These differences matter. If your league uses OBP, Joey Votto turns from a fantasy player who’s solid in four categories (including average, where his impact is limited because he walks all the time) to a guy with a truly elite skill. Maybe it’s easy for you to account for the relative value of a Joey Votto, but how well can you project the 25th through 35th outfielders? Some might be much better or worse in your league. If you have batter strikeouts, as in my league, how do you value Mark Trumbo and his home run power against the elite contact skills of Norichika Aoki?

Generating your own rankings

One answer, and the one I opted for, is to generate rankings based on your own league’s stats. Now, this may sound a bit too work-intensive and time-consuming for most of you (especially those of you with relatively normal priorities), but in reality it wasn’t as time-consuming as I expected.*

First of all, there’s no need to reinvent the wheel. There are lots of projection systems out there that are available to the public, and some of them are quite good. I decided I would simply download all the projections listed on FanGraphs, and average them out. And then, after thinking for a little while about the costs and benefits of that approach, I decided I wouldn’t do that at all, and instead would use the results of just one projection system. But which one should I use? Luckily, that’s yet another bit of analysis we don’t need to bother with, because the Interwebs are full of crazy mathematicians who love baseball and have nothing better to do. After searching for a few articles that evaluate projection systems, like this one and this meta-one, I decided that the forecasts I trusted most (and were easiest to obtain) were Steamer for batters and FanGraphs fans for pitchers. (The high accuracy of the latter shocked me at first, but then I realized that fans assimilate the results of all the projection systems into their own player projections, departing from them only as dictated by common sense, inside scoop, and hope.)

Operationalizing the Solution

Here’s where it gets tricky. What advanced data manipulation packages and techniques are best for downloading reams of data from the FanGraphs site into your spreadsheet? Certainly there was no need for me to copy and paste the data 50 players at a time like someone living the dark ages, was there? No, of course not. And I probably never really did that.

Instead – bear with me if you’re not technically inclined – I hit the gray “Export Data” button to the upper right of my chosen projection page. This involved a lot of loading the correct page, hovering my mouse over the text, and clicking, but in the end it was worth all the work, because 5 minutes of sweat, plus a beer, had finally paid off in spreadsheets full of data.

*If you’re not interested in these details, the fun stuff is posted in a couple of tables towards the end. (I like writing, so this is likely to go on for a while.)

Z-scoring your data points

Z-scoring batter projections is easy. The problem lies in determining what set of players to use in order to calculate means and standard deviations.

This is an important question, at least to the extent that any question in fantasy baseball is important. For example, if you must use every hitter in the league, including the guys projected for 8 at-bats, you create the illusion that lots of players bat .220 or score only 4 runs, as opposed to your league’s reality in which .270 with 70 runs is pretty ordinary. For a little math fun, I compared the results generated using means and deviations 500 players deep (the equivalent of a 25-team league that rosters 20 position players) versus one with more reasonable assumptions. It caused huge increases in variance in runs and rbi’s, so a guy who drove in and scored 100 compared no better to the mean either way (~2+ standard deviations), but smaller increases in the variance in SB’s, HR’s, and OPS, which, together with the lower means, meanings this system overvalues guys who produce in these categories. Martin Prado and Torii Hunter were made sad, whereas Billy Hamilton was elevated to a demigod (or at least a top-40 hitter).

So how do you generate values that represent your player pool?

One method – and a very reasonable one – is to use the final statistics compiled by your league the previous year. With this data, it’s easy to generate per-slot averages based on last year’s performance, and to compare projected performance against it. But I did not choose this method. A more savvy number-cruncher might say that projection systems, while designed to be as accurate as possible for each player, may be systematically biased on the whole, and therefore determining the value of this year’s projections based on last year’s actual statistics is tantamount to comparing apples and oranges.

I was more worried about lazy owners. Any league can have a couple of careless owners who are in it just for fun (the gall!), or who keep BJ Upton when he can’t even see the Mendoza line, because of that one time his cousin shook BJ’s hand at a Jay-Z concert. I know of what I speak. If your goal is to win your league, you want to base your evaluation on the best players available, rather than the happenstance of which Atlanta outfielders spent the whole year on someone’s roster.

I generated means using very precise data, plus a random stab in the dark. First, I looked up the exact number of players at each position in my league from the previous year. Then I mostly ignored this data. Although it’s true that player values vary greatly between leagues depending on how many players start, and how many are rostered, this is the sort of thing you can keep track of during the draft. Don’t draft another first baseman if you already have three of them and no shortstop, and don’t draft a first baseman just because he’s ranked ahead of a shortstop if there are another seven first basemen ranked close behind.

My league rostered only 123 regulars last year. Not a deep league. I used a lot more than 123 in my calculations in an effort to lower the means a bit, to account for the existence of catchers and second basemen. I then haphazardly created sort variables so I could bring the best 150 to 180 players to the fore, with the goal of getting a fair representation of the quality of players in my league. I tried various formulas like [(HR+1) * R * RBI * (SB +1) * AVG * OPS] (adding 1’s so as not to exclude players projected for 0 HR’s or SB’s ) and PA * wOBA. Virtually every one of them produced a good representation of the best hitters projected for regular playing time. In the end, the best way to evaluate the sort is to look at the list and see if the guys near the cutoff are fringe players who are familiar from last year’s waiver wire.

Calculating projected player values

Once you determine which players you want to include, Excel is happy to instantaneously calculate averages and standard deviations for each stat. Once you have these values, you can re-include the entire player pool, or as much of it as you wish, and the formula for each player in each category is simply (his projected value – the average projected value)/standard deviation.

The next challenge is to generate ranks from the Z-scores. The simplest way is simply to add them together (being sure to subtract ones where lower scores are better, such as pitcher walks or batter strikeouts). But here, I discovered another issue. A potential superstar who might not have a full-time job could end up ranked about the same or below a mediocre player who was guaranteed to start. If I wanted my draft rankings to make sense at a glance when I have just 90 seconds to pick a player while eating a sandwich, I needed to distinguish accumulators from guys with potential.

Ranking performance and potential

It matters whether a player is an okay guaranteed performer or a unpredictable potential star. If I find myself with no second basemen in the 22nd round, I might want to take the best guy who’s pretty much guaranteed 140 days in the starting lineup, like an Anthony Rendon or a Howie Kendrick. If my roster’s pretty much set, I might prefer a hitter who has a better chance to bust out and hit 45 home runs, like Chris Carter (unless I’m in my league, in which his 80% strikeout rate falls 37 standard deviations below the mean).

What I decided to do was generate two rankings for each batter, one based on projected totals, and one based on projections per plate appearance. Luckily, Steamer has already done the work for us by projecting everyone in both ways. For instance, Everth Cabrera is projected as the 479th-best player by wOBA, with 74 runs and 45 stolen bases. At the other extreme, Colorado’s Kris Parker is projected to be the 50th-best hitter in the league, just ahead of Dustin Pedroia, with a .279 batting average and .465 slugging percentage, despite getting only one plate appearance, and not getting a hit.

At this point, there are 2 sets of columns for each batter: 1 set of columns for his Steamer projections for each relevant stat, and 1 for the associated Z-scores. To this, I added 2 more sets of columns: 1 for per plate-appearance projections for each stat, and 1 for those associated Z-scores. (Dividing hits into plate appearances rather than at-bats feels unnatural, but that’s what you need to do if your league counts total hits.) Calculating per-PA quality is then easy, as you can just add the Z-scores (or subtract for negative statistics). But once you have projected rate statistics in your per-PA rankings, it becomes apparent that it doesn’t make sense to include the exact same values in your projected accumulated totals.

To handle this, I weighted the Z-scores for the rate stats. I multiplied the Z-score for AVG by projected AB’s/average projected AB’s, and you can do the same for OBP, using PA’s. My league uses OPS, a value generated by adding two fractions with different denominators (aka OBP & SLG), so to weight those Z-scores I multiplied them by projected (AB’s + PA’s)/average projected (AB’s + PA’s). I then added these weighted Z-scores to the other Z-scores for projected totals. The result of adding these weights is that a player who is one standard deviation above average in both AVG and OPS, and who has an average number of AB’s and PA’s, would get +2 from these categories in the variable used to rank projected totals. By the same lights, the aforementioned Kyle Parker’s AVG and OPS would essentially get no weighting at all, and have no effect at all on his projected totals, just as in real life his performance is not expected to have any effect at all on the rate stats of your team.

The Fun Stuff

And that’s about it. Once you have Z-scores, it’s very easy to rank players, to change the formulas to rank them by different systems, or to sort players by certain categories to see who stands out the most.

Two common variations on the traditional 5 stats are to include OBP instead of AVG, or to play in a points league. (For a points league, just change the Z-score weighting to reflect the point system). Here are the top players in these alternate systems using this evaluation method (I threw my own league in too, just for kicks):

Rank Trad 5 OBP 5 Points Crazy 8s
1 Miguel Cabrera Miguel Cabrera Miguel Cabrera Miguel Cabrera
2 Mike Trout Mike Trout Mike Trout Mike Trout
3 Carlos Gonzalez Carlos Gonzalez Joey Votto Carlos Gonzalez
4 Yasiel Puig Paul Goldschmidt Paul Goldschmidt Andrew McCutchen
5 Paul Goldschmidt Jose Bautista Andrew McCutchen Troy Tulowitzki
6 Andrew McCutchen Prince Fielder Prince Fielder Adrian Beltre
7 Troy Tulowitzki Andrew McCutchen Carlos Gonzalez Prince Fielder
8 Ryan Braun Edwin Encarnacion Troy Tulowitzki Yasiel Puig
9 Prince Fielder Jose Abreu Giancarlo Stanton Paul Goldschmidt
10 Jose Abreu Yasiel Puig Jose Bautista Edwin Encarnacion
11 Chris Davis Giancarlo Stanton Yasiel Puig Albert Pujols
12 Edwin Encarnacion Chris Davis Edwin Encarnacion Ryan Braun
13 Jose Bautista Troy Tulowitzki Ryan Braun Robinson Cano
14 Adrian Beltre Ryan Braun Chris Davis Adrian Gonzalez
15 Giancarlo Stanton Joey Votto Shin-Soo Choo Jacoby Ellsbury
16 Albert Pujols Shin-Soo Choo Jose Abreu Buster Posey
17 Jacoby Ellsbury Albert Pujols David Ortiz Jose Bautista
18 Wilin Rosario David Ortiz Adrian Gonzalez Joey Votto
19 David Ortiz Adrian Beltre Adrian Beltre Jose Abreu
20 Adam Jones Evan Longoria Albert Pujols Eric Hosmer
21 Joey Votto Bryce Harper Anthony Rizzo Billy Butler
22 Carlos Beltran Jacoby Ellsbury Robinson Cano David Ortiz
23 Shin-Soo Choo Anthony Rizzo Evan Longoria Carlos Beltran
24 Adrian Gonzalez Carlos Beltran Buster Posey Chris Davis
25 Robinson Cano David Wright David Wright Anthony Rizzo
26 Bryce Harper Matt Holliday Matt Holliday Giancarlo Stanton
27 Anthony Rizzo Adrian Gonzalez Billy Butler Shin-Soo Choo
28 Evan Longoria Robinson Cano Joe Mauer Adam Jones
29 Eric Hosmer Jason Heyward Freddie Freeman Jose Reyes
30 Michael Cuddyer Adam Jones Carlos Beltran Allen Craig
31 Carlos Gomez Billy Butler Bryce Harper Matt Holliday
32 David Wright Freddie Freeman Allen Craig Norichika Aoki
33 Matt Holliday Carlos Gomez Eric Hosmer Pablo Sandoval
34 Billy Butler Eric Hosmer Pablo Sandoval David Wright
35 Buster Posey Justin Upton Michael Cuddyer Dustin Pedroia
36 Alex Rios Wilin Rosario Jacoby Ellsbury Michael Cuddyer
37 Matt Kemp Buster Posey Alex Gordon Wilin Rosario
38 Hanley Ramirez Matt Kemp Jason Heyward Joe Mauer
39 Freddie Freeman Michael Cuddyer Carlos Santana Martin Prado
40 Jose Reyes Jay Bruce Justin Upton Bryce Harper

(Note: I evaluated points leagues the same way as the other leagues, generating both a points total and a points/PA score for each player. I scaled the two values to give them approximately equal weight, and ranked players by the mean of the two.)

I expected Joey Votto to be a stud in OBP leagues, but in reality Joey Bats benefits more. Jason Heyward too. Meanwhile, CarGo is top 3 in every other system, but falls to the bottom half of the first round in a points league. In my own crazy league, Norichika Aoki projects as a contact-hitting top-40 stud, while Mark Trumbo’s contact deficiencies show up in strikeouts and hits, as well as AVG, and he drops to 82nd.

I also thought it would be cool to see which players project to be affected most under different scoring systems. Here are the players with the largest variation in ranks between systems (weighted to prefer higher-ranked and therefore more interesting players):

Player Trad 5 OBP 5 Points
Billy Hamilton 42 45 166
Joey Votto 21 15 3
Carlos Santana 101 46 39
Carlos Gonzalez 3 3 7
Carlos Gomez 31 33 69
Yasiel Puig 4 10 11
Alex Rios 36 60 90
Jose Bautista 13 5 10
Adam Jones 20 30 46
Rajai Davis 102 115 208
Joe Mauer 67 57 28
Wilin Rosario 18 36 43
Leonys Martin 58 72 121
Jacoby Ellsbury 17 22 36
Ben Zobrist 93 68 45
Starling Marte 45 67 92
Troy Tulowitzki 7 13 8
Matt Carpenter 125 119 62
Jose Abreu 10 9 16
Martin Prado 88 105 53
Josh Willingham 121 71 73
Jean Segura 51 81 96
Jonathan Villar 139 132 220
Pablo Sandoval 52 63 34
Miguel Montero 197 155 110
Ryan Braun 8 14 13
Allen Craig 41 55 32
Yoenis Cespedes 46 47 72
Giancarlo Stanton 15 11 9
Mike Napoli 99 58 89
Mark Teixeira 71 42 59
Drew Stubbs 135 126 197
George Springer 206 184 293
Jason Heyward 48 29 38
Prince Fielder 9 6 6
Shin-Soo Choo 23 16 15
Nick Swisher 107 79 68
Adam Dunn 239 151 230
Coco Crisp 56 51 78
Alfonso Soriano 90 93 133

Billy Hamilton projects to be a one-category stud in any system that ranks stolen bases, but many people doubt whether he’ll be an especially good ballplayer in 2014, and the points system shares their skepticism. Carlos Santana will benefit enormously from any league using deeper measures than AVG, while Adam Dunn jumps from irrelevance to potential rosterability in OBP leagues only. A couple more notable players: Alex Rios is vastly more valuable in leagues with the standard five categories, and least valuable in points league, and Adam Jones follows a very similar, if somewhat less drastic, pattern.

And there you have it – the results of one approach to generating player values for leagues with alternative categories.


Does Pitching Deep into Games Lead to More Wins?

Predicting pitcher wins is a capricious exercise, and few factors have been shown to have any correlation whatsoever with win percentage (W%). To predict wins, one should consider a pitcher’s ERA, offensive support, strength of bullpen, quality of defense behind the mound, and, innings pitched (IP) in a season.

In fact, research has shown that IP and ERA are the only two factors that have a correlation above .30, and the two are very close. In a sample of pitchers from 2003-2013, the correlation for both eclipsed .40.

Obviously, pitching more games leads to more wins in a season, but many fantasy experts insist that pitching deep into games is an important part of earning a win as well. The theory, which I’ve seen taken for granted by experts at ESPN, CBS, Baseball Prospectus, and Rotographs, is that a starting pitcher who pitches into the 8th or 9th inning and leaves with a lead intact is more likely be credited with the W.

However, to earn a win a starter must pitch only 5 innings. Since we know that starters are often less effective after 75 pitches or so, pulling a pitcher early and relying a fresh bullpen that is at least league average should, in theory, be more effective than keeping a starter in the game. Dave Cameron articulated this point when creating a gameplan for the Pirates’ all-important play-in game in October 2013 when he suggested Liriano be pulled after only 3 innings. The chart below reinforces the obvious point that, except for walk rate, relievers generally eclipse starters in most skill metrics.

Figure 1

In 2013 Shelby Miller started 31 games and came away with the W a total of 15 times, earning a W% that ranked 22nd in the majors right behind Clayton Kershaw and Anibal Sanchez. That’s impressive, but also consider that the innings-limited rookie pitched an average of 5.5 innings per start—he only racked up 13 quality starts (QS), ranked 86th in the league. QS, after all, require putting in 6 innings of work with at least a 4.50 ERA.

Why, then, are innings pitched per start (IP/GS) so important, relatively, when considering W%? I hypothesized that pitchers who are given the leeway to pitch deep into games, and hence give their bullpen a rest, were generally better at run prevention than their peers, i.e. sported a lower ERA.

In healthcare research, where we don’t write particularly well, we love simple diagrams to explain hypothesized effects. Below is a diagram showing how one might view the relationship between various factors like ERA, IP, defense, offensive support and bullpen ERA. The perceived link between IP/GS and Pitcher Wins is confounded by ERA, which has an effect on both factors.

DAG
Pitch Efficiency

Before examining the theory that ERA accounts for the correlation between IP and W%, lets look at another possible explanation. Perhaps pitch efficiency is the key. Jordan Zimmermann was the 3rd most efficient starter (14.5 P/IP) in the majors last year, and was tied for the 8th highest W% (.68). However, the table below shows the correlation between W per game started (GS) and P/IP, ERA, and IP/GS among starters between 2009-2013:

 

W% and…

R2

     ERA

0.39

     IP/GS

0.36

     P/IP

0.08

While ERA and IP/GS appear to be almost equally correlated, the squared correlation coefficient for P/IP was negligible at .08. Variance in pitch efficiency has little to do with variance in W%.

IP/GS: How to Measure a Confounder

There are 2 straightforward ways to determine if the relationship between 2 variables is actually being skewed by a third factor, in this case ERA. The first is to stratify the sample by ERA and see if the relationship between IP/GS and W% still stands. If ERA is not a confounder, we would expect the correlation between each tier to remain relatively stable. As we can see in the chart below, it follows no clear trend.

Figure 3

Interestingly, only the best tier of pitchers, those with an ERA less than 3.65, show any discernible relationship between W% and IP/GS, supporting the theory that those starters who have demonstrated a strong ability to prevent runs are given the chance to pitch more innings.  Among more middling pitchers, the relationship between pitching deep into games and W% is negligible.

The second way to measure confounding is using a regression model. If you create a model examining how factor X predicts factor Y, introducing factor Z should not change the coefficient for X by more than 10% if Z does not have a strong pull on the relationship. For example, if we run a model that shows that smoking doubles your chance of getting lung cancer, then introducing tea drinking into the equation should not really change that smoking-lung cancer connection by more than 10%, unless we believe that drinking tea can also affect lung cancer and/or smoking.

I’m with MGL that regression is often unnecessary in baseball research, as its results can be difficult to interpret and unnecessarily complicated. I might add that even simple linear regression rests on a series of assumptions that are not always met. With that caveat, the data in this sample are normally distributed and I kept the model as simple as possible. Model 1 examines the relationship between W% and IP/GS. Model 2 adds a third variable, ERA.

Parameter

Coefficient (%)

P-Value

Model 1

IP/S

11.13

<.01

Model 2

IP/S

5.71

<.01

Model 2

ERA

-4.77

<.01

All results are statistically significant. Model 1 indicates that for each 1-inning increase in IP/GS, we would expect an 11% increase in W%. Once we control for ERA, we see that each 1-inning increase would result in an even weaker relationship— we would expect a 6% increase in W%. The new coefficient, .057, is more than 10% different from .111 and we can safely conclude that ERA is confounding this relationship, just as we found in the stratified analysis above.

Predicting Wins?

Here at FanGraphs we might mock the idea of pitcher wins, since they are mostly a byproduct of an era when pitchers did pitch deep into games and bullpens were not utilized as often or as effectively. However, when it comes to predicting wins, Will Larson has shown that projection systems like Steamer and CAIRO do a pretty good job, and are on average within 3.5-4 wins of the actual end-of-season results.

In fact, projection systems across the board are better at capturing player-to-player variation (ranking players) in counting statistics like W and strikeouts than rate stats ERA and WHIP.

Figure 4

While I have previously shown that QS correlate much better than W with pretty much every measure of pitcher skill we have, W% is still somewhat predictable. As long as we have yet to #killthewin, we might as well keep trying to forecast the future.