Showing posts with label Robust. Show all posts
Showing posts with label Robust. Show all posts

Wednesday, May 2, 2012

Take-the-Best Statistical Model

Why do we do multiple regression?


Multiple regression is the workhorse of econometrics.  Almost every empirical paper in economics relies on it, and it also forms the basis for a large majority of political science research.  But is this a valid model for prediction?  How sensitive are the results?

As it turns out, the answer is "very".  Multiple regression is sensitive to a host of issues, including normality, linearity, and low error data.  But if the real world doesn't always fit these assumptions, why do we try to use the model to predict the real world?  One thing to notice when reading the empirical papers is that they often tell you that a certain coefficient is statistically significant, but rarely does one see a given confidence interval for that coefficient.  Of course, listing confidence intervals for every coefficient, especially when there are so many, is quite cumbersome.  Yet this convenient omission often leads us to be too confident about our estimates.  How much do we really know?

Behavioral economists have long criticized this model of human decision making because there's no feasible way that we can run a regression in our head and then make a decision.  At least, I know I don't.   Although many of my friends may have used excel spreadsheets to decide where to go to college, they did not end up basing their decision on some kind of complex regression model.  It's just too computationally intractable for everyday use.

Furthermore, one key problem of multiple regression is ecological validity.  We know that the regression model predicts the sample pretty well, but does it predict the future with any accuracy?  Especially if the future is highly variable and uncertain, why should we trust our Gaussian methods that are highly sensitive to outliers? According to Gerd Gigerenzer, the most accurate rules are often not the high powered intensive statistics methods.  Rather, fast and frugal algorithms that actually limit the information they evaluate can create more accurate results.

One of the prototypical fast and frugal algorithms Gigerenzer describes is Take the Best.  While multiple regression would look at all the data and perform various tests on individual data's contribution to the dependent variable, Take the Best does a sequential evaluation of a list of key determinants.  For example, multiple regression would decide between two restaurants by looking at all the data: food quality, wait time, location, parking spots.  It would then weight each input carefully according to an equation, and then look at the results of the equations for the two restaurants.  Whichever restaurant returns the higher value is the restaurant that's chosen.  On the other hand, Take the Best would look at whether the food quality gap is high.  If so, then pick the restaurant with better food.  If the gap is not high enough, move on to the next rule and repeat this simple process.  Computationally, this would require M+1 evaluations, which is in linear time and computationally quite tractable.

Gigerenzer applied this heuristic to predicting Chicago high school dropout rates.  Given two high schools and all the associated statistics: attendance rate, proportion of low income students, social science test scores, and more, what's the most accurate way to predict which high school had a higher dropout rate?   Gigerenzer and his fellow researchers took half the population of schools and built a Take the Best model and a multiple regression model to explain the data.  No surprise, multiple regression did better, predicting over 70% of the pairs correctly while Take the Best only managed around 65%.  Yet when the two models were tested on the other half of the population, Take the Best had about a 60% accuracy rate and multiple regression barely manged around the low 50%'s.

Surprising?  Gigerenzer in Gut Feelings explains:
But why did ignoring information pay in this case? High school dropout rates are highly unpredictable-in only 60 percent of the cases could the better strategy correctly predict which school had the higher rate, (Note that 50 percent would be chance.) Just as a financial adviser can produce a respectable explanation for yesterday's stock results, the complex strategy can weigh its many reasons so that the resulting equation fits well with what we already know. Yet, as Figure 5-2 clearly shows, in an uncertain world, a complex strategy can fail exactly because it explains too much in hindsight. Only part of the information is valuable for the future, and the art of intuition is to focus on that part and ignore the rest. A simple rule that relies only on the best clue has a good chance of hitting on that useful piece of information.
Looking backwards can hurt; you might end up blindsided by what the future can hold. The data might show trends that are only valid for the sample, and not the population as a whole.  Your results won't be ecologically valid if all the data is taken into consideration.  Especially since correlations change substantially over time, Gaussian methods are more likely to offer the pretense of knowledge than knowledge itself.

This has massive policy implications.  From Gigerenzer:
According to the complex strategy, the best predictors for a high dropout rate were the school`s percentage of Hispanic students, students with limited English, and black students-in that order. In contrast, Take the Best ranked attendance rate first, then writing score, then social science test score. On the basis of the complex analysis, a policy maker might recommend helping minorities to assimilate and supporting the English as a second language program. The simpler and better approach instead suggests that a policy maker should focus on getting students to attend class and teaching them the basics more thoroughly. Policy, not just accuracy, is at stake.
Yet with these policy issues at stake, it's surprising that fast and frugal algorithms aren't used more in economics research.  One disadvantage of take the best is that it doesn't give much quantitative accuracy.  It only tells which value is higher, but not by how much.  But how much does that matter?  While multiple regression may give more statistically significant coefficients, do we really have the power to tune the economy that much?  Even the best of natural experiments don't result in parameters that don't change through time.  Romer and Romer beautifully estimate tax elasticity, but how arrogant would a person need to be to build our entire tax policy based on an estimation from an almost 80 year old data set?  DSGE quantitative accuracy is such a joke that peripheral ad-hoc models are needed to make them even somewhat useful.

These alternative statistical tools are likely to add much to our insight of models.  The development of robust heuristics will be critical in a complex world, in which calculation becomes increasingly difficult and Gaussian methods increasingly fragile.

Tuesday, March 13, 2012

Nominal GDP Targeting and Complexity

The Complexity View

I've recently started rereading passages of The Black Swan: The Impact of the Highly Improbable and I find it fascinating. The prose is fluid, and the arguments are powerful. Much of the book mocks economic theory, as models tend to minimize the role of large shocks that defy normal distributions.  In the book, Taleb inserts the following chart that shows how much these "outliers" influence the stock market.

Taleb places the blame for these large swings in the market on the shoulders of the Federal Reserve.  He argues that stabilization policy actually makes the economy more fragile, making it more likely to go down in a dramatic fashion once the "big one" hits.  He sums up this argument in the following quote from a section titled "Beware Manufactured Stability" in a supplementary essay.

...fear of volatility, leading to interference with nature to impose "regularity" makes us more fragile across so many domains.  Preventing small forest fires sets the grounds for more extreme ones; giving out antibiotics when it is not very necessary makes us more vulnerable to severe epidemics... 
Which brings me to another organism: economic life. Our aversion to variability and desire for order and our acting on it has helped precipitate severe crises... Another thing we saw in the 2008 debacle: the U.S. government (or, rather, the Federal Reserve) had been trying for years to iron out the business cycle, making us exposed to a severe disintegration. This is the sort of reasoning I have against "stabilization" policies and manufacturing a nonvolatile environment ...

In a sense, the reduction of volatility in the Great Moderation was only an illusion of stability.  We were, as Taleb would say, "sitting on a pile of dynamite," unaware of the risk that lay underneath.

Impact on NGDP Targeting Policy

This kind of critique seems rather damning against nominal GDP targeting.  The typical analysis of NGDP targeting hinges on the assertion that low volatility implies high stability.  But what if this isn't true?  What if these times of low volatility are just times of high fragility?  Some analysis of the arguments for NGDP targeting even suggest mechanisms by which this is the case.  Debt problems are waved away because NGDP is stable, financial opacity becomes a non-issue because monetary policy compartmentalizes it,  perceptions of "safe" assets  change because expectations of nominal growth are maintained.  Stable expectations permit these innovations because agents can plan ahead, allowing for higher growth.

However, this higher efficiency comes at the cost of redundancy.  Taleb jokes in an interview with Russ Roberts that:
An economist would never design a human being with two lungs and two kidneys. It's wasteful. Deadweight loss.  
He follows up with:
So, the opposite of spare parts would be debt. And nature doesn't like debt. Nature likes redundancies. This mechanism of overreaction is redundancy.
And this is what terrifies me about NGDP targeting.  The incredibly stable regime creates an environment in which redundancy is eschewed in favor of fragility.  Perhaps it would be better to have a more resilient economy that wouldn't be able to accumulate as much capital, but one that has lower levels of debt.  The cost of a mistake in an NGDP targeting world would be incredible.  Even if, theoretically, under a stable monetary regime, there are no demand-side recessions, would you be willing to bet the stability of the entire global financial system on it?  Even if it were true, can you guarantee the Fed will be able to maintain a "stable monetary regime" for perpetuity?

I'm not trying to say the current monetary system is ideal; the dismal employment numbers firmly reject that view.  But when we look onto NGDP targeting as the solution to the global economic malaise, we need to be careful that we don't put all of our eggs into one basket.  NGDP targeting is an incredible tool for monetary policy; but it can't be a panacea for all of these troubles.  

This critique of NGDP targeting brings up another key issue for the design of policy.  Optimal policy has to do more than maximize welfare, it must also be robust to errors.  While in the game playing, platonic world of models NGDP targeting should create incredible reductions in volatility and instability, what are the possible effects on global fragility?  Policy engineering needs to take into account Murphy's Law: "If anything can go wrong, it will."  The only question is how we prepare.