Showing posts with label Econometrics. Show all posts
Showing posts with label Econometrics. Show all posts

Tuesday, December 21, 2021

Estimating the Effect of Physical Infrastructure on Economic Growth

I have a new working paper coauthored with Govinda Timilsina of the World Bank and my PhD student Debasish Das. It is a panel data study of the effect of various forms of infrastructure on the level of GDP. 

Compared to existing studies, we use more recent data, include new types of infrastructure such as mobile phones, and provide separate estimates for developing and developed countries. We find larger effects than most previous studies. We also find that infrastructure has a larger effect in more recent years (1992-2017) than in earlier years (1970-1991), and the effects of infrastructure are higher in developing economies than in industrialized economies. The long-run effects seem to be much larger than the initial impact. We also tried to estimate the effect of infrastructure on the rate of economic growth. Controlling for the initial level of GDP per worker we found a null result. So, we can't say that having more infrastructure means a more rapid rate of economic growth.

Getting good quality data that is comparable across countries is really a problem in this area of research. Many types of infrastructure only have data available for a few years. The ones that have more panel-like data often suffer from differences in definition across countries – such as what is a road or a motorway – or unexplained jumps in individual countries. So, our results are subject to a lot of measurement error.

Our main analysis uses data on five types of infrastructure – roads, railways, electric generation capacity, fixed line telephones, and mobile telephones*:

Following some previous research, we aggregate the individual types of infrastructure using principal component analysis. We use two principal components. One factor seems to be related to transport infrastructure and the other to electricity and telecommunications. Still, we can recover estimates of the effect of each individual type of infrastructure.

Also following some previous research, we use the Pooled Mean Group estimator to estimate a dynamic panel regression model. This allows us to test for the weak exogeneity of the explanatory variables, allowing us to give the results a somewhat causal interpretation.

The table shows the percentage change in GDP per worker for a 1% change in each infrastructure type. Getting standard errors for these estimates would be rather tricky.** Interestingly, the PMG estimates are mostly much larger than the static fixed effects estimates. Static fixed effects can be expected to converge to a short-run estimate of the effects while PMG should be a better estimate of long-run effects. Fixed effects also tends to inflate the effects of measurement error

Maybe the most innovative thing in the paper is that we plot the impulse response functions of GDP with respect to a 1% increase in each of the two main types of infrastructure:

PC1 is electricity and communications and PC2 transport infrastructure.*** Long-run effects of infrastructure are much larger than the short-run effects. In the short run, transport infrastructure even has a negative impact.

* Note that the graphs show the country means of these variables, while we actually use the deviations from those means over time in each country

** We only estimate the GDP-infrastructure relationship, but I think we would need time series models for each of the explanatory variables in order to sample from those models' residuals in a bootstrapping procedure. Bootstrapping is needed because we first carry out the principal components analysis and then estimate the PMG model in a second stage. These elasticities are combinations of the parameters from those two models.

*** We could get a confidence interval for these impulse response functions if we assume that the explanatory variables in the PMG model are deterministic as this analysis assumes...


Wednesday, May 12, 2021

Fifth Francqui Lecture: Econometric Modeling of Climate Change

The video of my fifth and final Francqui lecture on the econometric modeling of climate change is now on Youtube:  


The lecture begins by introducing the issue of global climate change. The first image of the Earth's energy balance is from an IPCC assessment report. Probably, the 4th Assessment Report. The graph of global temperature is the Berkeley Earth combined land and sea series. The graph of CO2 concentration is based on the data we used in our Journal of Econometrics paper updated with recent observations from Hawaii. The original source of the global CO2 emissions series is the now defunct CDIAC website updated from the BP Statistical Review of World Energy. Following that are three charts from the IPCC 5th Assessment Report. World sulfur dioxide emissions are from the CEDS datasite.

The next section – "Why Econometrics" – opens with a graph of the relationship between economic growth and CO2 emissions, which I put together from World Bank, International Energy Agency, and BP data sources.

The following section – "Do GHG Emissions Cause Climate Change?" starts with original research using the temperature and CO2 time series in the previous graphs. The CO2 concentration acts as a proxy variable for all radiative forcing in this analysis. It then goes on to present results from my 2014 paper with Robert Kaufmann published in Climatic Change. Details of the data are given in that paper.

Finally, I presented my paper coauthored with Stephan Bruns and Zsuzsanna Csereklyei, which was published in the Journal of Econometrics.

Thursday, October 15, 2020

Climate Econometrics and the Carbon Budget

Though I recently abandoned a follow up paper on our Journal of Econometrics climate modeling paper, we are now working on a different one. I'm scheduled to give a presentation (remotely) on it at the American Geophysical Union conference in December. In the course of our research, I ran some simple simulations on our Journal of Econometrics model. This model is a two equation vector autoregression of global surface temperature and radiative forcing with energy balance restrictions imposed. This is done using the concept of multicointegration. But it is still a simple time series model once the complicated estimation is complete.

I ran three scenarios that all have the same peak level of radiative forcing equivalent to doubling CO2:

Single Shock: Forcing is doubled in one year and then the system is allowed to move to equilibrium.

Shock and Maintain: Forcing is doubled suddenly and then that level of forcing is maintained forever. This is the scenario in our published paper and is used to estimate the equilibrium climate sensitivity in general circulation models.

Transient: Forcing is increased linearly for 70 years until the CO2 equivalent would be doubled. Then emissions are cut to zero.

This is what happens to temperature in the three scenarios:

Under the Shock and Maintain scenario, we reach the equilibrium climate sensitivity of 2.78ºC. Under the Transient scenario, the temperature increases by 1.85ºC when emissions are cut to zero and then continues to increase by about 0.3ºC before flatlining. Under the Single Shock scenario, temperature increases quite rapidly, reaching equilibrium in around 40 years with only a 0.98ºC increase.

This is what happens to radiative forcing in the three scenarios:

Under the Single Shock scenario there is a steep fall in forcing after the single pulse of greenhouse gases. A new equilibrium concentration and temperature is reached. Under the Transient scenario the equilibrium level of forcing is much higher even though in both cases emissions are cut to zero. Of course, much more carbon would need to be pumped into the atmosphere to achieve the Transient path as all the time carbon is also being absorbed. This shows the importance of the carbon budget. The total amount of emissions, not just the peak concentration matters. It is interesting that our very simple model seems to pick this up from the data without imposing any information about the carbon budget on the model.



Tuesday, February 19, 2019

Energy Efficiency Improvements Do Not Save Energy

I have a new working paper out, coauthored with Stephan Bruns and Alessio Moneta, titled: "Macroeconomic Time-Series Evidence That Energy Efficiency Improvements Do Not Save Energy". It's another paper from our ARC funded project: "Energy Efficiency Innovation: Diffusion, Policy and the Rebound Effect". We estimate the economy-wide effect on energy use of energy efficiency improvements in the U.S. We find that the rebound is around 100%, implying that in the long run energy efficiency improvements do not save energy or reduce greenhouse gas emissions.


At the micro level, we might naïvely expect a 1% improvement in energy efficiency to reduce energy use by 1%. But people adjust their behavior. Efficiency improvements reduce the cost of energy services like heating, transport, or lighting. Because these are now cheaper to produce, people consume more of them, and so the percentage reduction in energy use is less than the improvement in efficiency. This is known as the direct rebound effect.

People might also redirect their spending to consume more of complementary goods, like larger houses in the case of residential heating improvements, and reduce their consumption of substitute goods and services, like bus rides or cycling, in the case of car fuel economy improvements. These changes have implications for the energy used to produce these goods and services. Additionally, the reduction in demand for energy should lower the price of energy further boosting the rebound in energy use. Finally, the improvement in energy efficiency is an increase in productivity, which should result in economic growth. Higher incomes mean higher demand for energy. Adding these indirect rebound effects to the direct rebound effect we get the economy-wide rebound effect.

The size of the economy-wide rebound effect is crucial for estimating the contribution that energy efficiency improvements can make to reducing energy use and greenhouse gas emissions. Our study provides the first empirical general equilibrium estimate of the economy-wide rebound effect. Previous studies use simulation models, known as computable general equilibrium models, or partial equilibrium econometric models that don't allow the price of energy to adjust. Some of the latter studies also measure rebound incorrectly, for example assuming that energy intensity – energy used per dollar of GDP – measures energy efficiency. In fact, the majority of the rebound effect happens when energy intensity rebounds as people shift to more energy intensive consumption after an energy efficiency improvement. Economic growth induced by the efficiency improvement is expected to contribute less to total rebound.

We use a structural vector autoregressive model, or SVAR, that is estimated using search methods developed in machine learning. We apply the SVAR to U.S. monthly and quarterly data. An SVAR explains changes in the vector of variables, x, in terms of its past values and a vector of serially and mutually uncorrelated shocks, ε:

In our basic model, the vector, x, contains three variables: primary energy use, GDP, and the price of energy. The first of the shocks is a shock to energy use, holding constant shocks to GDP and the price of energy and the past values of all three variables. We think this is a reasonable definition of an energy efficiency shock. The other two shocks are income and price shocks.

The matrix, B, which transmits the shocks to the dependent variables cannot be estimated without imposing some restrictions or conditions on the model. Usually economists use economic theory to impose restrictions on the coefficients in B (short-run restrictions) and the Π_i (long-run restrictions). Alternatively, they sample a range of models, rejecting only those that don't meet qualitative "sign restrictions" on the matrix B. Instead, we use independent component analysis, an approach that is relatively new to econometrics. This imposes conditions on the nature of the shocks instead and estimates B without direct restrictions. Unlike the short- and long-run restrictions approach, it doesn't impose a priori restrictions on the data, and unlike the sign restrictions approach, it estimates a unique model.

Using the estimated SVAR model we compute the impulse response functions of the dependent variables to the shocks:


The top left graph shows the effect of an energy efficiency shock on energy use. The grey shading is a 90% confidence interval, the x-axis is in months, and the y-axis in log units.

Initially, an energy efficiency shock strongly reduces energy use, but this effect wears off over the following years as consumers and the economy adjusts. Eventually, there is no change in energy use so that rebound is 100%.

The other graphs in the first column show the effect of the energy efficiency shock on GDP and the price of energy. The second column shows the effect of a shock to GDP, and the final column an energy price shock.

The implications for policy are that encouraging energy efficiency innovation is unlikely to make a contribution to reducing greenhouse gas emissions. This is one reason why I am skeptical of projections that predict that energy intensity will fall much faster in the future than in the past because of energy efficiency policies.

On the other hand, if these policies raise rather than reduce the costs of producing energy services then the direct rebound (and presumably the economy-wide rebound) will be negative rather than positive. As, apart from their environmental effects, these would reduce economic welfare, it seems that there would be better options to reduce emissions by switching to low carbon energy.

Saturday, February 10, 2018

A Multicointegration Model of Global Climate Change

We have a new working paper out on time series econometric modeling of global climate change. We use a multicointegration model to estimate the relationship between radiative forcing and global surface temperature since 1850. We estimate that the equilibrium climate sensitivity to doubling CO2 is 2.8ºC – which is close to the consensus in the climate science community – with a “likely” range from 2.1ºC to 3.5ºC.* This is remarkably close to the recently published estimate of Cox et al. (2018).

Our paper builds on my previous research on this topic. Together with Robert Kaufmann, I pioneered the application of econometric methods to climate science – Richard Tol was another early researcher in this field. Though we managed to publish a paper in Nature early on (Kaufmann and Stern, 1997), I became discouraged by the resistance we faced from the climate science community. But now our work has been cited in the IPCC 5th Assessment Report and recently there is also a lot of interest in the topic among econometricians. This has encouraged me to get involved in this research again.

We wrote the first draft of this paper for a conference in Aarhus, Denmark on the econometrics of climate change in late 2016 and hope it will be included in a special issue of the Journal of Econometrics based on papers from the conference. I posted some of our literature review on this blog back in 2016.

Multicointegration models, first introduced by Granger and Lee (1989), are designed to model long-run equilibrium relationships between non-stationary variables where there is a second equilibrium relationship between accumulated deviations from the first relationship and one or more of the original variables. Such a relationship is typically found for flow and stock variables. For example, Granger and Lee (1989) examine production, sales, and inventory in manufacturing, Engsted and Haldrup (1999) housing starts and unfinished stock, Siliverstovs (2006) consumption and wealth, and Berenguer-Rico and Carrion-i-Silvestre (2011) government deficits and debt. Multicointegration models allow for slower adjustment to long-run equilibrium than do typical econometric time series models because of the buffering effect of the stock variable.

In our model there is a long-run equilibrium between radiative forcing, f, and surface temperature, s:
The equilibrium climate sensitivity is given by 5.35*ln(2)/lambda. But because of the buffering effect of the ocean, surface temperature takes a long time to reach equilibrium. The deviations from equilibrium, q, represent a flow of heat from the surface to (mostly) the ocean. The accumulated flows are the stock of heat in the Earth system, Q. Surface temperature also tends towards equilibrium with this stock of heat:
where u is a (serially correlated but stationary) random error. Granger and Lee simply embedded both these long-run relations in a vector autoregressive (VAR) time series model for s and f. A somewhat more recent and much more powerful approach (e.g. Engsted and Haldrup, 1999) notes that:
where F is accumulated f and S is accumulated s. In other words, S(2) = s(1)+s(2), S(3) = s(1)+s(2)+s(3) etc. This means that we can estimate a model that takes into account the accumulation of heat in the ocean without using any actual data on ocean heat content (OHC) ! One reason that this is exciting, is because OHC data is only available since 1940 and data for the early decades is very uncertain. Only since 2000 is there a good measurement network in place. This means that we can use temperature and forcing data back to 1850 to estimate the heat content. Another reason that this is exciting is that F and S are so-called second order integrated variables (I(2)) and estimation with I(2) variables, though complicated, is super-super consistent – it is easier to get an accurate estimate of a parameter despite noise and measurement error issues in a relatively small sample. The I(2) approach combines the 2nd and 3rd equations above into a single relationship which it embeds in a VAR model that we estimate using Johansen's maximum likelihood method. The CATS package which runs on top of RATS can estimate such models as can the Oxmetrics econometrics suite. The data we used in the paper is available here.

This graph compares our estimate of OHC (our preferred estimate is the partial efficacy estimate) with an estimate from an energy balance model (Marvel et al., 2016) and observations of ocean heat content (Cheng et al, 2017):


We think that the results are quite good, given that we didn't use any data on OHC to estimate it and that the observed OHC is very uncertain in the early decades. In fact, our estimate cointegrates with these observations and the estimated coefficient is close to what is expected from theory. The next graph shows the energy balance:


The red area is radiative forcing relative to the base year. This is now more than 2.5 watts per square meter – doubling CO2 is equivalent to a 3.7 watt per square meter increase. The grey line is surface temperature. The difference between the top of the red area and the grey line is the disequilibrium between surface temperature and radiative forcing according to the model. This is now between 1 and 1.5 watts per square meter and implies that, if radiative forcing was held constant from now on, that temperature would increase by around 1ºC to reach equilibrium.** This gap is exactly balanced by the blue area, which is heat uptake. As you can see, heat uptake kept surface temperature fairly constant during the last decade and a half despite increasing forcing. It's also interesting to see what happens during large volcanic eruptions such as Krakatoa in 1883. Heat leaves the ocean, largely, but not entirely, offsetting the fall in radiative forcing due to the eruption. This means that though the impact of large volcanic eruptions on radiative forcing is short-lived, as the stratospheric sulfates emitted are removed after 2 to 3 years, they have much longer-lasting effects on the climate as shown by the long period of depressed heat content after the Krakatoa eruption in the previous graph.

We also compare the multicointegration model to more conventional (I(1)) VAR models. In the following graph, Models I and II are multicointegration models and Models IV to VI are I(1) VAR models. Model IV actually includes observed ocean heat content as one of its variables, but Models V and VI just include surface temperature and forcing. The graph shows the temperature response to permanent doubling of radiative forcing:


The multicointegration models have both a higher climate sensitivity and respond more slowly due to the buffering effect. This mimics, to some degree, the response of a general circulation model. The performance of Model IV is actually worse than the bivariate I(1) VARs. This is because it uses a much shorter sample period than Models V and VI. In simulations that are not reported in the paper, we found that a simple bivariate I(1) VAR estimates the climate sensitivity correctly if the time series is sufficiently long - much longer than the 165 years of annual observations that we have. This means that ignoring the ocean doesn't strictly result in omitted variables bias as I previously claimed. Estimates are biased in a small sample, but not in a sufficiently large sample. That is probably going to be another paper :)

* "Likely" is the IPCC term for a 66% confidence interval. This confidence interval is computed using the delta method and is a little different to the one reported in the paper.
** This is called committed warming. But, if emissions were actually reduced to zero, it's expected that forcing would decline and that the decline in forcing would about balance the increase in temperature towards equilibrium.

References

Berenguer-Rico, V., Carrion-i-Silvestre, J. L., 2011. Regime shifts in stock-flow I(2)-I(1) systems: The case of US fiscal sustainability. Journal of Applied Econometrics 26, 298—321.

Cheng L., Trenberth, K. E., Fasullo, J., Boyer, T., Abraham, J., Zhu, J., 2017. Improved estimates of ocean heat content from 1960 to 2015. Science Advances 3(3), e1601545.

Cox, P. M., Huntingford, C., Williamson, M. S., 2018. Emergent constraint on equilibrium climate sensitivity from global temperature variability. Nature 553, 319–322.

Engsted, T. Haldrup, N., 1999. Multicointegration in stock-flow models. Oxford Bulletin of Economics and Statistics 61, 237—254.

Granger, C. W. J., Lee, T. H., 1989. Investigation of production, sales and inventory relationships using multicointegration and non-symmetric error correction models. Journal of Applied Econometrics 4, S145—S159.

Kaufmann R. K. and D. I. Stern (1997) Evidence for human influence on climate from hemispheric temperature relations, Nature 388, 39-44.

Marvel, K., Schmidt, G. A., Miller, R. L., Nazarenko, L., 2016. Implications for climate sensitivity from the response to individual forcings. Nature Climate Change 6(4), 386—389.

Siliverstovs, B., 2006. Multicointegration in US consumption data. Applied Economics 38(7), 819–833.

Wednesday, October 26, 2016

The Ocean in Climate Econometrics

Third excerpt (previous excerpts):


Most studies of global climate change using econometric methods have ignored the role of the ocean. Though these studies sometimes produce plausible estimates of the climate sensitivity, they universally produce implausible estimates of the rate of adjustment of surface temperature to long-run equilibrium. For example, Kaufmann and Stern (2002) find that the rate of adjustment of temperature to changes in radiative forcing is around 50% per annum even though they estimate an average global climate sensitivity of 2.03K. Similarly, Kaufmann et al. (2006) estimate a climate sensitivity of 1.8K, while the adjustment coefficient implies that more than 50% of the disequilibrium between forcing and temperature is eliminated each year. Furthermore, the autoregressive coefficient in the carbon dioxide equation of 0.832 implies an unreasonably high rate of removal of CO2 from the atmosphere. The methane rate of removal is also very high.

Simple AR(1) I(1) autoregressive models of this type assume that temperature adjusts in an exponential fashion towards the long run equilibrium. The estimate of that adjustment rate tends to go towards that of the fastest adjusting process in the system, if, as is the case, that is the most obvious in the data. Schlesinger et al. (no date) illustrate these points with a very simple first order autoregressive model of global temperature and radiative forcing. They show that such a model approximates a model with a simple mixed layer ocean. Parameter estimates can be used to infer the depth of such an ocean. The models that they estimate have inferred ocean depths of 38.7-185.7 meters. Clearly, an improved time series model needs to simulate a deeper ocean component.
Stern (2006) used a state space model inspired by multicointegration. The estimated climate sensitivity for the preferred model is 4.4K, which is much higher than previous time series estimates and temperature responds much slower to increased forcing. However, this model only used data on the top 300m of the ocean and the estimated increase in heat content in the pre-observational period seems too large.

Pretis (2015) estimates an I(1) VAR for surface temperature and the heat content of the top 700m of the ocean for observed data for 1955-2011. The climate sensitivity is 1.67K for the preferred model but 2.16K for a model, which excludes the level of volcanic forcing from the radiative forcing aggregate, entering only as first differences. With two cointegrating vectors it is not possible to “read off” the rate of adjustment of surface temperature to increased forcing and Pretis does not simulate impulse or transient response functions.

References

Kaufmann, R. K., Kauppi, H., Stock, J. H., 2006. Emissions, concentrations, and temperature: a time series analysis. Climatic Change 77(3-4), 249-278.

Kaufmann, R. K., Stern, D. I., 2002. Cointegration analysis of hemispheric temperature relations. Journal of Geophysical Research 107(D2), Article No. 4012.

Pretis, F., 2015. Econometric models of climate systems: The equivalence of two-component energy balance models and cointegrated VARs. Oxford Department of Economics Discussion Paper 750.

Schlesinger, M. E., Andronova, N. G., Kolstad, C. D., Kelly, D. L., no date. On the use of autoregression models to estimate climate sensitivity. mimeo, Climate Research Group, Department of Atmospheric Sciences, University of Illinois at Urbana-Champaign, IL.

Stern, D. I., 2006. An atmosphere-ocean multicointegration model of global climate change. Computational Statistics and Data Analysis 51(2), 1330-1346.  

Sunday, September 11, 2016

Progress on New Climate Modeling Paper

As I mentioned a couple of posts ago, we are working on a new climate modeling paper. We just started estimating models. This graph shows predicted ocean heat content in units of 10^22 Joules, blue, and a 5 year moving average of observed heat content in the top 2000m of the ocean (which is only available from 1959*):


We only used data on global temperature and radiative forcing and the most basic estimator possible to produce this prediction. It's in the right ballpark in terms of the increase in heat content and even some of the wiggles match up (the levels are "arbitrary"). Diagnostic statistics look fairly good too. I think we can only improve on this prediction using more sophisticated estimators. Watch this space :)

* NOAA assign the middle year of each 5 year window as the date of the data. We assign the last year of each 5 year window instead.

Monday, May 23, 2016

Should We Test for Cointegration Using the Johansen Procedure If We Want to Estimate a Single Equation Static Regression?

A student from Cuba asked me:

"I want to apply the DOLS methodology... I have read several books and research works about DOLS but none of them explain clearly how to test cointegration in this case.... I asked some professors about this issue and one of them told me that I should apply the Johansen cointegration test."

It's quite easy to find papers that do this - first test for cointegration using the Johansen procedure, report only the cointegration test statistics, and if they can be used to reject the null hypothesis of non-cointegration then use some other method such as Dynamic Ordinary Least Squares (DOLS) to estimate a static single equation regression model. These researchers aren't actually interested in the complete vector autoregression (VAR) system, which is OK. I've reviewed quite a lot of papers that use this approach.

If your model has more than two variables (one dependent variable and one explanatory variable) then this is a very bad idea. The cointegration test statistics from the Johansen procedure (if they reject the null) say nothing about the cointegration properties of your single equation regression model.

The following simple example shows why. Imagine we have three variables, X1, X2, and X3 with the following "data generation process":


where epsilon 1 is a stationary stochastic process and epsilon 2 and 3 are simply white noise. Variables X2 and X3 follow simple random walks. Variable X1 cointegrates with X2. But X3 is a random walk that has nothing to do with the other two variables. If you estimate a VAR with these variables and do the Johansen cointegration test, you should expect to find that there is one cointegrating vector. But the following regression:


will not cointegrate. It is a spurious regression because it includes X3 which is an unrelated random walk. We cannot rely on finding that the VAR "cointegrates" to assume that this regression also cointegrates. Only X1 and X2 cointegrate in this example. Of course, it is possible that X1, X2, and X3 are jointly cointegrated but as this example shows, that doesn't have to be the case.

How can we avoid this? The cointegrating vector in this case is [1, -beta1]. We could test within the Johansen procedure whether we can restrict the cointegrating vector to not include a coefficient for X3. Unlike gamma3 in the static regression, if X3 does not belong in the cointegrating relationship, then this coefficient is expected to be zero. We can and should also test the residuals of the static regression to see if they cointegrate.

Friday, May 13, 2016

"Replicating" the Climate Contest

My previous post discussed Doug Keenan's climate contest. I wondered how accurate we could actually expect to be in such a situation. I assume that the temperature series is a simple random walk, possibly with a constant drift term. We want to see how accurately we can determine whether there is a drift term in the random walk or not.

So, again just using Excel, I created 1000 series of 134 observations each distributed as Normal(mu, 0.11), where mu is the drift term. For 250 series I set mu to 0.01, for 250 series I set it to -0.01 and for 500 to 0. I then compute the usual t-test for the significance of the sample mean for each series.

Only 127 t-tests were significant at the 5% level and 201 at the 10% level. Using a 10% significance level, statistical power - correct rejection of the incorrect null hypothesis of no drift - is 29%. Using a 5% significance level, power is 20%. There is no distortion of the actual "size" of the test - the number of incorrect rejections of the true null.

So, combining this information, if you use this method and a 10% significance level you will get 595 correct classifications of whether a random walk has a drift or does not have a drift, which is far below the 900 required to win the contest.

Of course, it seems that Keenan's data is a bit more complicated than this and may or may not have any relevance to the actual nature of climate data or the nature of the climate change problem.

You can download my data here. The first column is the drift term used and the first row indicates the years and the statistics columns.

More Mathiness in Climate Econometrics: Doug Keenan's Climate Change Contest

My colleague Robert Kaufmann got an e-mail from Doug Keenan inviting him to participate in his "climate change contest" without the usual $10 submission fee. I hadn't heard about this contest and went to the site to investigate. So, Keenan has produced 1000 time series of 135 observations each that are somehow derived from random numbers and then added a plus 1 or minus 1 per 100 observations trend to some of these. The series have been calibrated so that the they could potentially reproduce in some way the observed global temperature time series from 1880 to the present without an added trend. The task of the contestant is to determine for each series whether it has an added trend or not. If any submission gets 90% of these or more right by 30th November, this year, that submission will win $100,000.

Keenan's idea is that no-one can validly detect with 90% accuracy whether there is a trend in temperature or not. Therefore, the IPCC's claim that temperature has definitely increased over the last century and it is very likely that this is due to human activity must be wrong.

I downloaded the data. Looking at some of the series it's pretty clear that they are some sort of random walk (stochastic trend). It is not simply a series of random numbers (white noise) with a linear trend added. I haven't bothered to write a program to test this. Assuming that they are simple random walks, I tested in Excel whether the mean of first difference was different to zero for each of the thousand series. Only 8 of the series have a mean first difference that is significantly different to zero at the 5% level using the standard calculation of the standard error of the mean, which assumes that the first differences are white noise. If they were normally distributed white noise and none of the original series had an added trend then we would expect about 50 of the means of the first differences to be significantly different to zero by the definition of statistical significance. So, something else seems to be going on here. I expect that statistical power to detect a non-zero drift term of 0.01 or -0.01 when the standard deviation of the first differences is 0.11 is in any case rather low. Perhaps we could use structural time series methods, but statistical power of 90% at a significance level of 10% is a lot to ask for in this situation. I created my own dataset to see how many series one could expect to correctly classify - statistical power using a simple data generating process and a simple test was 29% for a 10% significance level test. This means that we can only correctly classify 595 of the 1000 series.

The real question to ask is whether Keenan's thought experiment makes sense. I would argue that it doesn't. His argument is that if temperature follows some kind of integrated process then it is very hard to determine whether it has a drift component or will sooner or later just stochastically trend down again. Therefore, we can't know if temperature has statistically significantly increased or not. But theory and climate models predict that global temperature should be stationary if radiative forcing is constant. If we detect a random walk or a close to random walk signal in the temperature data then something else is happening. Research can then try to determine if it is likely to be due to anthropogenic factors or not. It is possible that we make a type 1 error - falsely rejecting the null hypothesis - but we can determine how likely that is. So, in my opinion, Keenan's contest is another case of mathiness in climate econometrics.

Tuesday, February 23, 2016

Mathiness in Climate Change Econometrics

Terence Mills has a "white paper" on the Global Warming Policy Foundation Website. It predicts little future increase in temperature. Not surprisingly, The Australian has published a totally positive article about it. I commented in the comments there:

"Mills assumes that past fluctuations in temperature are purely random and of unknown causes and ignores greenhouse gases, or the sun, or volcanic eruptions, or any other specific factor that might drive climate change. He then fits simple statistical models based on this assumption to the data. Not surprisingly, if you assume that there isn't any specific factor driving the climate, your best forecast for the future is for not much change because you don't know what random shocks will show up to change the climate in the future. A more sensible approach is to test which of the various proposed drivers might actually have an effect and how large that effect has been. There are a lot of refereed academic papers that do just that including some I published myself. It's pretty easy to show that greenhouse gases have an effect on the climate, it's quite big (but fairly uncertain how big), and if emissions continue on a business as usual path there will be a lot of increase in temperature."

More technically: Mills fits univariate ARIMA models to HADCRUT,  RSS global lower troposphere series (only available since 1980) and Central England Temperature series. These include models with no deterministic component (an ARIMA(0,1,3) model of HADCRUT) and a model with a deterministic trend with breakpoints chosen based on "eyeballing" the temperature graph. None of these models predicts any future warming, because there is no trend in the trendless model and because the "hiatus" means there is no recent trend in the segmented trend model. Of course, a model with just a single linear deterministic trend fitted to HADCRUT data would forecast a lot of warming in the 21st Century, though with a very wide forecast error envelope. But that model isn't estimated, for some reason...

This is a prime case of "mathiness" I think - lots of math that will look sophisticated to many people used to build a model on silly assumptions with equally silly conclusions.

In other news, my paper coauthored with Luis Sanchez on drivers of greenhouse gas emissions is now published in Ecological Economics. It is open access till 12th April.

P.S. This post was cited in the Daily Mail.

Wednesday, January 13, 2016

Between and Within

This will be obvious to anyone with a good understanding of econometrics, but it is quite stunning really to think that all the information you see in the first set of graphs in my previous post on the EKC is thrown away by fixed effects panel estimators. That is because the graphs plot the mean value over time in each country of the dependent variable against the mean value over time in each country of the explanatory variable. Fixed effects estimation first deducts these means from the data and then estimates the regression of the two residuals using ordinary least squares. This is why fixed effects is also called the "within estimator" because the "between (country) variation" you see in these graphs is ignored. Of course, you can estimate a model that just exploits this between variation using the between estimator.*

The reason the latter estimator is rarely used is because researchers are worried about omitted variables bias. Any omitted variables are subsumed in the error term while the fixed effects estimator eliminates their country specific means and so reduces the potential bias. Hauk and Wacziarg (2009), however, found that when there is also measurement error in the explanatory variables (which can also bias the regression estimates) the between estimator performs well compared to alternatives. Fixed effects estimation tends to inflate the effect of the measurement error.

Differenced estimators sweep out any country fixed effects in the differencing operation.** So they also remove all the between variation in the data. However, they do allow us to include country characteristics that are constant over time to explain differences in growth rates across countries, which standard fixed effects does not allow.***

* The linked paper was eventually published in Ecological Economics.
** For a two period panel, fixed effects and first differences produce identical results.
*** There are variations of fixed effects that can allow this.

Sunday, August 9, 2015

The Extent and Consequences of P-Hacking in Science

Interesting paper from my biology colleagues at ANU on the effects of "p-hacking" - searching for more significant results by looking at various statistical models or samples and picking the more significant ones to report - on reported science. They conclude that when there is a strong real effect it can be detected despite p-hacking by looking at the "p-curve". The p-curve is the distribution of p-values across all the studies collected in a meta-analysis. If the curve is skewed right - there is a peak at very high significance levels (numbers a lot smaller than 5%) then there is a real effect. However, p-hacking can inflate the estimated size of the effect if we use a simple average of effect sizes in the literature. The main novelty of their paper I think is that they collected a large number of p-values from various fields of science using text-mining to test these ideas in the empirical literature.

In meta-analysis in economics, a popular approach is to test the effect of degrees of freedom or precision (inverse of the standard error) on the values of the reported test statistics using regression analysis. This effect is called the power-trace. The idea is that if there is a true effect, then, due to increasing statistical power, reported test statistics will be more significant the higher the degrees of freedom in the underlying study.* Some of these methods can also be used to estimate the true effect size adjusted for publication bias.

In our meta-analysis of energy-GDP Granger causality tests we also present graphs of the distribution of the test-statistics. These seemed to be roughly normal with a mean of about 1, which means there is excess significance in this literature but that the mean test statistic is not statistically significant (the solid histogram in the background is the standard normal distribution):


To help interpret these graphs, note that a normal test statistic (-probit(p)) of zero means that the original Granger causality test p-value was 0.5. A test statistic of 1.65 implies that the original p-value was 0.05 and a test statistic of -1.65 implies that the p-value was 0.95. The econometric analysis in the paper showed that there was no statistically significant relationship between these test statistics and degrees of freedom, also suggesting that there was no genuine effect. We showed in the paper that there did seem to be a robust effect from GDP to energy when underlying studies controlled for energy prices.

We didn't report the actual p-values though, and so I am curious what the p-curves look like. First I made a couple of histograms with bins for each 1% increment of p-values:



Uh-oh! The mode is for 0-1%! According to Head et al.'s methodology that means there is a true effect in each direction of causality. When I broke down the range from p=0 to p=0.1 into 100 bins, again the mode was for the smallest value. So, what does it mean when the overwhelming majority of studies find results that are less significant than the 1% or 0.1% level and yet the mode is for 0-1% or 0-0.1%? And when these results are for not particularly large sample sizes? Either the p-curve or the meta-regression/power trace method is wrong here. One hypothesis is that non-stationarity in macro-economic time series and the over-fitting problem discussed in our paper result in many spuriously significant test statistics in relatively small samples that wouldn't arise with more classically behaved data.

* Though this method can detect a "genuine effect" there is no guarantee that this is a "causal effect". If no studies control for the relevant variables or effects to identify a causal effect then the meta-analyst won't be able to detect a causal effect either. Similarly, if the meta-analyst doesn't control for all the relevant variables included in the underlying studies they may also fail to identify a causal effect when some papers do identify one. All the meta-analyst can find is a robust partial correlation in the underlying studies if one exists.

Wednesday, June 17, 2015

Meta-Granger Causality Testing

I have a new working paper with Stephan Bruns on the meta-analysis of Granger causality test statistics. This is a methodological paper that follows up on our meta-analysis of the energy-GDP Granger causality literature that was published in the Energy Journal last year. There are several biases in the published literature on the energy-output relationship, which we document in the Energy Journal paper:

1. Publication bias due to sampling variability - the common tendency for statistically significant results to be preferentially published. This is either because journals reject papers that don't find anything significant or more likely because authors don't bother submitting papers without significant results. So they either scrap studies that don't find anything significant or data-mine until they do. This means that the published literature may over-represent the frequency of statistically significant tests. This is likely to be a problem in many areas of economics, but especially in a field where results are all about test statistics and not about effect sizes.

2. Omitted variables bias - Granger causality tests are very susceptible to omitted variables bias. For example, energy use might seem to cause output in a bivariate Granger causality test because it is highly correlated with capital. This is a very serious problem in the actual empirical Granger causality literature, which I noted in my PhD dissertation.

3. Over-fitting/over-rejection bias - In small samples, there is a tendency for vector autoregression model fitting procedures to select more lags of the variables than the true underlying data generation process has. There is also a tendency to over-reject the null hypothesis of no causality in these over-fitted models. This means that a lot of Granger causality results from small sample studies are spurious. We realized in our Energy Journal paper that this was also a serious problem in the empirical Granger causality literature. The following graph illustrates this using studies from the energy-output causality literature:


Each graph shows normalized test statistics for causality in one of the two directions. Rather than fit models with more lags in larger samples, researchers tend to deplete the degrees of freedom by adding more lags. Therefore, there tend to be fewer degrees of freedom for studies with three lags than with two, and fewer for those with two than with one. Also, we see that the average significance level increases as the lags increase and degrees of freedom reduces.

Of course, the second two types of biases give researchers additional opportunities to select statistically significant results for publication and so, more generally, "publication bias" includes selection of statistically significant results from those provided by sampling variability and by various biases.

The standard meta-regression model used in economics deals with the first of the three biases by exploiting the idea that if there is a genuine effect then studies with larger samples should have more statistically significant test statistics than smaller studies. If there is no real effect then there will be either no relation or even a negative relation between significance and sample size. Meta-analysis can test for the effects of omitted variables bias by including dummy variables and interaction terms for the different variables included in primary studies. Finally, in our Energy Journal paper we controlled for the over-fitting/over-rejection bias by including the number of degrees of freedom lost in model fitting in our meta-regression.

The new paper focuses on the latter issue and examines both the potential prevalence of over-fitting and over-rejection and the effectiveness of controlling for over-fitting. The approach used in this paper is a little different to the Energy Journal paper - here we include the number of lags selected as the control variable. We show by means of Monte Carlo simulations that, even if the primary literature is dominated by false-positive findings of Granger causality, the meta-regression model correctly identifies the absence of genuine Granger causality. The following graphs show the key results:

Power is the probability of reject the null hypothesis of no causality when it is incorrect - so here we have set up a simulated VAR where there is causality from energy to GDP. Mu is the mean sample size of the primary studies in our simulation and var is the variance. So, the lefthand graph is a simulation of mostly small sample studies. The middle one has a mixture of small and large studies, and the right hand graph has mostly large studies (but a few small ones too). The meta sample size is the number of studies that are brought together in the meta-analysis. DGP2a is a data generating process with a small effect size - DGP2b has a larger effect size.

So, what do these graphs show? When the samples in primary studies are small and we only have a meta sample of 10 or 20 studies, it is hard to detect a genuine effect, whatever we do. When the effect size is small it is still hard to detect an effect even when we have 80 primary studies using the traditional economics meta-regression model ("basic model"). Our "extended model" which controls for the number of lags really helps a lot in this situation. With large primary study sizes it is quite easy to detect a true effect with only 20 studies in the meta-analysis and our method adds little value. However, the energy-GDP causality literature has mostly small similar sized samples and is trying to detect what is quite a small effect in the energy causes GDP direction (elasticity of 0.05 or 0.1). Our approach has much to offer in this context.

Wednesday, December 17, 2014

Global Energy Use: Decoupling or Convergence?

I have another new working paper coauthored with Zsuzsanna Csereklyei out titled: "Global Energy Use: Decoupling or Convergence?". This paper follows up on our paper on the the stylized facts of energy and growth using the methods from our paper on modeling the emissions-income relationship using long-run growth rates.
In particular, we focus on the stylized facts that there is a stable relationship between energy use and GDP per capita over time and that there has been convergence in energy intensity over time.

The following graph shows the relationship between the long-run growth rates of energy use and GDP per capita from 1971 to 2010:
As we saw in our previous paper, higher economic growth rates are associated with higher rates of growth in energy use, though there is quite a lot of variation in individual countries. As the intercept of the regression line is zero, on average there is no time effect that is reducing energy use across all countries in the absence of economic growth. On the other hand, many developed countries like the UK, Sweden, or even the United States have seen declines or little increase in energy use per capita in recent decades. So, what can explain that?

Our model adds additional explanatory variables to the regression illustrated above. To test whether there is decoupling, so that growth has less effect or even a negative effect on energy use at higher income levels we include an interaction term between the growth rate and level of income per capita. This is the same idea as the test for the environmental Kuznets curve in the Anjum et al. paper. We find that the regression coefficient for this interaction term is actually positive! So, actually growth has larger rather than smaller effects on energy use in higher income countries. However, we find that the level of income has a negative effect on the growth rate of energy use. This means that at higher income levels there is an improving energy efficiency effect so that energy use declines over time ceteris paribus. We call this "weak decoupling". The following graph shows the contribution of these and other effects to the growth in energy use in different groups of countries:
Though, as shown by the black dots, the growth rate of energy use  per capita is highest in upper middle income countries, the contribution of economic growth is greater in high income countries. But this positive effect of growth is offset by "weak decoupling" and other effects.

The most important of these other effects globally is convergence. Countries that had high energy intensities at the beginning of the period saw declines in energy intensity, ceteris paribus. But, this effect was most important in the low income countries, some of which were the most energy intensive in our sample in 1971. In the high income countries, convergence raised energy use on average. But in the US and Canada it contributed -1.0% and -0.9% p.a., respectively, to reducing energy use. So, there is a lot of variation across countries.

Projections and forecasts of future energy use should not, therefore, assume that economic growth will be associated with decreased energy use in the future. Instead, the scale effect seems to be alive and well. On the other hand, there appear to be improvements in energy efficiency across high income countries irrespective of their growth rates or their initial level of energy intensity. These would tend to moderate the growth in energy use as countries get richer at the upper end of the income continuum. At the lower end of the income continuum the same effects serve to raise energy intensity. But, some of the major reductions in energy intensity in countries, such as in the United States and China, have probably been the result of convergence towards the global mean, and so are unlikely to be reproduced in the future.

On the more technical side there are a couple of innovations in this paper too. We extend the method we used in the growth rates paper to allow for a spatially correlated error term. If there are omitted variables which are spatially correlated and the explanatory variables are also spatially correlated then it is likely that the two are correlated and regression coefficient estimates will be inconsistently estimated. To deal with this problem we use an approach called spatial filtering. This removes the spatial autocorrelation by adding additional variables to the regression which model the spatial process. These are in fact the eigenvectors of a transformation of the spatial weighting matrix. With 93 countries there are 93 eigenvectors, so the tricky part is deciding which of these eigenvectors to include in the regression. Tiefelsdorf and Griffith (2007) developed a procedure to do this, which we use.

Another issue is that if we want to give our model a causal interpretation, then we have to assume that GDP growth causes growth in energy use and not vice versa. Obviously changes in energy use might also cause GDP. We argue though that the latter effect is probably quite small compared to the effect of GDP on energy and so the estimated effect of GDP on energy is only biased upwards by a relatively small amount. The income elasticity of energy is likely to be close to unity whereas the elasticity of GDP w.r.t. to energy might be only 0.05. This is possibly one reason why it has been so hard for researchers to find robust signs of Granger causality from energy use to GDP. On other hand, this is maybe why a simple naive regression of GDP on energy use appears to find a very large effect of energy on GDP. The estimated regression coefficient is biased upwards by the effect of income on energy demand.

Wednesday, November 12, 2014

Article Published Today in PLoS ONE

My paper "High-Ranked Social Science Journal Articles Can Be Identified from Early Citation Information" was published today in PLoS ONE. For background on the paper, check out my blogpost or this story on the Crawford School website.

I'm travelling tomorrow to Kassel where I will be working with Stephan Bruns on what we think is a really awesome (Stephan's words :)) extension to this paper. We are also hoping to finish an econometric theory paper we are writing.

Tuesday, April 8, 2014

RATS Command RESTRICT

I've spent the last couple of days trying to replicate results in RATS that Chunbo produced using STATA. In the process we found a lot of bugs in our data-processing and computer codes but now we can replicate each others results. We are estimating a translog cost share system together with the cost function. Previously, I have used the RATS command SUR to estimate a system of seemingly unrelated regressions and then the command RESTRICT(replace) to impose the restrictions and SUR(CREATE) to produce the restricted SUR estimates. This did not reproduce the same results as STATA at all. Instead using NLSYSTEM in RATS (despite the fact that the system is linear), I managed to reproduce the same results as STATA. This now seems to be the recommended way to do this type of analysis according to the RATS User Guide. So, I really don't know what the estimates produced by RESTRICT in RATS represent and I strongly recommend not to use them. This shows yet again that it is very important to know what the computer code you are using is actually doing.

Monday, December 2, 2013

Blunt Instruments

"Blunt Instruments: Avoiding Common Pitfalls in Identifying the Causes of Economic Growth" by Samuel Bazzi and Michael Clemens is an interesting read. I have long thought that it is hard to find valid instrumental variables in macroeconomics, this paper provides lots of evidence for this.

So-called "endogeneity" of explanatory variables in regression analysis is mainly due to the following:
  • Reverse causality - Feedback from the dependent variable to the explanatory variable.
  • Omitted variables bias - When variables that are correlated with the included variables and the dependent variable are excluded from the regression, the included variables are attributed as explaining too much or too little of the variance in the dependent variable.
  • Measurement error - When there is error in measuring the explanatory variables their regression coefficients tend to be biased towards zero.
All of these lead to correlation between the explanatory variables and the true underlying error term (but by design not to the actual estimated regression residuals). One approach to dealing with these issues is instrumental variables regression where variables that are correlated with the explanatory variables but supposedly not with the true residuals are introduced. The instrumental variable estimates effectively only use the part of the explanatory variables associated with the variation in the instrument in estimating the regression coefficient. Bazzi and Clemens point out that there are many pairs of econometric studies that claim to have found the same strong and valid instruments for different explanatory variables but don't include the explanatory variables included in the other studies in their models. If a study finds an instrumental variable is strongly correlated with an explanatory variable and it is used as an instrument in another study that doesn't include the explanatory variable in question then the instrument must be invalid as it will be correlated with the error term. Bazzi and Clemens show that many studies published in top journals suffer from these problems.

Of course, there is much more in the paper, but I think that is the key point.


Wednesday, October 16, 2013

Econometric Approach to Detection and Attribution of Climate Change in IPCC AR5 Report

It seems odd to put the full Working Group 1 report on the open web but then say that it shouldn't be quoted or cited. But, anyway, it's nice to see that a fairly extensive discussion of the econometric approach to detecting and attributing climate change made it into the final draft of the report. The text of the final draft of the Working Group 3 report is just in the process of being submitted to the TSU. Last Friday was supposed to be the deadline. The government approval session will take place in April next year. So, some way to go to publication for us.

Thursday, September 26, 2013

Meta-Analysis Paper Accepted at The Energy Journal

My paper with Stephan Bruns and Christian Gross on meta-analysis of the energy-GDP causality literature has been accepted (subject to some minor editing corrections) at the Energy Journal. Back in March, I wrote a blogpost about this paper when we put out the working paper version. In this final version we mainly added some separate results for OECD and non-OECD countries, which while different do not change the overall results very much. As you will see, both my coauthors have nicely developing research track records and, if you are hiring, they are looking for jobs (probably in Germany or nearby countries)!