Showing posts with label Sociology of Science. Show all posts
Showing posts with label Sociology of Science. Show all posts

Sunday, August 7, 2022

Trends in RePEc Downloads and Abstract Views

For the first time in a decade, I updated my spreadsheet on downloads and abstract views per person and per item on RePEc.


The downward trends I identified ten years ago have continued, though there was an uptick during the pandemic, which has now dissipated. There was more of an increase in abstract views than in downloads in the pandemic.

Since the end of 2011 both abstract views and downloads per paper have fallen by about 80%. Total papers rose by around 260%, while total downloads fell 38% and total abstract views 27%. 

I'd guess that a mixture of the explanatory factors I suggested last time has continued to be in play.


Wednesday, May 12, 2021

Public Policy Schools in the Asia-Pacific Ranked

I have a new paper with my Crawford School colleague Bjoern Dressel published in Asia & the Pacific Policy Studies (open access). The data and figures for the article are on Figshare. Bjoern has been interested for a while in ranking public policy schools in the Asia-Pacific region.  But a comprehensive ranking seemed hard to achieve. Recently, I came across an article by Ash and Urquiola (2020) that ranks US public policy schools according to their research output and impact. Well, we thought, if they can rank schools just by their research output and not by their education and public policy impact then so can we 😀. Research is the easiest component to evaluate.

We compare the publication output of 45 schools with at least one publication listed in Scopus between 2014 and 2018, based on affiliations listed on the publications rather than current faculty. We compute the 5-Year impact factor for each school. This is identical to the impact factor reported for academic journals, but we compute it for a school rather than a journal. It is the mean number of citations received in 2019 by a publication published between 2014 and 2018. This can be seen as an estimate of research quality. We also report the standard error of the impact factor as in my 2013 article in the Journal of Economic Literature. If we treat the impact factor as an estimate of the research quality of a school then we can construct a confidence interval to express how certain or uncertain we are about that estimate. This graph shows the schools ranked by impact factor with a 90% confidence interval:

Peking and Melbourne are the two top-ranked schools but the point estimates have a very wide confidence interval. This is because their research output is relatively small and the variance of citations is quite large. The third ranked school – SGPP in Indonesia – only had two publications in our target period. After that there are several schools with much narrower confidence intervals. These mostly have more publications.


Here we can see the impact factors on the y-axis and the number of publications of each school on the x-axis. Three schools clearly stand out at the right: Crawford, Lee Kwan Yew, and Tsinghua. These schools are also top-ranked by total citations, which combines the quality and quantity variables. The three top schools account for 54% of publications and 63% of citations from the region.

In general, the elite schools are in China and Australia. Australia has three out of the top ten schools ranked by impact factor and total citations, despite its small population size. China, on the other hand has at least five schools ranked in the top ten across both rankings, which is remarkable given that many of these schools have been established only in the last 15 years (though linked to well-established research universities).

We found more schools that had no publications in Scopus in the target period. Perhaps in some cases they are too new, or faculty use their other affiliations, but clearly there is a lot of variation in research-intensiveness. Somewhat surprising is the low ranking of public policy schools in Japan and India – both countries with a considerable number of public policy schools, but none in the top ten schools when ranked by 5-year citation impact factor or total number of citations. 

One reason for the strong performance of the Chinese schools is that they focus to some degree on environmental issues, and particularly climate change, where citation numbers tend to be higher. We did not adjust for differences in citations across fields in this research, but this is something that future research should address.



Wednesday, August 5, 2020

Abandoning a Paper

Now and then it's time to give up on a project. In September 2018, I attended a climate econometrics conference at Frascati near Rome. For my presentation, I did some research on the performance of different econometric estimators of the equilibrium climate sensitivity (ECS) including the multicointegrating vector autoregression (MVAR) that we used in our paper in the Journal of Econometrics. The paper included estimates using historical time series observations (from 1850 to 2014), a Monte Carlo analysis, estimates using output of 16 Global Circulation Models (GCMs), and a meta-analysis of the GCM results.


The historical results, which are mostly also in the Journal of Econometrics paper, appear to show that taking energy balance into account increases the estimated climate sensitivity. By energy balance, we mean that if there is disequilibrium between radiative forcing and surface temperature the ocean must be heating or cooling. Surface temperature is in equilibrium with ocean heat, and in fact follows ocean heat much more closely than it follows radiative forcing. Not taking this into account results in omitted variables bias. Multicointegrating estimators model this flow and stock equilibirum. The residuals from a cointegrating relationship between the temperature and radiative forcing flows are accumulated into a heat stock, which in turn cointegrates with surface temperature. If we have actual observations on ocean heat content or radiative imbalances we can use them. But available time series are much shorter than those for surface temperature or radiative forcing. The results also suggested that using a longer time series increases the estimated climate sensitivity.

The Monte Carlo analysis was supposed to investigate these hypotheses more formally. I used the estimated MVAR as the model of the climate system and simulated the radiative forcing series as a random walk. I made 2000 different random walks and estimated the climate sensitivity with each of the estimators. This showed that, not surprisingly, the MVAR was an unbiased estimator. The other estimators were biased using a random walk of just 165 periods. But when I used a 1000 year series all estimators were unbiased. In other words, they were all consistent estimators of the ECS. This makes sense, because in the end equilibrium is reached between forcing and surface temperature. But it takes a long time.

Each of the GCMs I used has an estimated ECS ("reported ECS") from an experiment where carbon dioxide is suddenly increased fourfold. I was using data from a historical simulation of each GCM, which uses the estimated historical forcings over the period 1850 to 2014. A major problem in this analysis is that the modelling teams do not report the forcing that they used. This is because the global forcing that results from applying aerosols etc depends on the model and the simulation run. So, I used the same forcing series that we used to estimate our historical models. This isn't unprecedented, Marvel et al. (2018) do the same.

In general, the estimated ECS were biased down relative to the reported ECS for the GCMs, but again, the estimators that took energy balance into account seemed to do better. In an meta-analysis of the results, I compared how much the reported radiative imbalance (=ocean heat uptake roughly) from each GCM increased to how much the energy balance equation said it should increase using the reported temperature series, reported ECS, and my radiative forcing series. A regression analysis showed, that where the two matched, the estimators that took energy balance into account were unbiased, while those that did not match, under-estimated the ECS.

These results seemed pretty nice and I submitted the paper for publication. Earlier this year, I got a revise and resubmit. But when I finally got around to working on the paper post-lockdown and post-teaching things began to fall apart.

First, I came across the Forster method of estimating the radiative forcing in GCMs. This uses the energy balance equation:

where F is radiative forcing, T is surface temperature, and N is radiative imbalance. Lambda is the feedback parameter. ECS is inversely proportional to it. The deltas indicate the change since some baseline period. Then, if we know N and T, both of which are provided in GCM results, we can find F! So, I used this to get the forcing specific to each GCM. The results actually looked nicer than in the originally submitted paper. These are the results for the MVAR for 15 CMIP5 GCMs:


The rising line is a 45 degree line, which marks equality between reported and estimated ECSs. The multicointegrating estimators were still better than the other estimators. But there wasn't any systematic variation in the degree of underestimation that would allow us to use a meta-analysis to derive an adjusted estimate of the ECS.

This is still OK. But then I read and re-read more research on under-estimation of the ECS from historical observations. The recent consensus is that estimates from recent historical data will inevitably under-estimate the ECS because feedbacks change from the early stages after an increase in forcing to the latter stages as a new equilibrium is reached. The effective climate sensitivity is lower at first and greater later.

OK, even if we have to give up on estimating the long-run ECS, my estimates are estimates of the historical sensitivity. Aren't they? The problem is that I used the long-run ECS to derive the forcing from the energy balance equation. So, the forcing I derived is wrong. It is too low. I could go back to using the forcing I used previously, I guess. But now I don't believe the meta-analysis of that data is meaningful. So, I have a bunch of estimates using the wrong forcing with no way to further analyse them.

I also revisited the Monte Carlo analysis. By the way I had an on-and-off again coauthor through this research. He helped me a lot with understanding how to analyse the data. But he didn't like my overly bullish conclusions on the submitted paper and so withdrew his name from it. But he was maybe going to get back on the revised submission. He thought that the existing analysis which used an MVAR to produce the simulated data was maybe biased unfairly in favour of the MVAR. So, I came up with a new data-generating process. Instead of starting with a forcing series I would start with the heat content series. From that I would derive temperature, which needs to be in equilibrium with heat content and then using the energy balance equation derive the forcing. To model the heat content I fitted a unit root autoregressive model (stochastic trend) to the heat content reported from the Community GCM with the addition of a volcanic forcing explanatory variable. The stochastic trend represents anthropogenic forcing. The Community GCM is one of the 15 GCMs I was using and it has temperature and heat content series that look a lot like the observations. I then fitted a stationary autoregressive model for temperature with the addition of the heat content as an explanatory variable. The simulated model used normally distributed shocks with the same variance as these fitted models and volcanic shocks.

As an aside, the volcanic shocks were produced by the model:
where rangamma(0.05) are random numbers drawn from a standard gamma distribution with shape parameter 0.05. This is supposed to produce the stratospheric sulfur radiative forcing, which decays over a few years following an eruption. Here is an example realisation:

The dotted line is historical volcanic forcing and the solid line a simulated forcing. My coauthor said it looked "awesome".

So, again, I produced two sets of 2000 datasets. One with a sample size of 165 and one with a sample size of 1000. Now, even in the smaller sample, all four estimators I was testing produced essentially identical and unbiased results! I ran this yesterday. So, our Monte Carlo result disappears. I can't see anything unreasonable about this data generating process, which produces completely different results to the one in the submitted paper. So, I don't see anything to justify one over the other. So, this was the point where I gave up on this project.

My coauthor, who is based in Europe, is on vacation. Maybe he'll see a way to save it when he comes back, but I am sceptical.

Saturday, January 6, 2018

How to Count Citations If You Must

That is the title of a paper in the American Economic Review by Motty Perry and Philip Reny. They present five axioms that they argue a good index of individual citation performance should conform to. They show that the only index that satisfies all five axioms is the Euclidean length of the list of citations to each of a researcher's publications – in other words, the square root of the sum of squares of the citations to each of their papers.* This index puts much more weight on highly cited papers and much less on little cited papers than simply adding up a researcher's total citations would. This is a result of their "depth relevance" axiom. A citation index that is depth relevant always increases when some of the citations of a researcher's less cited papers are instead transferred to some of the researcher's more cited papers. In the extreme, it rewards "one hit wonders" who have a single highly cited paper, over consistent performers who have a more extensive body of work with the same total number of citations.

The Euclidean index is an example of what economists call constant elasticity of substitution, or CES, functions. Instead of squaring each citation number, we could raise it to a different power, such as 1.5, 0.5, or anything else. Perry and Reny show that the rank correlation between the National Research Council peer-reviewed ranks of the top 50 U.S. economics departments and the CES citation indices of the faculty employed in those departments is at a maximum for a power of 1.85:



This is close to 2 and suggests that the market for economists values citations in a similar way to the Euclidean index.

RePEc acted unusually quickly to add this index to their rankings. Richard Tol and I have a new working paper that discusses this new citation metric. We introduce an alternative axiom: "breadth relevance", which rewards consistent achievers. This axiom states that a citation index always increases when some citations from highly cited papers are shifted to less cited papers. We also reanalyze the dataset of economists at the top 50 U.S. departments that Perry and Reny looked at and a much larger dataset that we scraped from CitEc for economists at the 400 international universities ranked by QS. Unlike Perry and Reny, we take into account the fact that citations accumulate over a researcher's career and so junior researchers with few citations aren't necessarily weaker researchers than senior researchers with more citations. Instead, we need to compare citation performance within each cohort of researchers measured by the years since they got their PhD or published their first paper.

We show that a breadth relevant index that also satisfies Perry and Reny's other axioms is a CES function with exponent of less than one. Our empirical analysis finds that the distribution of economists across departments is in fact explained best by the simple sum of their citations, which is equivalent to a CES function with exponent of one, that favors neither depth nor breadth. However, at lower ranked departments – departments ranked by QS from 51 to 400 – the Euclidean index does explain the distribution of economists better than does total citations.


In this graph, the full sample is the same dataset that Perry and Reny used in their graph. The peak correlation is for a lower exponent – tau or sigma** – simply because we take into account cohort effects by computing the correlation for a researcher's citation index relative to the cohort mean.*** While the distribution across the top 25 departments is similarly to the full sample, with a peak at a slightly lower exponent that is very close to one, we don't find any correlation between citations and department rank for the next 25 departments. It seems that there aren't big differences between them.

Here are the correlations for the larger dataset that uses CitEc citations for the 400 universities ranked by QS:


For the top 50 universities, the peak correlation is for an exponent of 1.39 but for the next 350 universities the peak correlation is for 2.22. The paper also includes parametric maximum likelihood estimates that come to similar conclusions.

Breadth per se does not explain the distribution of researchers in our sample, but the highest ranked universities appear to weight breadth and depth equally, while lower-ranked universities do focus on depth, giving more weight to a few highly cited papers.

A possible speculative explanation of behavior across the spectrum of universities could be as follows. Lowest-ranked universities, outside of the 400 universities ranked by QS, might simply care about publication without worrying about impact. Having more publications would be better than having fewer at these institutions, suggesting a breadth relevant citation index. Our exploratory analysis that includes universities outside of those ranked by QS supports this. We found that breadth was inversely correlated with average citations in the lower percentiles.

Middle-ranked universities, such as those ranked between 400 and 50 in the QS ranking, care about impact; having some high-impact publications is better than having none and a depth-relevant index describes behavior in this interval. Finally, among the top-ranked universities such as the QS top 50 or NRC top 25, hiring and tenure committees wish to see high-impact research across all of a researcher's publications and the best-fit index moves towards. Here, adding lower-impact publications to a publication list that contains high-impact ones is seen as a negative.

* As monotonic transformations of the index also satisfy the same axioms, the simplest index that satisfies the axioms is simply the sum of squares.

** In the paper, we refer to an exponent of less than one as tau and an exponent greater than one as sigma.

*** The Ellison dataset that Perry and Reny use, uses Google Scholar data and truncates each researcher's publication list at 100 papers. With all working paper variants, it's not hard to exceed 100 items. This could bias the analysis in favor of depth rather than breadth. We think that the correlation computed for researchers with 100 papers or less only is a better way to test whether depth or breadth best explains the distribution of economists across departments. The correlation peaks very close to one for this dataset.

Tuesday, October 10, 2017

What Do Crawford School Economists Do?

I'm doing quite a bit of background work for our School Review, a review of the Future of Asia-Pacific Economics etc. The following table is based on the self-identified "Fields of Research" of core Crawford economics faculty. Most people chose more than one field. If, for example, someone chose three fields, then I attributed 1/3 of an FTE to each for that person. The result looks like this:


Our research foci are economic development and growth, environmental and resource economics, and international economics and finance. The (non-geographical) fields that we rank best in globally in RePEc are: Environment 7, Energy 7, Resources 11, Agriculture 23, Growth 29, International Trade 32, Development 39. So, this focus also is where we perform well.

Most Crawford economists have countries that they focus on. Using a similar approach I put together this table:


Naturally, Australia is number one, then follow China, Japan, Indonesia, and Vietnam. In RePEc, we rank 4th in the SE Asia ranking, 18th in Central/Western Asia (which actually includes South Asia), and 39th in the China subject ranking. This reflects more of our historical focus, while the current faculty is more focused on NE Asia. We don't have any current faculty with a professed interest in Thailand, for example! Of course, there is also less competition in research on SE Asia than on China and so that will also affect our ranking.

Tuesday, March 28, 2017

Cohort Size and Cohort Age at Top US Economics Departments

I'm working on a new bibliometrics paper with Richard Tol. We are using Glenn Ellison's data set on economists at the top 50 U.S. economics departments as a testbed for our ideas. I had to compute the size of each year cohort for one of our calculations, and thought this graph of the number of economists at the 50 departments in each "academic age" year was interesting:


There isn't as sharp a post-tenure drop-off in numbers as you might expect, given the supposed strict tenure hurdle these departments impose. But as we can see the cohorts increase in size up to year 5, which might be explained by post-docs and other temporary appointments, or people even moving up the rankings after a few years at a lower ranked department. So, as a result, the tenure or out year would be spread over a few years too. On the other hand, as the data were collected in 2011, the Great Recession might also explain lower numbers for the first few years.

A post-retirement drop-off only really seems to occur after 39 years. The oldest person in the study by academic age was Arnold Harberger.

Thursday, December 29, 2016

Ranking Economics Institutions Applying a Frontier Approach to RePEc data

Back in 2010 I posted that the RePEc ranking of economics institutions needed to be adjusted by size. Better quality institutions do tend to be bigger but as RePEc just sums up publications, citations etc rather than averaging them larger institutions also get a higher RePEc ranking even if they aren't actually better quality. In the post, I suggested using a frontier approach. The idea is that the average faculty member at Harvard perhaps is similar to one at Chicago (I haven't checked this), but because Harvard is bigger it is better. So, looking at average scores of faculty members might produce a misleading ranking.

A reader sent me an e-mail query about an updated version of this and I thought that was a good idea for a new post:


The chart shows the RePEc rank for 190 top-level institutions (I deleted NBER) against their number of registered people on RePEc. I drew a concave frontier by hand. How have things changed since 2010? The main change is the appearance of Stanford on the frontier. Also, the Federal Reserve is now listed as one institution, so the Minnesota Fed has dropped off the frontier. Dartmouth is now slightly behind the frontier and Tel Aviv looks like it has also lost a little ground. Otherwise, not much has changed.

Monday, July 25, 2016

Data and Code for Our 1997 Paper in Nature

I got a request for the data in our 1997 paper in Nature on climate change. I didn't think I'd be able to send the actual data we used as I used to follow the practice of continually updating the datasets that I most used rather than keeping an archival copy of the data actually used in a paper. But I found a version from February 1997, which was the month we submitted the final version of the paper. I got the RATS code to read the file and with a few tweaks it was producing the results that are in the paper. These are the results for observational data in the paper, not those using data from the Hadley climate model. I have now put up the files on my website. In the process I found this website - zamzar.com - that can convert .wks to .xls files. Apparently, recent versions of Excel can't read the .wks Lotus 1-2-3 files that were a standard format 20 or more years years ago. For those that don't know, Lotus 1-2-3 was the most popular spreadsheet program before Microsoft introduced Excel. I used it in the late 80s and early 90s when I was in grad school.

Tuesday, July 12, 2016

Legitimate Uses for Impact Factors

I wrote a long comment on this blogpost by Ludo Waltman but it got eaten by their system, so I'm rewriting it in a more expanded form as a blogpost of my own. Waltman argues, I think, that for those that reject the use of journal impact factors to evaluate individual papers, such as Lariviere et al., there should be then no legitimate uses for impact factors. I don't think this is true.

The impact factor was first used by Eugene Garfield to decide which additional journals to add to the Science Citation Index he created. Similarly, librarians can use impact factors to decide on which journals to subscribe or unsubscribe from and publishers and editors can use such metrics to track the impact of their journals. These are all sensible uses of the impact factor that I think no-one would disagree with. Of course, we can argue about whether the mean number of citations that articles receive in a journal is the best metric and I think that standard errors - as I suggested in my Journal of Economic Literature article - or the complete distribution as suggested by Lariviere et al., should be provided alongside them.

I actually think that impact factors or similar metrics are useful to assess very recently published articles, as I show in my PLoS One paper, before they manage to accrue many citations. Also, impact factors seem to be a proxy for journal acceptance rates or selectivity, which we only have limited data on. But ruling these out as legitimate uses doesn't mean rejecting the use of such metrics entirely.

I disagree with the comment by David Colquhoun that no working scientists look at journal impact factors when assessing individual papers or scientists. Maybe this is the case in his corner of the research universe but it definitely is not the case in my corner. Most economists pay much, much more attention to where a paper was published than how many citations it has received. And researchers in the other fields I interact with also pay a lot of attention to journal reputations, though they usually also pay more attention to citations as well. Of course, I think that economists should pay much more attention to citations too.


Thursday, January 14, 2016

People's Ability to Delude Themselves is Amazing

The Australian reported a couple of days ago that the University of Wollongong gave a PhD for a thesis by an anti-vaccination activist, Judy Wilyman. It's the comments on the article where the delusion is amazing. Many people comment that it is totally outrageous that the University of Wollongong gave this PhD because obviously anti-vax is total nonsense and a conspiracy theory. At the same time, some of them are complaining that climate scepticism doesn't get sufficient respect from academia. Of course, climate scepticism is just as much nonsense and a conspiracy theory as anti-vax.* But these people believe that one of these theories is totally correct and the other totally bogus. I'm a bit surprised, as I thought that both these theories were right-wing anti-government theories. Apparently not in Australia?

This is, of course, exactly the same as people who are convinced that their religion is true and all other religions are false.

* I am open-minded about both anti-vaccination and climate change sceptical hypotheses. Also UFOs, yetis...

Sunday, August 9, 2015

The Extent and Consequences of P-Hacking in Science

Interesting paper from my biology colleagues at ANU on the effects of "p-hacking" - searching for more significant results by looking at various statistical models or samples and picking the more significant ones to report - on reported science. They conclude that when there is a strong real effect it can be detected despite p-hacking by looking at the "p-curve". The p-curve is the distribution of p-values across all the studies collected in a meta-analysis. If the curve is skewed right - there is a peak at very high significance levels (numbers a lot smaller than 5%) then there is a real effect. However, p-hacking can inflate the estimated size of the effect if we use a simple average of effect sizes in the literature. The main novelty of their paper I think is that they collected a large number of p-values from various fields of science using text-mining to test these ideas in the empirical literature.

In meta-analysis in economics, a popular approach is to test the effect of degrees of freedom or precision (inverse of the standard error) on the values of the reported test statistics using regression analysis. This effect is called the power-trace. The idea is that if there is a true effect, then, due to increasing statistical power, reported test statistics will be more significant the higher the degrees of freedom in the underlying study.* Some of these methods can also be used to estimate the true effect size adjusted for publication bias.

In our meta-analysis of energy-GDP Granger causality tests we also present graphs of the distribution of the test-statistics. These seemed to be roughly normal with a mean of about 1, which means there is excess significance in this literature but that the mean test statistic is not statistically significant (the solid histogram in the background is the standard normal distribution):


To help interpret these graphs, note that a normal test statistic (-probit(p)) of zero means that the original Granger causality test p-value was 0.5. A test statistic of 1.65 implies that the original p-value was 0.05 and a test statistic of -1.65 implies that the p-value was 0.95. The econometric analysis in the paper showed that there was no statistically significant relationship between these test statistics and degrees of freedom, also suggesting that there was no genuine effect. We showed in the paper that there did seem to be a robust effect from GDP to energy when underlying studies controlled for energy prices.

We didn't report the actual p-values though, and so I am curious what the p-curves look like. First I made a couple of histograms with bins for each 1% increment of p-values:



Uh-oh! The mode is for 0-1%! According to Head et al.'s methodology that means there is a true effect in each direction of causality. When I broke down the range from p=0 to p=0.1 into 100 bins, again the mode was for the smallest value. So, what does it mean when the overwhelming majority of studies find results that are less significant than the 1% or 0.1% level and yet the mode is for 0-1% or 0-0.1%? And when these results are for not particularly large sample sizes? Either the p-curve or the meta-regression/power trace method is wrong here. One hypothesis is that non-stationarity in macro-economic time series and the over-fitting problem discussed in our paper result in many spuriously significant test statistics in relatively small samples that wouldn't arise with more classically behaved data.

* Though this method can detect a "genuine effect" there is no guarantee that this is a "causal effect". If no studies control for the relevant variables or effects to identify a causal effect then the meta-analyst won't be able to detect a causal effect either. Similarly, if the meta-analyst doesn't control for all the relevant variables included in the underlying studies they may also fail to identify a causal effect when some papers do identify one. All the meta-analyst can find is a robust partial correlation in the underlying studies if one exists.

Sunday, July 12, 2015

Increasing Requirements for Publication

From "Accelerating Scientific Publication in Biology":

Somewhat tongue – in - cheek, let’s imagine a contemporary editorial decision on the 1953 Watson and Crick papers (assuming that they were submitted together):

“Dear Jim and Francis: Your two papers have now been seen by three referees. Based upon these reviews, I regret to say that we cannot offer publication at this time. While your model is very appealing, referee 3 finds that it is somewhat speculative and premature for publication. Indeed, your model proposing a semi-conservative replication of DNA raises many obvious questions. As two of the referees point out, it should be possible to determine experimentally if the two strands can separate and serve as templates. This would address referee 3’s concern that strand separation is not feasible thermodynamically. I regret to say that without such experimental evidence, we will not be able to publish your work in Nature and suggest publication in a more specialized journal. Should you be able to furnish more direct experimental evidence, we would be willing to reconsider such a revised paper. Naturally we would need to consult our referees once again. Furthermore, since space in our journal is at a premium, if you do decide to resubmit, then we recommend that you combine your two submitted papers into a single and more cohesive Article, potentially including the X-ray studies of your colleagues at Cambridge. Thank you again for submitting your papers to Nature. I am sure that this revision will delay your Nobel Prize and the discovery of the genetic code by only one or two years."

Sunday, February 22, 2015

How Has Research Assessment Changed the Structure of Academia?

Does measuring something change it?  In quantum mechanics measurement disturbs what is being measured, which is referred to as the observer effect. The same is often true in social systems, especially of course when measurement is attached to rewards. The UK and Australia have been conducting periodical research assessment exercises - the REF and ERA. In the case of the UK, research assessment started almost three decades ago. In Australia, the first research assessment was only conducted in 2010 but the founding of the ARC in 1988 and its independence in 2001 are both milestones in the road to increased emphasis on competition in research in Australia.

Johnston et al. (2014) show that the total number of economics students has increased in UK more rapidly than the total number of all students, but the number of departments offering economics degrees has declined, particularly in post-1992 universities. Also, the number of universities submitting to the REF under economics has declined sharply with only 3 post-1992 universities submitting in the latest round. This suggests that the REF has driven a concentration of economics research in the more elite universities in the UK. BTW the picture above is of the Hotel Russell, which the Russell Group of British universities is named after.

Neri and Rodgers (2014) investigate whether the increased emphasis on research in Australia has had the desired effect in the field of economics. They investigate the output of top economics research by Australian academics from 2001 to 2010. By constructing a unique database of 26,219 publications in 45 top journals, they compare Australia’s output internationally, determine whether Australia’s output increased, and rank Australian universities based on their output. They find that Australia’s output, in absolute and relative terms, and controlling for differences in page size and journal quality, increased and, on a per capita basis, is converging to the levels of the most research-intensive countries. Finally, they find that the historical dominance of the top four universities is diminishing. The correlation between the number of top 45 journal articles published in 2005-2010 and the ERA 2012 ranking is 0.83 (0.78 for 2003-8 and ERA 2010).

References

Johnston, J., Reeves, A. and Talbot, S. (2014). ‘Has economics become an elite subject for elite UK universities?’ Oxford Review of Education, vol. 40(5), pp. 590-609.

Neri, F. and Rodgers, J. (2014). ‘The contribution of Australian academia to the world’s best economics research: 2001 to 2010’, Economic Record.

Monday, August 4, 2014

Online Accessibility of Scholarly Literature, and Academic Innovation

Kevin Staub presented this paper today at ANU. They download data on all papers in fifty core economics journals over the decades of the 1990s and 2000s during which academic journals gradually went online. They test the effect of the fraction of references cited in article being online on the diversity of references cited. Diversity is measured by the average citation difference between articles in the reference list. Imagine I write an article and in the reference list I cite this paper by Kevin Staub and Martin Weitzman's article on recombinant growth cited by Staub then the distance between those two articles is 1. If I also cite Paul Romer's article "The Origins of Endogenous Growth" cited by Weitzman then the distance between Staub's paper and Romer's is 2 and between Weitzman and Romer is 1. Controlling for time and journal fixed effects and some other control variables they found that there were significant increases in the share of references with distances of 3 or greater the more of the reference list was available online. This seems expected to me as people find it easier to search for literature beyond the reference lists or even the forward citations of articles they have already seen as more of the literature is searchable online.

But the paper contains another result which I found much less expected and much more interesting that I think deserves a paper of its own. They found that papers with higher average distances between items on their reference lists received higher numbers of citations 20, 30, or 40 years down the track than papers with less diverse reference lists. So, this supports the notion that papers that bring together articles that were not previously cited together are more innovative. One might expect papers with more eclectic references to be produced by less professional more dilettantish authors. Of course, these papers were all published in the fifty core economics journals, so that probably acts to filter out the more outlandish papers or papers written by "outsiders" that are doomed to be ignored.


Thursday, June 19, 2014

High-Ranked Social Science Journal Articles Can Be Identified from Early Citation Information

I have posted a new bibliometric working paper , which investigates how well we can predict future cumulative citations from the first citations received by a paper in the disciplines of economics and political science.

It is usually assumed that citations accumulate too slowly in social sciences apart from psychology to be useful for short-term research assessment. For this reason, the Australian Government’s Excellence in Research for Australia (ERA) exercise, which attempts to assess the research quality of universities in the previous 5 years, uses peer review in social science disciplines apart from psychology for this reason but uses citation analysis for psychology and all natural sciences. This peer review process seems to me to be a wasteful duplication of effort to review research outputs that have already passed through a peer review process once.

I show that, surprisingly, citations received by journal articles in the social sciences in the first one to two years after publication are strongly predictive for citations received in future years. By contrast, I show that journal impact factors are mostly useful in the year of publication and their contribution to predicting citations declines rapidly thereafter.

If it is actually possible to predict citations fairly reliably in social science disciplines, then it should also be easy to predict them in the natural sciences. This means that it should be possible to expand bibliometric analysis in research evaluation exercises to all disciplines apart from the humanities and arts. It also means that we should pay attention to the early citations received by papers when we evaluate individual academics for hiring and promotion. Impact factors are reflective of journal selectivity, which we frequently do not have easily available data on. But they only explain about 16-17% of the variation in rankings of papers six years later conditional on the citations already received in the year of publication. The latter explain 13-14% of the variation. But at the end of the year following publication, accumulated citations explain 52-53% of the variation in cumulative citations after 6 years and 73% at the end of the second year after publication.

These models could be improved by adding information on the characteristics of the articles themselves and their authors, but that was much too time consuming to do for the almost 12,000 articles in my sample.

I have submitted a copy of my paper to the HEFCE inquiry on the use of metrics in research assessment.

Friday, May 23, 2014

Unionization in Australian Universities


After seeing that Alison Booth's paper from the Quarterly Journal of Economics was the most downloaded paper from ResearchGate at Crawford this week, I was curious what fraction of employees at Australian universities belonged to the National Tertiary Education Union. Apparently NTEU has 26,000 members. It also seems that there are 113,000 employees in the university sector. On that basis the unionization rate would be only 23%. Of course, quite a lot of those are casuals or PhD students working as lecturer A etc.* But the total number of full-time staff is 86,000. That implies 30%. Also there were 67,933 staff on continuing contracts. If the latter is the real target market for the union then the rate is 38%. Based on this, social custom doesn't work well in the university sector to overcoming free-riding. Let me know if any of my assumptions are wrong as this is the first time I've ever looked at this issue.

* There were 41,730 academics at levels B and above, but the union also represents non-academic staff.

Thursday, April 10, 2014

John List to Take Up Fractional Appointment at Monash





A coup for Monash University -  John List to take up fractional appointment at Monash!

One motivation for this move would be the ERA. But the census date for ERA 2015 is 31 March 2014. Staff need to be affiliated at that date for their prior publications to be counted. Also the ARC is cracking down on institutions claiming the publications of affiliates. For those employed in less than a 0.4 fractional position, at least one publication must list the institution as an affiliation on the publication. As a reviewer for ERA 2012, I think some institutions really abused the system with their claims of affiliates' publications in ERA 2012. So, this move by Monash is either a long term plan, or has nothing to do with the ERA.



Wednesday, February 12, 2014

Working in Policy and Working in Academic Research

Interesting blogpost on the differences. These are of course the extreme poles between someone doing solo-authored work in economic theory and someone work hands on in government or international policy. Academic research in economics is increasingly done in teams. Most of my ongoing projects are coauthored at the moment. Despite Deirdre McCloskey's criticisms, we are also very interested in the magnitude of effects - for example the size of the rebound effect or the climate sensitivity. And if you want to get a grant (at least in Australia) you have to convince academics outside your discipline. If you want to have a policy influence you have to convince non-academics. As someone at a school of public policy that is an important part of our mission. I also prefer to answer important questions even if it is hard to give a good answer to them, rather than less important questions which can be answered better. Of course, if we can't say anything novel enough to publish we have to drop the topic. There is a "sweet spot" where the question is both important and can be answered well, but that is difficult to find.

Saturday, December 7, 2013

Researchers Work Times Vary Around the World

If downloading papers from Springer = working then this paper by Wang et al has fascinating evidence on when researchers are working around the world. They got several days data on downloads of academic articles from Springer by location and time of day and composed download curves across the day for both weekdays and the weekend. Most of the cultural stereotypes hold up - late lunch in Spain and almost no lunch break in the US and UK. Australians tend to have a more defined workday than other English speakers. Americans, Chinese, and British work particularly hard at the weekend compared to other countries.