Alpha in Academia

Alpha in Academia

A Drift, Not an Event

[WITH CODE] The standard way of measuring publication-driven arbitrage returns a number whether or not publication does anything

Alpha in Academia's avatar
Alpha in Academia
Sep 18, 2026
∙ Paid

Hello and welcome back to another paid post!

Part 2 of 2. Part 1 measured what 210 published anomalies earned inside and outside the window their own papers examined, and found returns run at roughly half outside it in both directions. This piece takes the one number that looked like an arbitrage effect and finds that the method producing it would have produced it anyway.

The standard design compares an anomaly’s returns in the window after its paper’s sample ends to its returns after the paper is published, and attributes the difference to investors learning from the literature. Run on 210 anomalies, that design gives a residual of 13.4% of the in-sample return. Run on the same anomalies with the real publication dates thrown out and fake ones substituted, it gives 13.4%.

Across 200 randomized runs the average gap was −0.0897. The one measured from genuine publication dates was −0.0903, sitting at the 49.5th percentile of the random distribution. The reason is that anomaly returns drift downward across their whole post-sample life, so splitting that history at any point produces a lower second half. The estimator is built to return a negative number, and calling that number arbitrage attributes a trend to an event.

Let’s dive right in.


Introduction

Part 1 cut each anomaly’s history into four regimes on its own clock: the decades before its paper’s sample begins, the paper’s own window, the gap between the end of that window and publication, and everything since. The headline was that returns outside the discovery window run at about half those inside it, and that pre-discovery and post-publication are statistically indistinguishable from each other.

One thing was deliberately left hanging. The decay is not flat across the two post-sample regimes. Returns fall 0.247 percentage points once the paper’s sample ends, and 0.338 once the paper is out. The difference, 0.090 percentage points or 13.4% of the in-sample return, arrives only after publication.

That gap is doing a lot of work in this literature. It is the quantity McLean and Pontiff built their headline on, attributing 32 points of it to investors learning about mispricing from academic research. The logic is clean and I find it persuasive on its face. Decay in the post-sample window cannot be publication-driven, because the paper is not out. Anything on top of that has to be something publication caused.

The logic has a hole in it, and the hole is not about these anomalies. It is about the estimator.

The argument assumes that if publication did nothing, the two windows would look alike. That assumption holds only if returns are flat across the post-sample period. If instead they drift downward throughout, then the later window is lower than the earlier one for reasons that have nothing to do with the paper, and the difference between them gets read as arbitrage regardless.

In plain terms. The published method works like this: an anomaly's returns fall after its paper comes out, and the portion of that fall which appears only after publication gets labelled arbitrage. That reasoning holds if returns would otherwise have stayed flat. If they were already sliding, then the second half of any comparison comes in below the first, and the method reports arbitrage for a decline that was happening regardless. The whole of this post is one question: Which of those two is going on here?

This post tests that directly. The cleanest way to find out whether a date is doing any work is to replace it with a date that cannot be, and see whether the answer changes.


Data and Methodology

Same dataset and construction as Part 1. The October 2025 release of Open Source Asset Pricing, 212 predictors with monthly long-short returns from January 1926 to December 2024, built the way each original paper built them, with 210 surviving a requirement of at least 24 months in each regime. Regressions are pooled with standard errors clustered by calendar month.

The companion notebook is standalone. It rebuilds the regimes from scratch rather than depending on Part 1’s, so it runs on its own once the data cache exists.

One measurement note that matters throughout. The natural way to express decay is as a share of the in-sample return, but that ratio explodes when an anomaly’s in-sample return sits near zero, and a handful of those dominate any average. The unwinsorized mean post-publication ratio is −0.797. Winsorized at the 5th and 95th percentiles it is −0.576, and the median is −0.617. So raw differences in percentage points are the headline here and ratios are the readable secondary view. Where they disagree, trust the raw difference.


Results

The regression is Part 1's. In-sample is 0.671% a month, the post-sample window comes in 0.247 lower at t=−3.77, and post-publication comes in 0.338 lower at t=−6.56. Both are significant, and neither is the quantity in question.

The test that matters is not either coefficient. It is the difference between them, which is the part publication could be responsible for:

post-publication minus post-sample
coefficient  −0.0903
std error     0.0700
z            −1.30
p             0.195
95% interval [−0.227, +0.046]

Not significant, and the interval contains zero comfortably. But a p-value of 0.195 on its own is a weak thing to argue around. Plenty of real effects fail to clear significance with 210 units and a noisy dependent variable, and failing to reject the null is not the same result as claiming there is no conclusion to be drawn.

The better test is to ask what this estimator returns when the date it depends on is meaningless. Assign each anomaly a fake publication year drawn from the same distribution of sample-end-to-publication gaps the real data shows, re-cut the regimes, and re-run. Two hundred times.

observed gap              −0.0903
randomized gaps, mean     −0.0897
randomized gaps, sd        0.0274
share at least as negative  0.495

The real publication dates produce a coefficient of −0.0903. Fake ones produce an average of −0.0897. The observed value sits at the 49.5th percentile of the randomized distribution, which is as close to the middle as this test can put it. The estimator returns the same answer either way.

Keep reading with a 7-day free trial

Subscribe to Alpha in Academia to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Alpha in Academia · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture