Locating a failed equity strategy by elimination
A cross-sectional long and short book on large-cap US equities returns 6.2 percent a year at a Sharpe ratio of 1.78 when the tradable universe is today's index membership applied backwards over history. The same code, the same features, the same model and the same costs return 0.1 percent a year at a Sharpe ratio of 0.06 when membership is instead taken as it was known on each decision date. The difference is worth 1.73 of Sharpe ratio at a t statistic of 7.99 across 1,043 weeks.
So the apparent edge was an artefact of how the universe was assembled, which is a well documented failure mode rather than a surprise. The interesting part is what follows. Given a flat honest baseline, a real edge might still have been hiding in the breadth of the universe, in the holding period, in how the legs are formed, or in what the model is allowed to see. This report tests all four and finds one small effect, which is not enough to trade.
| Was the apparent edge real | Survivorship accounts for all of it. 1.73 of Sharpe, t 7.99 | artefact |
| Is the large cap universe too efficient | A universe twice as broad is worth about a point of Sharpe and still lands flat | refuted |
| Is the three week holding period wrong | Five periods from five to 126 days all land between 0.16 and 0.39 | refuted |
| Are the legs an accidental sector bet | Sector neutral legs produce a book 0.949 correlated with the original | refuted |
| Do point in time fundamentals add anything | Yes, about 1.3 percent a year at t 2.60, robust to the obvious choices | survived |
| Is that enough to trade | No. It improves a Sharpe ratio of −0.91 to −0.36 | insufficient |
Every equity figure this project produced before the work described here was measured on a universe of companies that still exist. That sounds innocuous and is not.
Roughly seventy two percent of US common stock securities that have ever listed are now delisted. In the data used here that is 11,584 of 16,096. A backtest that silently omits them is not slightly optimistic, it is long a portfolio selected on the outcome, because every omitted name is a company whose history ended in a way that a surviving company's did not. Brown and colleagues established the shape of this problem for fund performance in 1992, and Shumway quantified the specific damage that delisting returns do to CRSP based equity studies in 1997.1,2 The mechanism is old. What is easy to underestimate is the magnitude when the universe is an index rather than a fund sample.
The practical obstacle to doing this correctly is duller than the statistical one. Exchange
symbols are recycled. SHLD belonged to Sears Holdings until 2018 and now belongs
to a defence sector exchange traded fund. BBBY belonged to Bed Bath and Beyond
until 2023 and the successor entity trades under a different symbol entirely. Joining prices
to filings on a symbol therefore produces a history that is not noisy but fictitious, and a
symbol based universe silently inherits whoever holds the symbol today.
What is needed is a security identifier that is never reissued, which is what CRSP calls a PERMNO and the vendor used here calls a permaticker. Acquiring one was the only step in this project that could not be substituted with effort. A free listing calendar is available from SEC EDGAR and it does measure the attrition, but EDGAR discards symbols for companies that stop filing, and reconstructing symbols from filing metadata reached about sixty five percent accuracy, which makes a universe worse than useless.
In this data each security is stored under the symbol it died with. Sears Holdings is
SHLDQ, Bed Bath and Beyond is BBBYQ, SVB Financial is
SIVBQ, all carrying the suffix US exchanges append on bankruptcy. Across all
30,941 securities with price history, not one symbol is shared by two securities.
The consequence runs opposite to intuition. Bed Bath and Beyond's 1998 prices sit under a symbol that did not exist until 2023. A symbol here is a label, and not even the label the security traded under at the time, so matching external data on a symbol as of a date remains wrong for a new reason.
With that in hand the comparison is straightforward and the result is not subtle.
A flat result produced by freshly written code is more often a bug than a finding, so the honest arm needs a control. The control here is the biased arm itself. An earlier feasibility study in this project, run on a hand written list of 159 large caps that are liquid today, returned 5.7 percent a year at a Sharpe ratio of 1.66. The survivor arm in Figure 1 was built mechanically from index membership with no hand selection and returns 6.2 percent at 1.78.
The biased construction reproduces the old biased answer through a different code path and a different universe, which is what says the pipeline computes what it claims to. The point in time arm also ranks more names per day than the survivor arm, 494 against 421, so the gap is not a small sample artefact in the honest direction.
The strategy under test is ordinary and that is deliberate. Rank a liquid universe cross sectionally on price derived features, hold the top decile long and the bottom decile short, rebalance on a fixed horizon, pay ten basis points a leg. Momentum and low volatility in this form are among the most studied patterns in the literature, going back to Jegadeesh and Titman in 1993 and Ang and colleagues in 2006.3,4 If the honest version of that is flat, the first question is whether the specification is wrong somewhere rather than empty everywhere.
The S&P 500 is the most heavily arbitraged universe in existence, so perhaps the signal is fine and the names are simply too efficient.
Index membership was replaced with a point in time liquidity screen over every common stock security on disk. On each date the universe is the top 500 by trailing twenty one day dollar volume among names that actually had bars that day. The screen uses only information available at the close on which the decision is taken, and a security enters when it becomes liquid and leaves when its bars stop, with no reference to whether it survived.
The comparison is tight. Both arms trade 494 names a day. The broad arm draws them from 2,361 securities rather than 1,148, so the breadth is matched and the turnover in which names are held is more than doubled.
Every number so far was measured at a twenty one day hold, the one parameter nothing had varied.
Five holding periods from five to 126 days were run on the broad universe, with one panel labelled for all of them so nothing differs except the label and the holding period. The first pass used the fourteen year sample and produced an apparent lead at sixty three days, a Sharpe ratio of 0.72 against 0.07 at three weeks, which looked like the most promising result in the whole subproject.
It was not. The sixty three day figure rested on roughly eighty one independent holding periods, its t statistic against the three week base was 1.21, and it was the maximum of four specifications compared at once. Re running the sweep on the full twenty eight year price history, which raises the decision day count from 4,445 to 6,981, collapses it.
Ranking across a whole universe may load the long leg into one sector and the short leg into another, so the book wagers on sectors and any stock level signal is drowned by sector variance.
The legs were instead formed inside each sector, taking the top and bottom decile of every sector with at least twenty names and pooling the result, which makes the book sector neutral by construction. Sector is known for 3,547 of 3,551 securities and the remaining four were dropped from both arms so the universes stay matched.
Three structural explanations, each measured, each refuted. Taken together they do more than accumulate negative results. They say the failure is not localised in a parameter, which is the difference between a specification that needs tuning and one that is empty.
The remaining axis is not a parameter of the strategy but a question about what the model is allowed to see. The crypto half of this project had concluded, across four separate experiments, that the model class is not the binding constraint. Listwise ranking objectives, a three class target, a cross asset attention encoder and a zoo of sequence models given raw lookback windows all landed within noise of plain rank regression. The strongest form of that claim is that richer inputs would beat richer models, and equities offer an input crypto does not, which is quarterly accounting data tied to a permanent issuer identifier.
A fact enters the panel on the day it was filed and never on the day the period ended. Using the period end would hand the model a quarter of earnings weeks before the market saw it. The join is therefore an as of merge on filing date, and the measured median lag between period end and first public disclosure is thirty seven days.
Getting that right took two attempts. The SEC XBRL frames interface returns one observation per company and period taken from the most recent filing that mentions it, so a quarter's balance sheet arrives dated by a filing up to a year later, because subsequent reports repeat it as a comparative. That is conservative rather than forward looking, but a fundamental that arrives a year late carries almost no information, and the median lag under that approach was 383 days. Reading every filing of every fact and keeping the earliest brought it to the thirty seven above.
Fundamentals are keyed on an SEC central index key, and a central index key survives bankruptcy while the security does not. In this universe 214 of them cover 430 distinct securities, because a company reorganises and the successor inherits the registrant. American Airlines Group and AMR Corporation share one. US Airways has three.
Joining on the registrant alone would hand a live company the final balance sheet of its bankrupt predecessor, which is a forward looking leak in the worst available direction given that those filings are maximally distressed. Every fact is therefore attached to a security only if it was filed inside that security's own price window. Forty three securities have genuinely overlapping windows and cannot be separated by date at all, so the result is reported with and without them.
XBRL reporting phased in across 2009 to 2011, so machine readable fundamentals do not exist for companies that delisted before then. Of the index securities that have since delisted, twenty nine of the 249 that died between 1998 and 2008 have fundamentals at all, against eighty seven of eighty eight for 2012 to 2016 and fifty of fifty two for 2022 onward. Live securities are at 629 of 629.
This bounds the test rather than inconveniencing it. A fundamentals comparison run before roughly 2012 would drop dead names for want of filings and leave the fundamentals arm trading survivors while the price arm trades everything, which reintroduces the exact bias the point in time universe was bought to remove, through a different door. The test window therefore starts in 2012, and that is a property of the data rather than a choice.
| Arm | Return | Sharpe | Max drawdown | Hit rate |
|---|---|---|---|---|
| Price features only | −2.0% | −0.91 | −25.3% | 0.45 |
| Price and fundamentals | −0.8% | −0.36 | −15.6% | 0.51 |
The difference is 0.0241 percent a week at a t statistic of 2.60, which implies 1.26 percent a year against a directly measured gap of 1.24 percent. Four checks matter more than the point estimate. The arm with fundamentals wins 53.4 percent of weeks, so the effect is broad rather than carried by a handful of them. The ten largest weekly differences account for fifteen percent of the total, so it is not driven by outliers. Split in half the effect is present in both, at 0.0307 percent a week before 2019 and 0.0175 after. And excluding the forty three securities whose registrant windows overlap moves the effect from 0.0241 to 0.0231 percent a week.
The effect does not move when an arbitrary preprocessing choice changes, which is the first time anything in this subproject has passed that test.
That last point is the one worth dwelling on. An earlier version of this experiment, run before the universe was fixed, produced t statistics of 2.77 and −3.21 depending on which of three universes it used. Two significant results pointing in opposite directions from the same code is not evidence about fundamentals. It measures how much freedom the specification has. The present result is modest and it is stable, and the ordering of those two properties is the right way round.
So richer inputs do beat richer models here, which is the first result in either half of this project to improve an outcome through what the model sees rather than how it is shaped. The finding is consistent with a long line of work establishing accounting variables as cross sectional predictors, Novy-Marx on gross profitability being a direct example.8 It is also consistent with that literature's more recent and less comfortable finding, which is that published predictors decay after publication and many do not replicate at the strength originally claimed.5,6,7 An effect of 1.3 percent a year that weakens across the sample fits that picture rather well.
And 1.3 percent a year does not make a losing book profitable. Both arms lose money. Turning a Sharpe ratio of −0.91 into −0.36 is a real improvement to something that should not be traded.
Four axes, fifteen configurations, one small effect. The useful output of this study is not the flat line, it is the shape of the space around it.
A negative result is worth little when the specification has many free parameters, because the obvious reply is that the wrong corner was searched. That reply is available here only for corners nobody has reason to prefer. Breadth, holding period and leg construction are the three structural explanations anyone would reach for, and each was measured rather than argued about. Universe breadth is worth a point of Sharpe ratio and arrives flat. Holding period is worth nothing across a factor of twenty five in horizon. Leg construction produces a book 0.949 correlated with the original, which refutes the premise rather than merely failing to confirm it.
The remaining free parameter is the signal family, and that is where the honest uncertainty sits. Everything tested here derives from prices and from eight accounting ratios. The space of documented cross sectional predictors is far larger, and the open source replication effort by Chen and Zimmermann catalogues more than two hundred of them.12 Analyst revisions, short interest and insider transactions are the obvious untried families, and none of them is a variation on what was tested. Pursuing one would be a new study rather than another sweep of this one.
It licenses a narrow claim. Cross sectional ranking of liquid US equities on price derived features, traded market neutrally in deciles at ten basis points a leg, has no edge at any holding period from five to 126 days on any survivorship free universe tried, over twenty eight years. Point in time fundamentals add about 1.3 percent a year to that, robustly and insufficiently.
It does not license the claim that equities are unprofitable, which would be absurd. The specification tested is one of many, it is a simple one, and its simplicity was the point when the goal was a second uncorrelated book rather than a novel factor. What the study does say is that this particular road is closed, and it says where the barrier is not, which is worth more than the flat number on its own.
One incidental result is worth carrying forward. The sector neutral experiment recomputed its cross sectional ranks over 3,547 securities rather than 3,551, a difference of four names, and the control arm's Sharpe ratio moved by about 0.04. Every holding period measured in this study landed between 0.16 and 0.39, a band only a few multiples of that precision wide. None of these figures should be read to two decimal places, and the five horizon results in particular should be read as one flat band rather than a ranking.
The same caution applies with more force to the deflated Sharpe ratio framework of Bailey and López de Prado, and to the multiple testing correction argued for by Harvey and colleagues.10,6 Fifteen configurations were compared here. The maximum of fifteen draws from a distribution centred near zero is positive by construction, and the single largest honest Sharpe ratio in the study, 0.39 at a forty two day horizon, is comfortably inside what that selection alone would produce.
Daily split and dividend adjusted bars for 20,994 US securities from December 1997 to October 2026, keyed on a permanent security identifier, purchased from Sharadar under a personal use licence. Of 16,096 common stock securities on US exchanges, 11,584 are delisted. Index membership comes from 115 quarterly point in time snapshots of the S&P 500 from March 1998, covering 1,148 distinct securities. Fundamentals are first disclosure observations assembled from the SEC XBRL company facts interface for 894 companies, 303,322 facts, with a median filing lag of thirty seven days.
All four prices are scaled by one adjustment factor per bar. Swapping in an adjusted close while leaving the open, high and low unadjusted would compare an adjusted close against unadjusted extremes in the range features, which for any name with a split history means two price scales inside one ratio.
Walk forward with purging and an embargo of twice the holding period on either side of each training window. Within each training window the last fifteen percent is held out behind the same embargo and used only for early stopping. The model is a gradient boosted ranker on the within date percentile rank of the forward return, which is a monotone transform of the price ratio and therefore identical whether returns are expressed in logs or simply. Features are per security price statistics ranked cross sectionally within each decision date, which removes the market level and leaves the relative position.
Returns are net of ten basis points per leg round trip. Legs are combined as arithmetic returns rather than logarithmic ones. The distinction is immaterial for ordinary moves and decisive for a leg that delists, where a ninety nine percent loss is −4.6 in logs against −0.99 arithmetically, and it is not symmetric between the two sides. A short position's simple return is the negative of the exponential of its log figure less one, so the naive conversion credits a short one hundred percent when an asset halves instead of fifty.
The arithmetic point above was found here rather than known in advance. The point in time universe is the first thing in this project to hold securities that go to zero, and that exposed an aggregation error in shared backtest code which had been invisible across twenty eight prior experiments on cryptocurrency, where individual legs are volatile but rarely terminal. Correcting it changed every figure in the other half of this project.
It is recorded here because the finding belongs to the equity work even though the damage was elsewhere, and because it is a concrete instance of a general point. Holding assets that can die is a different arithmetic regime, not merely a wider distribution.
Every figure in this report is generated from the stored weekly return series by a script, so no figure can drift from the result it illustrates. The raw vendor data cannot be redistributed under the licence, so the derived series, the counts and the test statistics are recorded in the project's research log rather than left implicit in a data file.