Report two · equities · survivorship and elimination

HydraTrade II. The Universe Was the Result

Locating a failed equity strategy by elimination

An engraving of a long colonnade receding into haze. The
      columns on the right stand intact with terracotta capitals that trace a rising line.
      Most of the colonnade, to the left, is fallen and drawn so faintly it almost
      disappears. A single small figure stands at the near end, facing only the part that
      is still standing.
Generated with OpenAI Image Generator

A cross-sectional long and short book on large-cap US equities returns 6.2 percent a year at a Sharpe ratio of 1.78 when the tradable universe is today's index membership applied backwards over history. The same code, the same features, the same model and the same costs return 0.1 percent a year at a Sharpe ratio of 0.06 when membership is instead taken as it was known on each decision date. The difference is worth 1.73 of Sharpe ratio at a t statistic of 7.99 across 1,043 weeks.

So the apparent edge was an artefact of how the universe was assembled, which is a well documented failure mode rather than a surprise. The interesting part is what follows. Given a flat honest baseline, a real edge might still have been hiding in the breadth of the universe, in the holding period, in how the legs are formed, or in what the model is allowed to see. This report tests all four and finds one small effect, which is not enough to trade.

Was the apparent edge real Survivorship accounts for all of it. 1.73 of Sharpe, t 7.99 artefact
Is the large cap universe too efficient A universe twice as broad is worth about a point of Sharpe and still lands flat refuted
Is the three week holding period wrong Five periods from five to 126 days all land between 0.16 and 0.39 refuted
Are the legs an accidental sector bet Sector neutral legs produce a book 0.949 correlated with the original refuted
Do point in time fundamentals add anything Yes, about 1.3 percent a year at t 2.60, robust to the obvious choices survived
Is that enough to trade No. It improves a Sharpe ratio of −0.91 to −0.36 insufficient
Part one

Why none of the earlier numbers counted

Every equity figure this project produced before the work described here was measured on a universe of companies that still exist. That sounds innocuous and is not.

Roughly seventy two percent of US common stock securities that have ever listed are now delisted. In the data used here that is 11,584 of 16,096. A backtest that silently omits them is not slightly optimistic, it is long a portfolio selected on the outcome, because every omitted name is a company whose history ended in a way that a surviving company's did not. Brown and colleagues established the shape of this problem for fund performance in 1992, and Shumway quantified the specific damage that delisting returns do to CRSP based equity studies in 1997.1,2 The mechanism is old. What is easy to underestimate is the magnitude when the universe is an index rather than a fund sample.

Ticker symbols are not identifiers

The practical obstacle to doing this correctly is duller than the statistical one. Exchange symbols are recycled. SHLD belonged to Sears Holdings until 2018 and now belongs to a defence sector exchange traded fund. BBBY belonged to Bed Bath and Beyond until 2023 and the successor entity trades under a different symbol entirely. Joining prices to filings on a symbol therefore produces a history that is not noisy but fictitious, and a symbol based universe silently inherits whoever holds the symbol today.

What is needed is a security identifier that is never reissued, which is what CRSP calls a PERMNO and the vendor used here calls a permaticker. Acquiring one was the only step in this project that could not be substituted with effort. A free listing calendar is available from SEC EDGAR and it does measure the attrition, but EDGAR discards symbols for companies that stop filing, and reconstructing symbols from filing metadata reached about sixty five percent accuracy, which makes a universe worse than useless.

A detail worth stating, because it inverts the obvious rule

In this data each security is stored under the symbol it died with. Sears Holdings is SHLDQ, Bed Bath and Beyond is BBBYQ, SVB Financial is SIVBQ, all carrying the suffix US exchanges append on bankruptcy. Across all 30,941 securities with price history, not one symbol is shared by two securities.

The consequence runs opposite to intuition. Bed Bath and Beyond's 1998 prices sit under a symbol that did not exist until 2023. A symbol here is a label, and not even the label the security traded under at the time, so matching external data on a symbol as of a date remains wrong for a new reason.

With that in hand the comparison is straightforward and the result is not subtle.

0.5× 1× 2× 4× point-in-time members 1.03× today's members, applied backwards 3.37× 2008 2011 2014 2017 2020 2023 2026
Figure 1. Cumulative growth per unit of capital, weekly, net of ten basis points per leg, over twenty walk forward windows and 1,043 weeks. The two lines run the same ranker over the same securities with the same features. The only difference is which names are eligible on each decision date. Taking membership as it was known leaves the strategy flat. Taking today's membership and applying it backwards produces an edge that was never available.

Why a null result from new code deserves suspicion, and why this one survives it

A flat result produced by freshly written code is more often a bug than a finding, so the honest arm needs a control. The control here is the biased arm itself. An earlier feasibility study in this project, run on a hand written list of 159 large caps that are liquid today, returned 5.7 percent a year at a Sharpe ratio of 1.66. The survivor arm in Figure 1 was built mechanically from index membership with no hand selection and returns 6.2 percent at 1.78.

The biased construction reproduces the old biased answer through a different code path and a different universe, which is what says the pipeline computes what it claims to. The point in time arm also ranks more names per day than the survivor arm, 494 against 421, so the gap is not a small sample artefact in the honest direction.

Part two

Three places a real edge could have been hiding

The strategy under test is ordinary and that is deliberate. Rank a liquid universe cross sectionally on price derived features, hold the top decile long and the bottom decile short, rebalance on a fixed horizon, pay ten basis points a leg. Momentum and low volatility in this form are among the most studied patterns in the literature, going back to Jegadeesh and Titman in 1993 and Ang and colleagues in 2006.3,4 If the honest version of that is flat, the first question is whether the specification is wrong somewhere rather than empty everywhere.

1Breadth

The S&P 500 is the most heavily arbitraged universe in existence, so perhaps the signal is fine and the names are simply too efficient.

Index membership was replaced with a point in time liquidity screen over every common stock security on disk. On each date the universe is the top 500 by trailing twenty one day dollar volume among names that actually had bars that day. The screen uses only information available at the close on which the decision is taken, and a security enters when it becomes liquid and leaves when its bars stop, with no reference to whether it survived.

The comparison is tight. Both arms trade 494 names a day. The broad arm draws them from 2,361 securities rather than 1,148, so the breadth is matched and the turnover in which names are held is more than doubled.

Refuted, with a little truth in it. Widening the universe is worth about a point of Sharpe ratio, from −0.91 on the index to 0.09 on the broad screen over the same 730 weeks. Small capitalisation names are where cross sectional anomalies have historically been strongest, so a real improvement is the expected sign. It arrives at flat.

2Holding period

Every number so far was measured at a twenty one day hold, the one parameter nothing had varied.

Five holding periods from five to 126 days were run on the broad universe, with one panel labelled for all of them so nothing differs except the label and the holding period. The first pass used the fourteen year sample and produced an apparent lead at sixty three days, a Sharpe ratio of 0.72 against 0.07 at three weeks, which looked like the most promising result in the whole subproject.

It was not. The sixty three day figure rested on roughly eighty one independent holding periods, its t statistic against the three week base was 1.21, and it was the maximum of four specifications compared at once. Re running the sweep on the full twenty eight year price history, which raises the decision day count from 4,445 to 6,981, collapses it.

Refuted. On the full sample every horizon from five to 126 days lands between 0.16 and 0.39, every absolute t statistic against the base is below 0.7, and the nominal best moves from sixty three days to forty two. A Sharpe ratio of 0.72 becomes 0.25. That is a ranking over noise being resampled.
0 0.25 0.5 0.75 5d +0.10 10d +0.05 21d +0.07 63d +0.72 2012 sample, 730 weeks 21d +0.23 42d +0.39 63d +0.25 84d +0.16 126d +0.26 full sample, 1,022 weeks 0.72 0.25 5d 10d 21d 42d 63d 84d 126d holding period
Figure 2. Sharpe ratio by holding period on the shorter sample and on the full one. The dashed line is the fourteen year window that produced the apparent quarterly lead. The solid line is twenty eight years of the same thing. Long horizons give up independent observations quickly, so a wide gap with a small t statistic is exactly what should be expected here, and the shorter sample's ordering does not survive.

3Construction

Ranking across a whole universe may load the long leg into one sector and the short leg into another, so the book wagers on sectors and any stock level signal is drowned by sector variance.

The legs were instead formed inside each sector, taking the top and bottom decile of every sector with at least twenty names and pooling the result, which makes the book sector neutral by construction. Sector is known for 3,547 of 3,551 securities and the remaining four were dropped from both arms so the universes stay matched.

Refuted, and the premise was false rather than the effect null. Sector neutral construction gives a Sharpe ratio of 0.19 against 0.20, a difference of −0.0055 percent a week at a t statistic of −0.72. The informative number is the correlation between the two books, which is 0.949. Forming legs inside sectors produces very nearly the same book, so the global ranking was never concentrating into sectors and there was no sector bet available to remove.
global deciles, weekly sector neutral, weekly correlation 0.949 the two books are the same book, so there was no sector bet to remove
Figure 3. Weekly return of the sector neutral book against the global book, one point per week over 1,044 weeks, with the identity line shown. A sector bet worth removing would appear as dispersion away from the diagonal. There is very little. Sector neutrality does buy three points of maximum drawdown, from −30.4 percent to −27.1 percent, at no cost in return, which is worth knowing if such a variant is ever wanted for other reasons.

Three structural explanations, each measured, each refuted. Taken together they do more than accumulate negative results. They say the failure is not localised in a parameter, which is the difference between a specification that needs tuning and one that is empty.

-1 +0 +1 +2 every honest configuration universe today's members, backwards +1.78 +1.78 point-in-time index, 2006 on +0.06 point-in-time index, 2012 on -0.91 -0.91 broad liquid top 500 +0.09 horizon 5 days +0.10 10 days +0.05 21 days +0.23 42 days +0.39 63 days +0.25 84 days +0.16 126 days +0.26 construction global deciles +0.20 sector neutral +0.19 inputs price only -0.91 -0.91 price and fundamentals -0.36
Figure 4. Every configuration measured in this study, grouped by the axis varied. Fifteen points. The shaded region spans every honest configuration, which runs from −0.91 to 0.39 and includes nothing tradable. The single point above it is the universe built from today's index membership. The only construction in the study that produces an edge is the one that could not have been traded.
Part three

The one thing that worked

The remaining axis is not a parameter of the strategy but a question about what the model is allowed to see. The crypto half of this project had concluded, across four separate experiments, that the model class is not the binding constraint. Listwise ranking objectives, a three class target, a cross asset attention encoder and a zoo of sequence models given raw lookback windows all landed within noise of plain rank regression. The strongest form of that claim is that richer inputs would beat richer models, and equities offer an input crypto does not, which is quarterly accounting data tied to a permanent issuer identifier.

Doing it without looking ahead

A fact enters the panel on the day it was filed and never on the day the period ended. Using the period end would hand the model a quarter of earnings weeks before the market saw it. The join is therefore an as of merge on filing date, and the measured median lag between period end and first public disclosure is thirty seven days.

Getting that right took two attempts. The SEC XBRL frames interface returns one observation per company and period taken from the most recent filing that mentions it, so a quarter's balance sheet arrives dated by a filing up to a year later, because subsequent reports repeat it as a comparative. That is conservative rather than forward looking, but a fundamental that arrives a year late carries almost no information, and the median lag under that approach was 383 days. Reading every filing of every fact and keeping the earliest brought it to the thirty seven above.

A registrant identifier is not a security identifier either

Fundamentals are keyed on an SEC central index key, and a central index key survives bankruptcy while the security does not. In this universe 214 of them cover 430 distinct securities, because a company reorganises and the successor inherits the registrant. American Airlines Group and AMR Corporation share one. US Airways has three.

Joining on the registrant alone would hand a live company the final balance sheet of its bankrupt predecessor, which is a forward looking leak in the worst available direction given that those filings are maximally distressed. Every fact is therefore attached to a security only if it was filed inside that security's own price window. Forty three securities have genuinely overlapping windows and cannot be separated by date at all, so the result is reported with and without them.

What the data allows, which is less than one would like

XBRL reporting phased in across 2009 to 2011, so machine readable fundamentals do not exist for companies that delisted before then. Of the index securities that have since delisted, twenty nine of the 249 that died between 1998 and 2008 have fundamentals at all, against eighty seven of eighty eight for 2012 to 2016 and fifty of fifty two for 2022 onward. Live securities are at 629 of 629.

This bounds the test rather than inconveniencing it. A fundamentals comparison run before roughly 2012 would drop dead names for want of filings and leave the fundamentals arm trading survivors while the price arm trades everything, which reintroduces the exact bias the point in time universe was bought to remove, through a different door. The test window therefore starts in 2012, and that is a property of the data rather than a choice.

Fundamentals against a price only baseline on the point in time index universe, 730 weeks from October 2012, fourteen walk forward windows. Both arms trade the identical universe of 1,106 securities and 3.45 million rows. Adding fundamentals adds eight scale free features to twenty one price features and removes no rows, so a security without filings carries missing values and the tree handles them.
ArmReturnSharpe Max drawdownHit rate
Price features only−2.0%−0.91 −25.3%0.45
Price and fundamentals−0.8% −0.36−15.6%0.51

The difference is 0.0241 percent a week at a t statistic of 2.60, which implies 1.26 percent a year against a directly measured gap of 1.24 percent. Four checks matter more than the point estimate. The arm with fundamentals wins 53.4 percent of weeks, so the effect is broad rather than carried by a handful of them. The ten largest weekly differences account for fifteen percent of the total, so it is not driven by outliers. Split in half the effect is present in both, at 0.0307 percent a week before 2019 and 0.0175 after. And excluding the forty three securities whose registrant windows overlap moves the effect from 0.0241 to 0.0231 percent a week.

The effect does not move when an arbitrary preprocessing choice changes, which is the first time anything in this subproject has passed that test.

That last point is the one worth dwelling on. An earlier version of this experiment, run before the universe was fixed, produced t statistics of 2.77 and −3.21 depending on which of three universes it used. Two significant results pointing in opposite directions from the same code is not evidence about fundamentals. It measures how much freedom the specification has. The present result is modest and it is stable, and the ordering of those two properties is the right way round.

0% 10% 20% 30% all 1,106 securities 13 ambiguous excluded t +2.40 before, t +1.31 after 2013 2016 2019 2022 2025
Figure 5. Cumulative difference between the arm with fundamentals and the arm without, which is the only quantity this experiment claims. Both variants are shown, with and without the securities whose registrant windows overlap. The accumulation is gradual rather than stepwise, and it flattens in the second half of the sample, where the t statistic falls from 2.40 to 1.31.

So richer inputs do beat richer models here, which is the first result in either half of this project to improve an outcome through what the model sees rather than how it is shaped. The finding is consistent with a long line of work establishing accounting variables as cross sectional predictors, Novy-Marx on gross profitability being a direct example.8 It is also consistent with that literature's more recent and less comfortable finding, which is that published predictors decay after publication and many do not replicate at the strength originally claimed.5,6,7 An effect of 1.3 percent a year that weakens across the sample fits that picture rather well.

And 1.3 percent a year does not make a losing book profitable. Both arms lose money. Turning a Sharpe ratio of −0.91 into −0.36 is a real improvement to something that should not be traded.

Part four

Reading the elimination

Four axes, fifteen configurations, one small effect. The useful output of this study is not the flat line, it is the shape of the space around it.

A negative result is worth little when the specification has many free parameters, because the obvious reply is that the wrong corner was searched. That reply is available here only for corners nobody has reason to prefer. Breadth, holding period and leg construction are the three structural explanations anyone would reach for, and each was measured rather than argued about. Universe breadth is worth a point of Sharpe ratio and arrives flat. Holding period is worth nothing across a factor of twenty five in horizon. Leg construction produces a book 0.949 correlated with the original, which refutes the premise rather than merely failing to confirm it.

The remaining free parameter is the signal family, and that is where the honest uncertainty sits. Everything tested here derives from prices and from eight accounting ratios. The space of documented cross sectional predictors is far larger, and the open source replication effort by Chen and Zimmermann catalogues more than two hundred of them.12 Analyst revisions, short interest and insider transactions are the obvious untried families, and none of them is a variation on what was tested. Pursuing one would be a new study rather than another sweep of this one.

What this does and does not license

It licenses a narrow claim. Cross sectional ranking of liquid US equities on price derived features, traded market neutrally in deciles at ten basis points a leg, has no edge at any holding period from five to 126 days on any survivorship free universe tried, over twenty eight years. Point in time fundamentals add about 1.3 percent a year to that, robustly and insufficiently.

It does not license the claim that equities are unprofitable, which would be absurd. The specification tested is one of many, it is a simple one, and its simplicity was the point when the goal was a second uncorrelated book rather than a novel factor. What the study does say is that this particular road is closed, and it says where the barrier is not, which is worth more than the flat number on its own.

A note on what the measurement precision will bear

One incidental result is worth carrying forward. The sector neutral experiment recomputed its cross sectional ranks over 3,547 securities rather than 3,551, a difference of four names, and the control arm's Sharpe ratio moved by about 0.04. Every holding period measured in this study landed between 0.16 and 0.39, a band only a few multiples of that precision wide. None of these figures should be read to two decimal places, and the five horizon results in particular should be read as one flat band rather than a ranking.

The same caution applies with more force to the deflated Sharpe ratio framework of Bailey and López de Prado, and to the multiple testing correction argued for by Harvey and colleagues.10,6 Fifteen configurations were compared here. The maximum of fifteen draws from a distribution centred near zero is positive by construction, and the single largest honest Sharpe ratio in the study, 0.39 at a forty two day horizon, is comfortably inside what that selection alone would produce.

Methods

How the numbers were produced

Data

Daily split and dividend adjusted bars for 20,994 US securities from December 1997 to October 2026, keyed on a permanent security identifier, purchased from Sharadar under a personal use licence. Of 16,096 common stock securities on US exchanges, 11,584 are delisted. Index membership comes from 115 quarterly point in time snapshots of the S&P 500 from March 1998, covering 1,148 distinct securities. Fundamentals are first disclosure observations assembled from the SEC XBRL company facts interface for 894 companies, 303,322 facts, with a median filing lag of thirty seven days.

All four prices are scaled by one adjustment factor per bar. Swapping in an adjusted close while leaving the open, high and low unadjusted would compare an adjusted close against unadjusted extremes in the range features, which for any name with a split history means two price scales inside one ratio.

Protocol

Walk forward with purging and an embargo of twice the holding period on either side of each training window. Within each training window the last fifteen percent is held out behind the same embargo and used only for early stopping. The model is a gradient boosted ranker on the within date percentile rank of the forward return, which is a monotone transform of the price ratio and therefore identical whether returns are expressed in logs or simply. Features are per security price statistics ranked cross sectionally within each decision date, which removes the market level and leaves the relative position.

Returns are net of ten basis points per leg round trip. Legs are combined as arithmetic returns rather than logarithmic ones. The distinction is immaterial for ordinary moves and decisive for a leg that delists, where a ninety nine percent loss is −4.6 in logs against −0.99 arithmetically, and it is not symmetric between the two sides. A short position's simple return is the negative of the exponential of its log figure less one, so the naive conversion credits a short one hundred percent when an asset halves instead of fifty.

Provenance of that last paragraph

The arithmetic point above was found here rather than known in advance. The point in time universe is the first thing in this project to hold securities that go to zero, and that exposed an aggregation error in shared backtest code which had been invisible across twenty eight prior experiments on cryptocurrency, where individual legs are volatile but rarely terminal. Correcting it changed every figure in the other half of this project.

It is recorded here because the finding belongs to the equity work even though the damage was elsewhere, and because it is a concrete instance of a general point. Holding assets that can die is a different arithmetic regime, not merely a wider distribution.

Reproducibility

Every figure in this report is generated from the stored weekly return series by a script, so no figure can drift from the result it illustrates. The raw vendor data cannot be redistributed under the licence, so the derived series, the counts and the test statistics are recorded in the project's research log rather than left implicit in a data file.

Register

Every experiment, and what it showed

  1. Point in time universe from SEC EDGAR
    A free listing calendar for 11,514 companies keyed on registrant, and it does measure the attrition, with forty seven percent of companies filing in the first half of 2019 having since stopped. Unusable as a universe because EDGAR discards symbols for companies that stop filing.
  2. Recovering symbols from filing metadata
    About sixty five percent accurate. The leading token of a filing name is often the filing agent rather than the company. A third of symbols wrong makes a universe fictitious rather than noisy, so this was abandoned rather than tuned.
  3. Acceptance test on the purchased data
    Three failures specified in advance. All three pass, though not as expected. Delisted names are stored under bankruptcy suffixed symbols, so the first query returned nothing for four of five test cases while the data was complete.
  4. Survivorship, measured
    Twenty windows, 1,043 weeks. Point in time membership gives 0.1 percent a year at a Sharpe ratio of 0.06. Today's membership applied backwards gives 6.2 percent at 1.78. Overstatement t 7.99, Sharpe inflated by 1.73.
  5. Fundamentals on the honest universe
    Fourteen windows, 730 weeks. Price only −2.0 percent at −0.91, price and fundamentals −0.8 percent at −0.36. Effect 0.0241 percent a week at t 2.60, and 0.0231 at t 2.05 excluding ambiguous registrants.
  6. Breadth beyond the index
    Point in time liquidity screen over all 16,096 securities, 2,361 of them ever eligible, same 494 names a day. Returns 0.3 percent at a Sharpe ratio of 0.09 against the index's −2.0 percent at −0.91.
  7. Holding period, short sample
    Four horizons on fourteen years. An apparent lead at sixty three days, 0.72 against 0.07, at t 1.21 and as the maximum of four. Flagged as a lead requiring a dedicated test rather than reported as a result.
  8. Holding period, full sample
    Five horizons on twenty eight years and 1,022 weeks. Everything from five to 126 days lands between 0.16 and 0.39, all absolute t statistics below 0.7, nominal best moves to forty two days. The 0.72 becomes 0.25.
  9. Sector neutral construction
    Legs formed inside each sector with at least twenty names. Sharpe ratio 0.19 against 0.20 at t −0.72, and the two books correlate 0.949. Buys three points of drawdown at no cost in return.
References
  1. Brown, S. J., Goetzmann, W., Ibbotson, R. G., and Ross, S. A. (1992). Survivorship bias in performance studies. Review of Financial Studies, 5(4), 553–580.
  2. Shumway, T. (1997). The delisting bias in CRSP data. Journal of Finance, 52(1), 327–340.
  3. Jegadeesh, N., and Titman, S. (1993). Returns to buying winners and selling losers. Implications for stock market efficiency. Journal of Finance, 48(1), 65–91.
  4. Ang, A., Hodrick, R. J., Xing, Y., and Zhang, X. (2006). The cross section of volatility and expected returns. Journal of Finance, 61(1), 259–299.
  5. McLean, R. D., and Pontiff, J. (2016). Does academic research destroy stock return predictability? Journal of Finance, 71(1), 5–32.
  6. Harvey, C. R., Liu, Y., and Zhu, H. (2016). And the cross section of expected returns. Review of Financial Studies, 29(1), 5–68.
  7. Hou, K., Xue, C., and Zhang, L. (2020). Replicating anomalies. Review of Financial Studies, 33(5), 2019–2133.
  8. Novy-Marx, R. (2013). The other side of value. The gross profitability premium. Journal of Financial Economics, 108(1), 1–28.
  9. Novy-Marx, R., and Velikov, M. (2016). A taxonomy of anomalies and their trading costs. Review of Financial Studies, 29(1), 104–147.
  10. Bailey, D. H., and López de Prado, M. (2014). The deflated Sharpe ratio. Correcting for selection bias, backtest overfitting and non normality. Journal of Portfolio Management, 40(5), 94–107.
  11. Linnainmaa, J. T., and Roberts, M. R. (2018). The history of the cross section of stock returns. Review of Financial Studies, 31(7), 2606–2649.
  12. Chen, A. Y., and Zimmermann, T. (2022). Open source cross sectional asset pricing. Critical Finance Review, 11(2), 207–264.
HydraTrade equities research report, October 2026