The data spine: research-grade market data for one venue
Live WebSocket collectors for what cannot be back-filled, S3 archive backfills for what can, labelled derivations for what must be reconstructed, and the storage contract born from the day a disk-pressure cleanup deleted the model weights a promotion needed: if a consumer might ever need it, it lives in the database — local files are cache, never contract.
Grid A/B/C: a promotion funnel that assumes you're fooling yourself
Our current three-gate protocol: CPCV signal screen, execution-cost sweep on locked predictions, and a 15-seed causal walk-forward with pre-registered distributional gates. Seed medians, never best seeds; multiplicity charged, never free; ties settled by paper trading, never by a third sweep. Built directly on the wreckage of our own selection-bias audit.
GRID-50: a pooled cross-sectional ranker on fifty coins
The first family through every gate: universe widened 30 → 51 perpetuals, one pre-registered LightGBM lambdarank family, 15/15 seeds positive on the causal walk-forward (+0.8–1.3 SR at 4 bps). Includes the finding that didn't cooperate — the construction ranking chosen on CPCV inverted on the causal panel, so two K-of-N books now trade side by side in paper to settle it.
Backtest → paper → live: where the edge evaporates
Three parallel ledgers — synthetic paper, live model-mirror, real fills — decompose the decay chain for our first live strategies. The verdict inverts the folk wisdom: execution matched the model to within a dollar; the damage was done at the promotion decision, which selected a mean-reverting hot streak at its peak. Why we retired directional single-coin bets from live.
An operating system for agent-driven research
Three months of LLM agents executing this program under written contracts: depth floors, numerical-plausibility flags, one canonical implementation per metric. Scored honestly — two bugs were caught by a human before the rules existed; the rules are codified hindsight. And the loop we chose not to close: agents propose, paper adjudicates, only a human promotes to live.
What we actually read: a grounded survey
The architecture families this program is built on — patch-attention transformers, MLP mixers, modern TCNs, gradient-boosted trees, LOB-specialised networks — and the primary literature each rests on. Where a paper's claims did not replicate on Hyperliquid data, we note it.
A funnel through 28 model families
Every architecture family in the modern time-series literature run through a single audited protocol on a 26-coin Hyperliquid panel: 116 hyperparameter cells, identical OOS windows, identical cost model. No favourites. Most fail — that is the point. The few that survive go to the next stage.
What transferred and what didn't: architecture lessons
Literature top-tier models (ModernTCN, PatchTST, TimesNet) do not transfer cleanly to short-horizon crypto perps. The strongest standalone cross-sectional signal on the held-out window came from a tuned GRU and a tabular ranker. We explain why sequence models struggle and what the tabular models see that they don't.
The full validation pipeline: Pass A, Pass B, and walk-forward
The program's first validation protocol: dual out-of-sample passes plus a held-out walk-forward block. Kept published as the historical record — a June 2026 audit found it still admitted selection on noise (best-seed picking, gate shopping), and it has been superseded by the pre-registered Grid A/B/C funnel described in Paper № 11.
Information coefficient is not Sharpe
IC measures rank correlation between predicted and realised returns — it says nothing about dollar PnL or trading costs. We published a +7.38 headline that became −3.53 once the pipeline was corrected to use raw forward returns at 4 bps per leg. The IC had looked great throughout. The bug, the fix, and why IC and dollar Sharpe must never be conflated.
Excess above noise: the signal quality floor
A positive IC is not enough — it must exceed its own standard error. We define excess IC as |IC| − SE(IC), where SE is estimated from the per-bar IC time series. Cells with excess IC ≤ 0 are statistically indistinguishable from noise regardless of their headline IC. Applying this floor as a promotion gate cut roughly half of otherwise-passing cells and meaningfully improved the quality of what reached capital.
Five ways Sharpe gets warped — and how we catch them
A Sharpe ratio is only as trustworthy as its inputs. We document five bugs that produced plausible-looking numbers on real data: future-leaking centring (one line of code turned −2 SR into +2 SR); cross-sectional rank targets passed as raw returns (+7.4 became −3.5); h-bar overlap in the cost function (+8 SR vanished); a mislabelled per-coin panel; and non-reproducible training. For each: what the bug looked like, who or what caught it, and the mechanical rule that now prevents it.
Cost-aware backtesting: the 4 bps convention
Why every figure on this site is quoted net of 4 basis points per leg, how that convention was chosen, why doubling it “to be conservative” double-counts, and worked examples of signals whose positive gross IC becomes a deeply negative net Sharpe once realistic turnover meets realistic cost.
Honest centring: one line of look-ahead, four points of Sharpe
The full case study behind Bug 1 of Paper № 07: a rolling-median centring window that included the bar being scored turned a −2 SR strategy into a +2 SR backtest. The causal centring contract that replaced it — trailing window, explicit warmup, shared code path between backtest and live scorer.
Debunked: regime-routed experts don't beat always-on
We built a three-state market-regime detector and routed validated per-coin specialists through it — a mixture-of-experts architecture. Out-of-sample it destroyed approximately 4.4 Sharpe versus simply leaving every specialist always on. The adaptive variant that looked compelling in-sample tripped every plausibility flag we had. We did not build the live router.
Debunked: stacked per-coin specialists
Training a dedicated model per coin and stacking their outputs into a portfolio signal looks appealing — more data per coin, no cross-sectional blending noise. In practice: sample starvation per coin, selection multiplicity across the grid, and correlated drawdowns exactly when diversification is needed most. Debunked head-to-head by the pooled cross-sectional ranker (Paper № 12). The earlier “ensemble beats every single arm” claim from Phase 1.7 was retired unconfirmed for the same reasons.
STRAT-04b: the first live LGBM-rank deployment
The write-up of our first real-money strategy, preserved as published. Its lifecycle — promotion, live decay, retirement — is exactly the arc decomposed in Paper № 13, and the reason live capital now runs market-neutral books only.
STRAT-48: the consensus specialist book
The multi-specialist consensus construction from the Phase 1.7 era, preserved as published. Superseded by the pooled cross-sectional K-of-N books of Paper № 12.
All headline Sharpe ratios on this site are annualised dollar Sharpe computed on raw forward returns at a uniform 4 basis points per leg cost convention, the operational estimate for our intended size on Hyperliquid. Information coefficient and rank-IR are reported separately and never substituted. Backtest performance is not a forecast of live results.