A forward-tracked paper index of U.S. quality-cash-flow companies. Rules are published, frozen, and third-party timestamped before the record begins — so they cannot be tuned after the fact. It does not promise to beat anything. It promises to be reproducible, and never backfilled.
Methodology: METHODOLOGY.md (v1.0) — config.yaml is its executable copy.
Chinese version: METHODOLOGY.zh.md
It does not promise to outperform the S&P 500 or any other index. That is not its goal, and it makes no such claim.
It promises exactly two things:
- Reproducibility — anyone with this repository's methodology and the free, public SEC EDGAR API can compute identical results. Every data source, every threshold, and every known XBRL pitfall (with its handling) is documented.
- No backfilling — the rules were published and timestamped before inception and are never revised to flatter results. Losing periods are published the same as winning ones; losers are never deleted.
This is an index, not a trading record. Like every index, it excludes slippage, taxes, and liquidity costs.
Backtests are not the product. Any backtest published later must be prominently labeled as carrying look-ahead and survivorship bias.
Any stock-selection rule that can be written down has already been productized by Wall Street (COWZ, QUAL, and dozens more), and factor premia decay by roughly half after publication. So the scarce thing here is not the rules — those are public, and you are welcome to copy them.
What is scarce is the forward, tamper-evident ledger: rules nailed down in public before any result existed, and every day since recorded under third-party timestamps. A snapshot can be copied. The ledger cannot.
| Layer | What it does | The trap it blocks |
|---|---|---|
| L1 Universe | S&P 500 snapshot, excluding balance-sheet financials (banks / insurers / consumer finance) and REITs | The FCF–margin–ROIC frame is meaningless for them |
| L2 Landmines | Drop on: earnings far exceeding operating cash flow · receivables growing >2× revenue · share count rising despite buybacks · net debt/EBITDA too high | Accounting games and financial engineering |
| L3 Quality | Require all of: FCF positive ≥7 consecutive years with low variance · gross margin ≥30% and not declining over 10 years · 5-year average ROIC ≥12% · asset growth ≤ revenue growth | Cyclical peaks masquerading as cash cows; empire-building |
| L4 Valuation | FCF yield ≥ 10-year Treasury or P/E ≤ 30 | Paying too much for a good business |
| L5 Count | Top 20 by composite score | — |
Weighting — score-weighted at entry, capped at 8% per name; never rebalanced afterwards, so winners are allowed to drift upward.
Reconstitution — first trading day of January and July only. A rank buffer (enter at ≤20, exit only past 40) keeps turnover low.
The choice of N = 20 is derived, not arbitrary: diversification benefits saturate between 15 and 25 names; a mechanical screen has low information coefficient per name, so by IR ≈ IC × √breadth it should not be concentrated; but sector concentration caps effective breadth near 25. The two forces cross at 20–25.
- U.S. fundamentals — SEC EDGAR
companyfactsAPI (free, no key, full XBRL history from ~2009) - Prices / valuation — yfinance (price, market cap, trailing P/E only)
Only us-gaap tags and annual 10-K / 10-K/A values are used. Foreign private
issuers filing solely 20-F are structurally out of scope — their statements
cannot be reproduced by this pipeline.
Missing data is flagged data_incomplete and excluded. Never filled with zero,
estimates, or industry averages.
Six known EDGAR pitfalls — tag changes across accounting standards, magnitude shifts within a series, stock-split discontinuities, stale values after a company stops reporting a subtotal, single-step income statements, and currency mismatches — are documented with their handling in METHODOLOGY.md §1.2.
python -m src.run_screen # run the L1–L5 funnel; writes candidates + full rejection ledger
python -m src.build_portfolio # score-weighted constituents and weights
python -m src.open_books # one-time: open the ledger on the inception date
python -m src.daily_level # compute today's index level (exits safely before inception)Data health checks (run these after touching extraction or metric logic):
python tests/probe_edgar.py # raw annual series
python tests/probe_metrics.py # derived signals + normalization checksNothing in the daily record depends on a human remembering to run anything:
- Daily level — GitHub Actions (
daily.yml), 21:30 UTC on weekdays, shortly after the US close, with an idempotent backup trigger at 23:30 UTC (scheduled events are best-effort; a dropped cron must not cost the day): computes the level, appends it to the ledger, commits asdata:, and anchors a Wayback snapshot. The row's date comes from the price bars' own US-Eastern trading day, never from the runner's clock (ERRATA, 2026-08-06). - Freshness monitor — GitHub Actions (
monitor.yml), shortly after midnight New York plus a 13:00 UTC backup sweep: compares the ledger's last row against the last completed US trading day read off SPY's own bars — never a wall clock, never a workflow's self-report — and fails loudly the same night a trading day goes missing or a future-dated row appears. - There is no silent path: either the day's
data:commit lands, or a workflow run turns red and GitHub raises a failure notification. An additional off-repo review recomputes each day's level against an independent price source; discrepancies are published to ERRATA.md, never patched into history.
- Public repository; commit times are attested by GitHub as a third party.
- Every data update is snapshotted to the Internet Archive.
- Commit types are strictly separated:
methodology:= rule changes,data:= data updates. Anyone can walk the history and verify that the rules went untouched across any given period. - A separate freshness monitor fails loudly if the ledger stops updating — a silent gap in the record is the one thing that would undermine all of the above.
| Inception | 2026-07-20 |
| Base level | 100 |
| Constituents | 19 — one seat short of target N = 20 due to a dual-class defect at inception, restored at the January 2027 review (ERRATA.md) |
| Reconstitution | 1st trading day of January and July |
| Files | data/ledger/ — constituents.csv, index_level.csv |
| Errata | ERRATA.md — dated corrections; ledger rows are never rewritten |
PolyForm Noncommercial 1.0.0 — free for noncommercial use; commercial use requires a separate license.
Everything in this repository is for informational and research purposes only. It is not investment advice and recommends no security. Markets carry risk; make your own decisions.