This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
This is ctbk.dev - a data pipeline and visualization dashboard for NYC Citi Bike trip data. The project combines:
- Python CLI (
ctbk) for ETL data processing - Next.js web dashboard at ctbk.dev
- Automated data ingestion via GitHub Actions
- DVC (Data Version Control) for data versioning with S3 backend
s3://tripdata (.csv.zip) → norm → cons ─┬→ agg (histograms: e_c, se_c, ymrgtb_cd, ymrgtbs_cd, ymrgtbe_cd)
├→ smh → sm → spj
└→ station-trips-json (per-station ymdgtb JSONs)
The monthly driver is ctbk update -S <YYYYMM> (ctbk/update.py) — norm reads the s3://tripdata .csv.zips directly (the old csv extract stage is orphaned). Stage outputs are content-addressed in the DVX cache (s3://ctbk/.dvc/files/md5/…, migrating to R2 behind data.ctbk.dev; see specs/s3-to-r2-migration.md).
Two serving stacks sit on top of these outputs:
- Rides (homepage +
/s/:slugcharts): the pyrmts rollup-pyramid —normalized/*.parquettiles built on AWS Batch, registered in Cloudflare D1, served by the CF api worker at/api/rides(rides/{start,end}: raws:<id>station leaves + materializedc:<canonical>rollups, canonical by default,?raw=1audit; content-hashed shard keys; monthlyctbk gbfs rides-extend—specs/rides-rekey.md). Superseded the legacyymrgtb_cd.json/ per-station-JSON flow (still reachable via?tsrc=legacy). The pyramid rebuild runs separately fromctbk update(R2-only writes). - Availability: the GBFS subsystem under
gbfs/(see below).
The pipeline processes raw Citi Bike .csv.zip files through multiple stages:
- Normalization: Read
.csv.zips straight froms3://tripdata; merge NYC/JC regions, harmonize columns, split by (source, start, end) months - Consolidation: Combine all records ending in a given month into a single parquet
- Aggregation: Generate histograms by various dimensions (time, station, user type, etc.)
- Station metadata: Compute canonical station info and ride counts between station pairs
The s3/ctbk/normalized/ directory contains two types of DVC-tracked outputs per month:
-
YYYYMM/(directory): Output ofctbk norm create, contains parquet files split by source month- Example:
202006/202005_202006.parquet= rides from May 2020 tripdata that ended in June 2020 - Tracked by
YYYYMM.dvc
- Example:
-
YYYYMM.parquet(single file): Output ofctbk cons create, the canonical consolidated month- Combines all records from any normalized directory that end in this month
- Tracked by
YYYYMM.parquet.dvc
Special case: Months 202001-202101 have additional "v0" input data (normalized/v0/) used for backfilling older columns (Gender, Birth Year, Bike ID) that were removed in 202102
/ctbk/- Python package with CLI and data processing logic/www/- Next.js frontend dashboard/s3/- Local mirror of S3 data structure with DVC tracking/nbs/- Jupyter notebooks for analysis/gbfs/- Real-time station availability subsystem (orthogonal to trips ETL)worker/— CFW cron* * * * *, polls GBFSstation_status.json→ R2 WAL JSONsloader/— CFW R2-event queue consumer, ingests WAL → D1 hot-cache (last 7 days)api/— CFW serving/api/stations/*,/api/query,/api/totals,/api/ridescompact-r2.py— GHA daily compaction: WAL JSONs → daily/per-station parquet on R2- See
docs/pipeline.mdGBFS section +specs/gbfs-r2-only.md(migration in flight)
pip install -e .This installs the ctbk command with CLI entry points.
# Process a new month of data (using update.sh)
./update.sh 202506
# Generate "station-pairs-json" data, locally, for all months
ctbk station-pairs-json create
# Process specific date ranges for different stages
ctbk normalized -d 202206-202209 create
ctbk aggregated -g "ymd" -a "c" -d 202206-202209 create
# View URLs that would be processed
ctbk normalized -d 202206-202209 urlsThe ctbk CLI has subcommands for each pipeline stage: zip, csv, normalized, aggregated, station-meta-hist, station-modes-json, station-pairs-json.
Most create commands support:
-G, --no-git- Skip git/DVC workflow integration-e, --engine- Parquet engine selection (for normalized)-g, --group-by- Grouping keys (for aggregated)-a, --aggregate-by- Aggregation keys (for aggregated)
cd www/
npm run dev # Development server
npm run build # Production build
npm run export # Static site export
npm run lint # ESLint
npm run tc # TypeScript check
npm run scrns # Generate screenshots- Python:
ctbk/tests/+ctbk/pyramid_cascade/tests/(hermetic; a test needing network/cloud gets thenetworkmarker). Run withpytest; CI runs them via.github/workflows/py-tests.yml. - Workers: each
gbfs/<worker>vitest suite gates its deploy (gbfs.yml). - Frontend: Playwright e2e (
www/e2e/) gates the www deploy (www.yml). - Live API contract:
ctbk gbfs api-check(goldens inctbk/api_check_goldens/for closed rides windows + invariants), daily viaapi-check.yml; after a deliberate data repair, rerun with-uand commit the golden diff. - Linting: Use
npm run lintfor frontend, no Python linting configured
- Local data: Stored under
s3/ctbk/directory mirroring S3 structure - DVC integration: All data files tracked with
.dvcfiles for version control - Git workflow: Commands automatically stage DVC changes unless
-G/--no-gitused
- Use
-d YYYYMM-YYYYMMfor date ranges - Each subcommand supports
urls(preview paths) andcreate(generate data) - Pipeline stages depend on predecessors (e.g.,
aggregatedrequiresnormalized) - Data stored locally in
s3/ctbk/with DVC tracking
- Local development: Data in
s3/ctbk/with.dvctracking files - Production: Data synchronized to
s3://ctbk/via DVC - Public access: Datasets served content-addressed from the DVX cache (
s3://ctbk/.dvc/files/md5/…), migrating to R2 behinddata.ctbk.dev(specs/s3-to-r2-migration.md). The frontend already readsdata.ctbk.dev; the browsablectbk.s3.amazonaws.com/index.htmllisting remains on S3.
- CI (
.github/workflows/ci.yml): Monthly data ingestion from s3://tripdata - Website (
.github/workflows/www.yml): Deploys ctbk.dev (CF Workers Assets) on pushes tomaintouchingwww/**(or dispatch; the monthly pipeline dispatches it once at its end);wwwbranch = marker of the live commit - Tests:
py-tests.yml(pytest),api-check.yml(daily live-API contract check), plus the per-deploy vitest/Playwright gates
CI (.github/workflows/ci.yml) polls for new Citi Bike data monthly and runs the whole pipeline for the new month via a single driver, ctbk update (ctbk/update.py):
ctbk update -S <YYYYMM> # -S skips station-harmonize (whole-history; CI runs it separately)which runs, in order: norm → cons → smh -gil / smh -gin → agg ×5 (-ge -ac, -gse -ac, -g ymrgtb -acd, -g ymrgtbs -acd, -g ymrgtbe -acd) → sm → spj → station-trips-json -a -d (per-station ymdgtb JSONs) → node www/scripts/gen-station-urls.js. The rides rollup-pyramid rebuild (R2-only, no DVX artifacts) runs afterward as a separate best-effort CI step. The root update.sh is a thin deprecated pointer to this command.
- Python uses Click for CLI framework
- Frontend uses Next.js 14 with TypeScript
- Data processing relies heavily on pandas/pyarrow
- Visualization uses Plotly.js and Leaflet maps
- Python: pandas, pyarrow, boto3, s3fs, dvc-s3, plotly, click
- Frontend: Next.js, React, @mui/material, plotly.js, leaflet
setup.py- Python package withctbkandymsCLI entry pointswww/package.json- Frontend dependencies and scriptsrequirements.txt- Python dependencies- ESLint/TypeScript configured for frontend code quality
The project uses DVC extensively:
- Data files tracked with
.dvcfiles in git - Remote storage in S3 (
s3://ctbk/) - Use
-G/--no-gitflag to skip DVC workflow during development