This project builds an end-to-end cloud data warehouse for the New York Times Article Archive using Google Cloud Platform.
The pipeline:
- Extracts historical articles from the NYT Archive API.
- Lands raw JSON into Google Cloud Storage (GCS).
- Loads and models data in BigQuery using Airflow and dbt.
- Performs data cleaning, SQL analytics, and machine learning for insights.
- Is designed as a portfolio-ready example of:
- Data Engineering & Cloud Warehousing
- Data & Business Analytics
- Basic ML on real-world news data
Real news archives are rich but messy. The goal of this project is to:
- Build a reproducible pipeline to ingest NYT articles at scale.
- Store data in a query-friendly warehouse schema on BigQuery.
- Enable analytical SQL for editors, analysts, and product teams.
- Explore ML use cases such as topic trends and article popularity.
Key questions the warehouse can answer:
- How has coverage of specific topics (e.g., “climate change”, “elections”) evolved over time?
- Which sections and keywords are most common in different decades?
- What patterns appear in article length, sections, or publication volume?
High-level architecture:
-
Ingestion (Python + NYT API)
scripts/archive_extraction.pycalls the NYT Archive API for each year/month.- Raw JSON responses are stored in GCS.
-
Landing & Staging (GCS + BigQuery)
scripts/bigquery_pipeline.pyloads raw JSON from GCS into BigQuery landing tables.- Schema inference and basic normalization.
-
Orchestration (Apache Airflow)
dags/nyt_archive_dag.pycoordinates:- Archive extraction
- GCS load
- BigQuery staging/transformations
-
Transformations (dbt + SQL)
dbt/models.sql(placeholder for dbt models) defines warehouse-friendly tables:- Cleaned articles
- Dimensions (section, keyword, date)
- Aggregated fact tables
-
Analytics & ML (Notebooks + SQL)
notebooks/data_cleaning.ipynbperforms data exploration & cleaning logic.notebooks/machine_learning_analysis.ipynbexplores ML use cases such as:- Topic clustering
- Article popularity modelling
sql/sql_query_analysis.sqlcontains curated analytical queries for stakeholders.
-
Documentation
reports/GROUP5_REPORT.pdf– full written project report.reports/GROUP5_DB_PRESENTATION.pptx– presentation slides.
Cloud stack:
- GCP: BigQuery, GCS
- Orchestration: Apache Airflow
- Modeling: dbt + SQL
- Analysis: Jupyter / Python / SQL
nyt-archive-cloud-warehouse/
├── dags/
│ └── nyt_archive_dag.py # Airflow DAG for end-to-end pipeline
├── scripts/
│ ├── archive_extraction.py # NYT Archive API → GCS
│ └── bigquery_pipeline.py # GCS → BigQuery landing/staging tables
├── notebooks/
│ ├── data_cleaning.ipynb # Data understanding & cleaning
│ └── machine_learning_analysis.ipynb
├── dbt/
│ └── models.sql # dbt models for cleaned & aggregated tables
├── sql/
│ └── sql_query_analysis.sql # Analytical SQL queries on the warehouse
├── reports/
│ ├── GROUP5_REPORT.pdf # Detailed project report
│ └── GROUP5_DB_PRESENTATION.pptx
├── README.md
├── requirements.txt # Python dependencies
└── .gitignore
Note: Raw data files and credentials are not stored in this repository for security and size reasons.
Responsibilities:
- Call the NYT Archive API for a range of years/months.
- Handle pagination, retry logic, and simple error handling.
- Write raw responses to GCS as JSON or newline-delimited JSON.
Key config (loaded via environment variables):
NYT_API_KEY– NYT API keyGCP_PROJECT_ID– GCP projectGCS_BUCKET_NAME– target bucket for raw dataGCS_CREDENTIALS_BLOB– service account JSON blob name (or use ADC)
Example run (local):
export NYT_API_KEY="YOUR_API_KEY"
export GCP_PROJECT_ID="your-project-id"
export GCS_BUCKET_NAME="your-bucket"
export GCS_CREDENTIALS_BLOB="service-account.json"
python scripts/archive_extraction.py \
--start_year 1980 \
--end_year 2020(Arguments are illustrative; adjust to match the script’s CLI.)
Responsibilities:
- Read raw JSON from GCS.
- Create or update landing and staging tables in BigQuery.
- Apply initial flattening and type casting.
Example run:
python scripts/bigquery_pipeline.py \
--dataset raw_nyt \
--table archive_rawdags/nyt_archive_dag.py wires together:
- Extract NYT archive → GCS
- Load raw JSON → BigQuery
- Run staging / transformation steps (optionally calling dbt)
Typical Airflow concepts used:
PythonOperator– for calling extraction & load functions.- Scheduling (e.g., monthly runs for new archives).
- Dependency management between tasks.
Once deployed to an Airflow environment (e.g., Cloud Composer or local Airflow), the DAG can be triggered manually or on a schedule.
Example logical models (conceptual):
articles_clean– cleaned article records with normalized date, section, and headline fields.dim_section– section dimension with human-readable labels.fact_article_counts– counts of articles by date, section, keyword, etc.
These models:
- Simplify downstream queries.
- Enforce consistent transformations.
- Support BI tools and analysts.
Contains example questions such as:
- Article count by section and year.
- Most frequent keywords over time.
- Distribution of article word counts.
- Trends for selected topics (e.g., “economy”, “climate”, “elections”).
This file is written in standard SQL targeting BigQuery.
The goal of the ML component is not to build a production model, but to show:
- How a cloud warehouse feeds into ML workflows.
- How to explore textual news data for insights.
- Inspects raw data structure (JSON fields, nested arrays).
- Handles missing values, long text fields, and normalization.
- Defines reusable cleaning logic applied later in dbt or scripts.
Possible analyses include:
- Topic modeling or clustering on article abstracts/lead paragraphs.
- Basic classification or regression (e.g., predicting popularity proxies).
- Time-series plots of topic frequency.
These notebooks connect directly to BigQuery using the google-cloud-bigquery client, treating the warehouse as the single source of truth.
-
Python 3.9+
-
A GCP project with:
- BigQuery enabled
- GCS bucket created
-
NYT Archive API key
-
Airflow environment (for DAG execution) – local or managed
-
(Optional) dbt installed and configured for BigQuery
From the project root:
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txtExample requirements.txt (simplified):
pandas
requests
google-cloud-bigquery
google-cloud-storage
apache-airflow
dbt-core
dbt-bigquery
jupyter
Adjust as needed to match the actual imports in your environment.
Set the following before running scripts or Airflow tasks:
export GCP_PROJECT_ID="your-project-id"
export NYT_API_KEY="your-nyt-api-key"
export GCS_BUCKET_NAME="your-bucket"
export GCS_CREDENTIALS_BLOB="your-service-account.json"In production, use a more secure mechanism (Secret Manager, Airflow Connections, etc.).
Data Engineering / Cloud
- Built a full ETL/ELT pipeline: API → GCS → BigQuery.
- Used Airflow for orchestration and scheduling.
- Designed warehouse layers (raw, staging, modeled) using SQL and dbt.
Data / Business Analytics
- Designed SQL queries to answer concrete editorial and product questions.
- Produced aggregated tables suitable for BI tools and dashboards.
- Showed how different dimensions (section, keyword, date) drive insights.
Machine Learning / Data Science
- Used warehouse data in notebooks for topic analysis and basic ML experiments.
- Demonstrated model-ready feature extraction from a cloud warehouse.
- Add automated tests for extraction and transformations.
- Expand dbt project into separate models for each dimension/fact table.
- Integrate a BI layer (e.g., Looker Studio) on top of BigQuery.
- Add incremental loading for newly published articles instead of full historic runs.
- Build alerting around pipeline failures and data quality checks.
- New York Times Article Search & Archive APIs.
- Google Cloud Platform (BigQuery, GCS).
- dbt, Airflow, and the Python ecosystem for data engineering.