Collection of MSNBC show transcripts from 2008–2022.
All data are available on Harvard Dataverse: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/UPJDE1
10,739 transcripts across 216 shows, May 2010 – October 2022.
| Year | Transcripts |
|---|---|
| 2010 | 43 |
| 2011 | 115 |
| 2012 | 233 |
| 2013 | 237 |
| 2014 | 217 |
| 2015 | 994 |
| 2016 | 906 |
| 2017 | 1,185 |
| 2018 | 1,467 |
| 2019 | 1,475 |
| 2020 | 1,288 |
| 2021 | 1,471 |
| 2022 | 1,108 |
Files:
msnbc_transcripts_api_2010-2022.tar.gz— HTML transcript files (186MB)msnbc_transcripts_api_2010-2022_metadata.csv— metadata for all transcriptsmsnbc_shows_api.csv— show list with transcript counts
5,369 transcripts from the now-defunct http://www.nbcnews.com/id/3719710, scraped via archive.org.
| Year | Transcripts |
|---|---|
| 2008 | 76 |
| 2009 | 434 |
| 2010 | 752 |
| 2011 | 1,042 |
| 2012 | 1,164 |
| 2013 | 1,177 |
| 2014 | 724 |
Files (on Dataverse):
- Raw HTML files and final CSV available at: https://doi.org/10.7910/DVN/ND1TCV
Local files:
data/legacy_2008-2014/legacy_2008-2014_links.csv— list of all transcript URLsdata/legacy_2008-2014/legacy_2008-2014_links_with_metadata.txt— URLs with show names and dates
Notes:
- Scripts from 2014
- Some transcripts had date typos (e.g., 'Thusday', 'Februrary') causing parsing failures
- 2003–2014 scrape: 16k transcripts from an earlier collection
- 2025 HTML scrape:
msnbc_transcripts_2022.csv.gz— transcripts from 2020–2025 scraped from listing pages
| Column | Description |
|---|---|
id |
Unique transcript ID |
date |
Publication date (ISO 8601) |
title |
Transcript title |
url |
Original URL |
slug |
URL slug |
guests |
Guest names |
show_ids |
Associated show IDs |
modified |
Last modified date |
| Column | Description |
|---|---|
id |
Show ID |
name |
Show name |
slug |
URL slug |
count |
Number of transcripts |
Each transcript is saved as an HTML file named {id}.html containing the full transcript text.
For reproducibility or extending the dataset:
- WordPress API Scraper — REST API scraper (recommended)
- HTML Scraper — scrapes transcript listing pages
- Quick Peek — preview data
- Upload to Dataverse
Legacy scripts (2008–2014):
- Legacy Crawl — crawls archive.org for transcript links
- Legacy Extract — extracts transcripts from HTML files