Data Sources#
Overview#
All data is pulled by code rather than committed to the repo: raw and
processed files live under _data/, which is git-ignored and fully
regenerable by rerunning the pipeline. The only data tracked in Git is
data_manual/, reserved for manually-collected inputs that cannot be
re-pulled automatically.
Datasets#
Dataset |
Source |
Frequency |
Description |
|---|---|---|---|
FRED macro/financial series |
FRED, via |
Monthly, Quarterly |
Moody’s Aaa/Baa seasoned corporate bond yields, the 10-year Treasury yield, the 3-month T-bill rate, CPI, population, real GDP, and the NBER recession indicator. |
Robert Shiller stock-market data |
|
Monthly |
S&P Composite price, dividends, earnings, CPI, GS10, and the cyclically adjusted price-earnings ratio (CAPE / P/E10). |
Greenwood-Hanson high-yield share (HYS) |
Published Greenwood & Hanson (2013) series (1926-2008) spliced with a Mergent FISD reconstruction via WRDS (2009-present) ( |
Annual |
Fraction of gross nonfinancial corporate bond issuance rated below investment grade. |
Data Pipeline#
FRED data (pull_fred.py)#
Pull. Downloads each series in
series_to_pull(Aaa/Baa yields, GS10, historical long-term/short-term bond-yield series used to extend coverage before the modern series begin, TB3MS, CPI, population, GDP, USREC) viapandas_datareader.data.DataReader.Store. Cached to
_data/raw_data/fred.parquet/.csv, with a data dictionary written to_data/data_dictionaries/fred_data_dictionary.md.Process.
process_fred_data_monthly.pyandprocess_fred_data_annual.pybuild the cleaned Baa/Aaa-Treasury credit spreads, GDP-per-capita growth, inflation, and the NBER recession indicator used throughout the replication, cached to_data/processed_data/fred_final_series_monthly.parquetandfred_final_series_annual.parquet.
Robert Shiller’s data (pull_shiller.py)#
Pull.
pull_shillerdownloadsie_data.xlsfrom https://shillerdata.com/ (default, overridable via theSHILLER_URLsetting). Since shillerdata.com serves the workbook from a versioned CDN link that changes whenever Shiller updates the file, the puller scrapes the current link from the landing page rather than hard-coding it, then parses theDatasheet.Structure. The
Datasheet is one row per month from 1871-01 onward. The rawDatecolumn encodes October asYYYY.1, so the monthly index is rebuilt from row order rather than parsed from that column.Process.
process_shiller_annualcollapses to annual frequency (year-end by default) and addsln_pe10 = ln(P/E10).Store. Cached to
_data/raw_data/shiller_data.parquet(monthly) and_data/processed_data/shiller_data_annual.parquet.
Notes#
All pulled and derived data live in
_data/, which is git-ignored, so no data is committed to the repo.Run all three pulls together with
doit pull_data, or run the full pipeline (pulls, processing, replication, notebooks, report, and site) with a singledoit.