HETUS Metadata
HETUS Database
HETUS Data Browser
HETUS Guidelines 2020
HETUS data delivery guidelines 2020
HETUS Data Delivery Guidelines
Codes Activités
Staggered Grid Example - Motion.dev
Animated Waffle Charts with D3 and GSAP
-
Set your environement:
python -m venv .venv_hf .\.venv_hf\Scripts\activate.bat pip install -r requirements.txt
-
Set your Hugging Face token and repository in
.env. -
Smoke run:
python .\scripts\db-hf-normalization.py --hetus-root .\hetus\2020 --max-files 1 --overwrite
-
Full run:
python .\scripts\db-hf-normalization.py --hetus-root .\hetus --overwrite
-
Validation:
python -c "import pandas as pd; f=pd.read_csv('hf_export/intermediate/files.csv'); s=f['lossless_validated']; print('all_lossless=', s.astype(str).str.strip().str.lower().eq('true').all()); print('rows_match=', (f['rows_long']==f['rows_expected_long']).all()); print('files=', len(f))"And you should see:
all_lossless= True rows_match= True -
To push to Hugging Face:
python .\scripts\db-hf-normalization.py --hetus-root .\hetus --overwrite --push-to-hub
Since the three tables have different schemas, the script publishes one Hub config per table:
observations,files,metadata. -
Load from Hugging Face (example):
python -c "from datasets import load_dataset; print(load_dataset('Bluefir/hetus-time-use', 'observations', split='train'))"
Generate the JSON summary used by the frontend:
python .\scripts\extract_json-data.pyPrerequisites:
- Install dependencies first:
pip install -r requirements.txt - If the dataset repository is private, set
HF_TOKENin.envor in your shell environment. - If you have a local export already, use
--local-dataset-dirto read the saved dataset instead of downloading remote parquet shards.
Examples:
python .\scripts\extract_json-data.py --repo-id Bluefir/hetus-time-use --config observations --split trainpython .\scripts\extract_json-data.py --geo DE --sex T --unit TIME_SPpython .\scripts\extract_json-data.py --local-dataset-dir .\hf_export\hf_dataset_unified_aclNotes:
- By default,
extract_json-data.pykeeps all availablegeoandsexvalues. - To restrict the extraction, pass explicit filters such as
--geo DEor--sex T. - If
--local-dataset-direxists, the script will use the local save-to-disk dataset first. - The script scans the remote
observations/trainparquet shards incrementally, so it can process the full HF split without loading all 1.68M rows in RAM at once.
Generate a small graph from the published HF dataset (observations config):
python .\scripts\plot_hf_age_category_year.py --repo-id Bluefir/hetus-time-use --geo DE --sex T --unit TIME_SP --max-categories 4 --output .\hf_export\age_category_by_year.pngNotes:
--geo DEis a good default for full 2000/2010/2020 coverage.- For a private dataset repo, set
HF_TOKENin.env(or environment). - You can also run against local exported data with
--local-dataset-dir .\hf_export\hf_dataset.
Generate a compatibility report to inspect:
- ACL00/ACL18 overlap and incompatibilities
- strict-null and wave-specific fields
- cross-wave domain overlap for dimension columns
- suggested merge candidates for wave-specific fields
python .\scripts\hf_harmonization_audit.py --local-dataset-dir .\hf_export\hf_dataset --output-json .\hf_export\harmonization_audit_report.json --output-md .\hf_export\harmonization_audit_report.mdIf --local-dataset-dir does not exist, the script falls back to loading from Hugging Face (--repo-id, --config, --split).
Use add_unified_acl18_column.py to create a new unified_acl_codes column in the observations table:
- if
acl18is present, keep it; - otherwise map
acl00using Full Review and Mappings fromacl00_acl18_mapping_need.md. - Mapping is embedded in the script and has been validated against the markdown source.
- After processing, the script prints statistics for rows with
acl18, rows with mappedacl00, unmappedacl00, and unique code counts.
Small local smoke run:
python .\scripts\add_unified_acl18_column.py --local-dataset-dir .\hf_export\hf_dataset --config observations --max-rows 2000 --output-dir .\hf_export\hf_dataset_unified_acl_smokeFull local run:
python .\scripts\add_unified_acl18_column.py --local-dataset-dir .\hf_export\hf_dataset --config observations --output-dir .\hf_export\hf_dataset_unified_aclPush updated split to Hugging Face Hub:
python .\scripts\add_unified_acl18_column.py --repo-id Bluefir/hetus-time-use --config observations --push-to-hub --target-repo-id Bluefir/hetus-time-use --target-config observations --target-split trainUse plot_unified_acl_categories.py to plot a few selected unified ACL18 activity codes with descriptions from mappings/activities_ACL18.yml:
python .\scripts\plot_unified_acl_categories.py --local-dataset-dir .\hf_export\hf_dataset --yml-file .\mappings\activities_ACL18.yml --categories AC11-12,AC0_X_021 --geo DE --sex T --unit TIME_SP --output .\hf_export\unified_acl_categories.pngIf --categories is omitted, the script selects the top categories by mean minutes automatically.