Skip to content

Repository files navigation

visualisation-app5

Hugging Face Dataset

HETUS Metadata
HETUS Database
HETUS Data Browser

Database Reference

HETUS Guidelines 2020
HETUS data delivery guidelines 2020
HETUS Data Delivery Guidelines
Codes Activités

Animation Reference

Staggered Grid Example - Motion.dev
Animated Waffle Charts with D3 and GSAP

Conversion

  1. Set your environement:

    python -m venv .venv_hf
    .\.venv_hf\Scripts\activate.bat
    pip install -r requirements.txt
  2. Set your Hugging Face token and repository in .env.

  3. Smoke run:

    python .\scripts\db-hf-normalization.py --hetus-root .\hetus\2020 --max-files 1 --overwrite
  4. Full run:

    python .\scripts\db-hf-normalization.py --hetus-root .\hetus --overwrite
  5. Validation:

    python -c "import pandas as pd; f=pd.read_csv('hf_export/intermediate/files.csv'); s=f['lossless_validated']; print('all_lossless=', s.astype(str).str.strip().str.lower().eq('true').all()); print('rows_match=', (f['rows_long']==f['rows_expected_long']).all()); print('files=', len(f))"

    And you should see:

    all_lossless= True
    rows_match= True
    
  6. To push to Hugging Face:

    python .\scripts\db-hf-normalization.py --hetus-root .\hetus --overwrite --push-to-hub

    Since the three tables have different schemas, the script publishes one Hub config per table: observations, files, metadata.

  7. Load from Hugging Face (example):

    python -c "from datasets import load_dataset; print(load_dataset('Bluefir/hetus-time-use', 'observations', split='train'))"

Export age/category/year JSON

Generate the JSON summary used by the frontend:

python .\scripts\extract_json-data.py

Prerequisites:

  • Install dependencies first: pip install -r requirements.txt
  • If the dataset repository is private, set HF_TOKEN in .env or in your shell environment.
  • If you have a local export already, use --local-dataset-dir to read the saved dataset instead of downloading remote parquet shards.

Examples:

python .\scripts\extract_json-data.py --repo-id Bluefir/hetus-time-use --config observations --split train
python .\scripts\extract_json-data.py --geo DE --sex T --unit TIME_SP
python .\scripts\extract_json-data.py --local-dataset-dir .\hf_export\hf_dataset_unified_acl

Notes:

  • By default, extract_json-data.py keeps all available geo and sex values.
  • To restrict the extraction, pass explicit filters such as --geo DE or --sex T.
  • If --local-dataset-dir exists, the script will use the local save-to-disk dataset first.
  • The script scans the remote observations/train parquet shards incrementally, so it can process the full HF split without loading all 1.68M rows in RAM at once.

Quick chart: time by age/category/year

Generate a small graph from the published HF dataset (observations config):

python .\scripts\plot_hf_age_category_year.py --repo-id Bluefir/hetus-time-use --geo DE --sex T --unit TIME_SP --max-categories 4 --output .\hf_export\age_category_by_year.png

Notes:

  • --geo DE is a good default for full 2000/2010/2020 coverage.
  • For a private dataset repo, set HF_TOKEN in .env (or environment).
  • You can also run against local exported data with --local-dataset-dir .\hf_export\hf_dataset.

Harmonization audit (ACL00 vs ACL18 + nullability)

Generate a compatibility report to inspect:

  • ACL00/ACL18 overlap and incompatibilities
  • strict-null and wave-specific fields
  • cross-wave domain overlap for dimension columns
  • suggested merge candidates for wave-specific fields
python .\scripts\hf_harmonization_audit.py --local-dataset-dir .\hf_export\hf_dataset --output-json .\hf_export\harmonization_audit_report.json --output-md .\hf_export\harmonization_audit_report.md

If --local-dataset-dir does not exist, the script falls back to loading from Hugging Face (--repo-id, --config, --split).

Add unified_acl_codes column (ACL18 harmonized)

Use add_unified_acl18_column.py to create a new unified_acl_codes column in the observations table:

  • if acl18 is present, keep it;
  • otherwise map acl00 using Full Review and Mappings from acl00_acl18_mapping_need.md.
  • Mapping is embedded in the script and has been validated against the markdown source.
  • After processing, the script prints statistics for rows with acl18, rows with mapped acl00, unmapped acl00, and unique code counts.

Small local smoke run:

python .\scripts\add_unified_acl18_column.py --local-dataset-dir .\hf_export\hf_dataset --config observations --max-rows 2000 --output-dir .\hf_export\hf_dataset_unified_acl_smoke

Full local run:

python .\scripts\add_unified_acl18_column.py --local-dataset-dir .\hf_export\hf_dataset --config observations --output-dir .\hf_export\hf_dataset_unified_acl

Push updated split to Hugging Face Hub:

python .\scripts\add_unified_acl18_column.py --repo-id Bluefir/hetus-time-use --config observations --push-to-hub --target-repo-id Bluefir/hetus-time-use --target-config observations --target-split train

Plot unified ACL categories with YAML labels

Use plot_unified_acl_categories.py to plot a few selected unified ACL18 activity codes with descriptions from mappings/activities_ACL18.yml:

python .\scripts\plot_unified_acl_categories.py --local-dataset-dir .\hf_export\hf_dataset --yml-file .\mappings\activities_ACL18.yml --categories AC11-12,AC0_X_021 --geo DE --sex T --unit TIME_SP --output .\hf_export\unified_acl_categories.png

If --categories is omitted, the script selects the top categories by mean minutes automatically.

Used by

Contributors

Languages