This guide covers how to build datasets and run analysis tools in hvantk.
Migration note: hvantk has retired the
mktableandmkmatrixCLIs. The unified replacement ishvantk reprocess <plugin>:<dataset>, which runs the full Phase B pipeline (download → parse → build → drift check → save with provenance). See section 1 below for the current pattern.
If you haven't installed hvantk yet, see the main README for install steps.
For downloading raw data files (built-in downloaders and manual steps), see Data Sources.
Every in-tree provider is a plugin under hvantk/skills/<provider>/ with a plugin.yaml manifest declaring its datasets, lifecycle (download / parse / build), drift probe, and artifact contract. The single CLI entry point is hvantk reprocess.
# List all loaded plugins and their datasets
hvantk plugins list
# Show one plugin's manifest in detail
hvantk plugins describe clinvar# ClinVar variants
hvantk reprocess clinvar:variants \
--raw-dir /data/clinvar/ \
--output /out/clinvar.ht
# HGNC lookup
hvantk reprocess hgnc:lookup \
--raw-dir /data/hgnc/ \
--output /out/hgnc.ht# Already-downloaded raw file; skip the download stage
hvantk reprocess clinvar:variants \
--skip-download \
--raw-dir /data/clinvar/ \
--output /out/clinvar.ht
# Pre-parsed intermediate file; skip both download and parse
hvantk reprocess gevir:metrics \
--skip-download --skip-parse \
--intermediate /data/gevir.tsv.bgz \
--output /out/gevir.htUse --plugin-arg KEY=VALUE (repeatable) to forward arguments that the plugin's download / parse / build functions accept. For example, to fetch a specific CPTAC cancer type:
hvantk reprocess cptac:expression \
--raw-dir /data/cptac/ \
--output /out/cptac_brca.h5ad \
--plugin-arg cancer_type=brcaCPTAC phosphoproteomics builds one AnnData per cancer type — the single
cancer_type flows to both the download and the build:
hvantk reprocess cptac:phospho \
--raw-dir /data/cptac/ \
--output /out/cptac_phospho_brca.h5ad \
--plugin-arg cancer_type=brcaPredict genetic ancestry for samples using PCA and Random Forest classification against a labeled reference panel.
# Predict ancestry using 1000 Genomes as reference
hvantk ancestry-inference \
-q /data/my_cohort.mt \
-r /data/1kg_phase3.mt \
--ancestry-col super_pop \
-o /out/ancestry_predictions.ht \
--generate-report \
--export-tsv# Conservative assignment with custom filtering
hvantk ancestry-inference \
-q /data/my_cohort.mt \
-r /data/1kg_phase3.mt \
--ancestry-col super_pop \
-o /out/ancestry_predictions.ht \
--min-af 0.05 \
--min-call-rate 0.99 \
--n-pcs 30 \
--n-pcs-classify 15 \
--min-prob 0.90 \
--generate-reportimport hail as hl
from hvantk.algorithms.ancestry import run_ancestry_inference
# Initialize Hail
hl.init()
# Load data
query_mt = hl.read_matrix_table("my_cohort.mt")
reference_mt = hl.read_matrix_table("1kg_phase3.mt")
# Run inference
result = run_ancestry_inference(
query_mt=query_mt,
reference_mt=reference_mt,
ancestry_col="super_pop",
min_prob=0.75,
)
# Get results
predictions = result.get_predictions_df()
print(predictions['predicted_ancestry'].value_counts())
# Generate visualizations
result.generate_report("ancestry_report.html")
fig = result.plot_pca()
fig.savefig("pca_plot.png", dpi=300)
# Annotate original MT with predictions
annotated_mt = result.annotate_matrixtable(query_mt)| File | Description |
|---|---|
predictions.ht |
Hail Table with ancestry predictions |
predictions.tsv |
TSV export (with --export-tsv) |
ancestry_report.html |
HTML report with visualizations |
rf_model.pkl |
Trained model (with --save-model) |
pca_loadings.ht |
PCA loadings (with --save-loadings) |
Full Ancestry Documentation | Examples
Inspect, summarize, and extract markers from expression AnnData (.h5ad) files.
# Inspect observation metadata fields
hvantk expression describe -m /data/heart_sc.h5ad
# Collapse into a per-group × per-gene summary AnnData grouped by cell type
hvantk expression summarize \
-m /data/heart_sc.h5ad \
--group-by cell_type \
-o /out/heart_celltype_summary.h5ad
# Multi-field grouping with pre-filtering
hvantk expression summarize \
-m /data/heart_sc.h5ad \
--group-by cell_type --group-by region \
--filter-by time_point=9wpc \
--min-cells 50 \
-o /out/heart_summary.h5ad
# Extract marker genes (Wilcoxon rank-sum test)
hvantk expression markers \
-m /data/heart_sc.h5ad \
--method wilcoxon \
--group-by cell_type \
--top-n 200 \
-o /out/heart_markers.json
# Extract markers using a t-test, pre-filtered to one region
hvantk expression markers \
-m /data/heart_sc.h5ad \
--method t-test \
--group-by cell_type \
--filter-by region=LV \
--top-n 200 \
-o /out/heart_LV_markers.jsonConvert plain-text gene panels into GeneSetCollection JSON files for use
with hvantk psroc, hvantk enrichex burden, and hvantk enrichex overlap.
# Format: headerless two-column TSV (gene_set_name<TAB>gene_symbol)
hvantk genesets prepare -i panels.tsv -o panels.json
# With HGNC validation and alias resolution (recommended)
hvantk genesets prepare -i panels.tsv -o panels.json --hgnc /data/hgnc.ht
# Filter small sets and provide explicit background
hvantk genesets prepare -i panels.tsv -o panels.json \
--hgnc /data/hgnc.ht --min-genes 5 --background bg_genes.txt
# Also export as GMT for GSEA compatibility
hvantk genesets prepare -i panels.tsv -o panels.json --export-gmt panels.gmt
# Extract gene sets from COSMIC Cancer Gene Census
hvantk genesets cosmic --ht /data/cosmic_cgc.ht -o cosmic_gene_sets.jsonfrom hvantk.core.utils.geneset_io import parse_geneset_tsv, validate_with_catalog
from hvantk.core.utils.gene_sets import load_gene_sets_from_dict
from hvantk.skills.hgnc.streamers import HGNCGeneCatalogStreamer
# Parse TSV
result = parse_geneset_tsv("panels.tsv")
# Optional: validate against the HGNC gene catalog (resolves aliases)
catalog = HGNCGeneCatalogStreamer.from_path("/data/hgnc.ht")
vr = validate_with_catalog(result.gene_sets, catalog)
# Build and save collection
collection = load_gene_sets_from_dict(vr.gene_sets, source="prepare-geneset")
collection.save("panels.json")from hvantk.skills.clingen.streamers import ClinGenGeneDiseaseTableStreamer
streamer = ClinGenGeneDiseaseTableStreamer("/data/clingen/clingen_gene_disease.ht")
# High-confidence genes
definitive = streamer.get_genes_by_classification("Definitive")
# Disease keyword search
cancer_genes = streamer.get_genes_by_disease(
["cancer", "carcinoma", "tumor"],
match_mode="contains",
min_classification="Moderate",
)
# Dataset stats + summary
stats = streamer.compute_stats()
summary = streamer.classification_summary()
# Export gene sets for EnrichEx
streamer.export_for_enrichex(
"/out/clingen_gene_sets.json",
min_classification="Moderate",
)
# Translate output to Ensembl IDs using the HGNC gene catalog
from hvantk.skills.hgnc.streamers import HGNCGeneCatalogStreamer
catalog = HGNCGeneCatalogStreamer.from_path("/data/hgnc/hgnc.ht")
ensembl_ids = streamer.get_genes_by_classification(
min_classification="Definitive",
gene_catalog=catalog,
output_id_type="ensembl_gene_id",
)
# Or translate to HGNC IDs
hgnc_ids = streamer.to_gene_set(
min_classification="Moderate",
gene_catalog=catalog,
output_id_type="hgnc_id",
)Hail supports standard gzip (.gz) and uncompressed files but processes them single-threaded. Block gzip (BGZF) .bgz files enable parallel import and are strongly recommended for large datasets. Convert with hvantk utils convert-bgz:
# Default: replaces .gz extension with .bgz
hvantk utils convert-bgz input.tsv.gz
# Custom output path and thread count
hvantk utils convert-bgz input.tsv.gz -o output.tsv.bgz --threads 4The command auto-detects whether the file is already BGZF and skips conversion if so.
hvantk annotate turns built datasets into one gene × feature matrix. It is three
commands run in order, because each stage is independently re-runnable and the
expensive middle stage is per-source:
| Stage | Command | Produces |
|---|---|---|
| 1. Spine | annotate spine |
the gene table every annotation joins onto |
| 2. Prepare | annotate prepare (once per axis) |
one source mapped onto the spine's gene_id |
| 3. Compose | annotate compose |
the joined matrix + a manifest JSON |
The spine fixes which genes exist and what gene_id means for the whole matrix.
hvantk annotate spine \
--gene-table ensembl_structure.ht \
--hgnc hgnc_lookup.ht \
--output spine.ht--biotype defaults to protein_coding; pass --biotype all to keep every biotype.
A feature spec names each axis, the dataset it comes from, the key to join on, and the columns to keep:
name: chd-features
layer1:
- {axis: constraint, source: gnomad-metrics:metrics, key: gene_id, columns: [mis_z]}
- {axis: gevir, source: gevir:metrics, key: gene_id, columns: [gevir_pct]}layer1: is a literal key, not a label you choose — it names the gene-level evidence
axes. Each axis label must be unique: it is how the later stages address the entry
(--axis constraint, --prepared constraint=…) and it names the {axis}_present
column. The join itself is on key:, which is gene_id here.
key accepts gene_id, hgnc_id, symbol, uniprot_id or variant. Anything but
gene_id needs prepare --hgnc to resolve the mapping — as does a variant entry
that aggregates to something other than gene_id. Run
hvantk annotate prepare --help for the authoritative rule.
prepare runs once per axis and is the stage worth parallelising — each invocation is
independent, so they can be separate cluster jobs:
hvantk annotate prepare --spec features.yaml --axis constraint \
--input gnomad_metrics.ht --spine spine.ht --output constraint.ht
hvantk annotate prepare --spec features.yaml --axis gevir \
--input gevir_metrics.ht --spine spine.ht --output gevir.htcompose left-joins every prepared axis onto the spine. It requires one --prepared
per axis in the spec and fails if any is missing, so a partial matrix cannot be produced
silently:
hvantk annotate compose --spec features.yaml --spine spine.ht \
--prepared constraint=constraint.ht \
--prepared gevir=gevir.ht \
--output features.htThe manifest JSON lands next to the output (features.ht.manifest.json unless
--manifest says otherwise) and records what went into the matrix.
Scheduling is deliberately external: the commands do their work in-process and hvantk never imports a scheduler, so an sbatch wrapper submits them.
hvantk cohort brings a cohort's own gene-level results alongside the annotation
matrix. A cohort is declared by a manifest, not by flags:
name: my-cohort
key: symbol # or gene_id / hgnc_id
key_column: gene # the column in `table` holding that key
table: cohort.tsv
prior:
column: minp
direction: lower_is_better
cohort_axes:
- axis: architecture
columns: [n_case_var, conc, driver_af]# Check the manifest against its table (and labels) before anything reads it
hvantk cohort validate --cohort cohort.yaml
# Originate a gene-level prior with a Fisher-exact burden test
hvantk cohort burden --help
# Join the cohort's declared columns onto the Layer-1 matrix
hvantk cohort attach --helpvalidate first is the point: it fails on a key column that is not present, or a
declared axis column missing from the table, rather than letting a silently-empty join
propagate into the matrix.
hvantk rerank does not read this manifest directly: it takes its own config,
which points at the manifest and adds the evidence axes and labels —
cohort: cohort.yaml # the manifest above
features:
- {name: constraint, path: constraint.tsv}
labels:
path: labels.txtThe manifest's prior.column is carried through to the output as prior_stat. See
examples/rerank for a
complete runnable configuration.
- Use
--overwriteto replace an existing output. Without it, builders abort if the output exists. - For UCSC, gene labels may be pipe-delimited (e.g., A|B);
--split-gene-fielddefaults to true. - Expression AnnData stores per-cell/sample metadata in
adata.obs; inspect it withhvantk expression describe -m <file.h5ad>oradata.obsin Python. - gzip vs BGZF: Hail reads standard gzip files single-threaded, which is significantly slower for large files. Pre-convert with
hvantk utils convert-bgz <file.gz>for parallel import before runninghvantk reprocess.
- Architecture – system design and extension points
- Contributing – development workflow and contribution guidelines
- Data Sources – available datasets and how to acquire them