Skip to content

Latest commit

 

History

History
384 lines (281 loc) · 22.8 KB

File metadata and controls

384 lines (281 loc) · 22.8 KB

Crates.io version Conda version Crates.io downloads Conda downloads European Galaxy server biorXiv preprint

Logo

Deacon

Deacon filters DNA sequences in FASTA/Q files and streams using SIMD-accelerated minimizer comparison with query sequence(s), emitting either matching sequences (search mode), or sequences without matches (deplete mode). Sequences match when they share enough distinct minimizers with the indexed query to exceed chosen absolute and relative thresholds. Query size has little impact on filtering speed, enabling ultrafast search and depletion with gene-, genome- and pangenome-scale queries using a laptop. Deacon filters uncompressed FASTA/Q at gigabases per second on recent AMD, Intel (x86_64), and Apple arm64 systems. Built with panhuman host depletion in mind—yet broadly useful for searching large sequence collections—Deacon delivers leading classification accuracy for host depletion and unrivalled speed using 5GB of RAM.

Default parameters are carefully chosen but easily changed. Classification sensitivity, specificity and memory requirements may be tuned by varying k-mer length (-k), window size (-w), absolute match threshold (-a) and relative match threshold (-r) . Minimizer k and w are chosen at query index time, while the match thresholds can be chosen at filter time. Matching sequences are those that share enough distinct minimizers with the indexed query to exceed both the absolute threshold (-a, default 2 shared minimizers) and the relative threshold (-r, default 0.01 [1%] shared minimizers). For paired sequences, hits in either mate counts towards a single match threshold for the pair. Deacon reports filtering performance during execution and optionally writes a JSON --summary upon completion. Sequences can optionally be renamed using --rename for privacy and smaller file sizes. Deacon fully supports stdin/out (uncompressed fastx) and natively handles .gz, .zst and .xz fastx file IO, as well as BINSEQ CBQ (.cbq & .cba).

Benchmarks for panhuman host depletion of complex microbial metagenomes are described in a preprint. Deacon with the panhuman-1 (k=31, w=15) index exhibited the highest balanced accuracy for both long and short simulated reads. Deacon was less specific only than Hostile for short reads.

Use cases

  • Depletion of human or other host genome sequences in FASTQ reads or streams.
  • Ultrafast binary classification of genes, genomes or pangenomes in terabase genome catalogues like AllTheBacteria without tedious pre-indexing.

Install

Conda/mamba/pixi Conda version

conda install -c bioconda deacon

Cargo Crates.io version

RUSTFLAGS="-C target-cpu=native" cargo install deacon

Important

Cargo installation requires Rust 1.88 or newer. Update using rustup update.

Docker Crates.io version

Containers are available from the BioContainers registry.

docker pull quay.io/biocontainers/deacon:0.15.0--hdd79491_0

Quickstart

Ultrafast panhuman host depletion

# Download validated 3GB human pangenome index (version 0.13.0 or later)
deacon index fetch panhuman-1

# Deplete long reads
deacon filter -d panhuman-1.k31w15.idx reads.fq -o filt.fq

# Deplete short paired reads
deacon filter -d panhuman-1.k31w15.idx reads.r1.fq.gz reads.r2.fq.gz -o filt.r1.fq.gz -O filt.r2.fq.gz

Ultrafast gene/genome/pangenome search

deacon index build amr-genes.fa > amr-genes.idx
deacon filter amr-genes.idx AllTheBacteria.fa.zst > hits.fa

Tip

If searching large sequence collections, consider encoding them as Binseq cbq for dramatically improved search throughput with Deacon

Prebuilt indexes

Prebuilt pangenome indexes can be downloaded using the links below, or with deacon index fetch <name>. These may alternatively be reproduced using workflows in deacon-indexes.

Name & URL Composition Minimizers Subtracted minimizers Size Date
panhuman-1 (k=31, w=15) Cloud, Zenodo HPRC Year 1 ∪ CHM13v2.0 ∪ GRCh38.p14 - bacteria (FDA-ARGOS) - viruses (RefSeq) 409,907,949 20,671 (0.0050%) 3.3GB 2025-04
panmouse-1 (k=31, w=15) Cloud, Zenodo GRCm39 ∪ PRJEB47108 - bacteria (FDA-ARGOS) - viruses (RefSeq) 551,041,865 9,866 (0.0018%) 4.4GB 2025-11

Note

Index compatibility. Deacon 0.11.0 and above uses index format version 3. Using version 3 indexes with older Deacon versions and vice versa triggers an error. Prebuilt indexes in legacy formats are archived in object storage and Zenodo to ensure reproducibility. To download indexes in legacy formats, replace the /3/ in any prebuilt index download URL with either /2/ or /1/ accordingly.

  • Deacon 0.11.0 and above uses index format version 3
  • Deacon 0.7.0 through to 0.10.0 used index format version 2
  • Deacon 0.1.0 through to 0.6.0 used index format version 1

Usage

Filtering

The main command deacon filter accepts an index path followed by up to two sequence file paths, depending on whether input sequences originate from stdin, a single file, or paired input files. Indexes are built with deacon index build. Paired inputs are supported as either two separate or one interleaved file/stream when using --interleaved, and may be written either to separate paired output files or one interleaved file. For paired sequences, distinct minimizer hits originating from either mate are counted. Paired read headers can be validated using --check-pairs. By default, input sequences must meet both an absolute threshold of 2 minimizer hits (-a 2) and a relative threshold of 1% of minimizers (-r 0.01) to pass the filter. Filtering can be inverted for e.g. host depletion using the --deplete (-d) flag. Gzip, Zstandard, and xz compressed FASTX formats are detected automatically by file extension. Single and paired BINSEQ CBQ files are natively supported for both input and output. CBQ input is detected by content rather than file extension, while a .cbq output path writes CBQ, and a .cba output path writes CBQ without quality.

Examples

# Keep only sequences matching a collection of genes
deacon index build genes.fa > genes.idx
deacon filter genes.idx sequences.fa.gz -o matches.fa.gz

# Host depletion using the panhuman-1 index
deacon filter -d panhuman-1.k31w15.idx reads.fq.gz -o filt.fq.gz

# High sensitivity host depletion with absolute threshold of 1 and no relative threshold
deacon filter -d -a 1 -r 0 panhuman-1.k31w15.idx reads.fq.gz -o filt.fq.gz

# High specificity 10% relative match threshold
deacon filter -d -r 0.1 panhuman-1.k31w15.idx reads.fq.gz > filt.fq.gz

# Stdin and stdout
zcat reads.fq.gz | deacon filter -d panhuman-1.k31w15.idx > filt.fq

# True multithreaded gzip decompression with rapidgzip
rapidgzip -dc reads.fq.gz | deacon filter -d panhuman-1.k31w15.idx > filt.fq

# Zstandard compression
deacon filter -d panhuman-1.k31w15.idx reads.fq.zst -o filt.fq.zst

# Paired reads
deacon filter -d panhuman-1.k31w15.idx r1.fq.gz r2.fq.gz > filt12.fq
deacon filter -d panhuman-1.k31w15.idx r1.fq.gz r2.fq.gz -o filt.r1.fq.gz -O filt.r2.fq.gz

# Validate matching Illumina record names while filtering paired sequences
deacon filter -d --check-pairs panhuman-1.k31w15.idx r1.fq.gz r2.fq.gz -o filt.r1.fq.gz -O filt.r2.fq.gz

# Interleaved paired reads (file or stdin)
deacon filter -d --interleaved panhuman-1.k31w15.idx r12.fq.gz > filt12.fq
zcat r12.fq.gz | deacon filter -d panhuman-1.k31w15.idx - - > filt12.fq

# Save summary JSON
deacon filter -d panhuman-1.k31w15.idx reads.fq.gz -o filt.fq.gz -s summary.json

# BINSEQ CBQ input and output; paired records are stored in one file
deacon filter -d panhuman-1.k31w15.idx reads.fq.gz -o filt.cbq
deacon filter -d panhuman-1.k31w15.idx r1.fq.gz r2.fq.gz -o filt12.cbq
deacon filter -d panhuman-1.k31w15.idx filt12.cbq -o filt12.fq.gz

# A .cba extension writes cbq without quality values, if present
deacon filter -d panhuman-1.k31w15.idx reads.fq.gz -o filt.cba

# Replace read headers with incrementing integers
deacon filter -d -R panhuman-1.k31w15.idx reads.fq.gz > filt.fq

# Only look for minimizer hits inside the first 1000bp per record
deacon filter -d -p 1000 panhuman-1.k31w15.idx reads.fq.gz > filt.fq

# Output FASTA regardless of input format (discards quality scores)
deacon filter -d -f panhuman-1.k31w15.idx reads.fq.gz > filt.fa

# Write per-record minimizer hits to a TSV
deacon filter -d --debug hits.tsv panhuman-1.k31w15.idx reads.fq.gz > filt.fq

Note

deacon filter uses 8 threads by default. Using more threads (e.g. --threads 16) can accelerate filtering given sufficient resources, especially with uncompressed sequences whose processing is not rate limited by decompression. Since version 0.13.0, Deacon writes gzipped output files (e.g -o out.fastq.gz) in parallel, providing particular practical benefit for gzipped paired reads. If output file(s) ending in .gz are detected, total --threads are allocated 1:1 to compression and filtering tasks respectively. Gzip compression thread allocation can be overriden with --compression-threads.

Indexing

# Index one FASTA/FASTQ file
deacon index build genome.fa.gz > genome.idx

# Index many FASTA/FASTQ files using stdin
zcat *.fa.gz | deacon index build - > genomes.idx

deacon index build accepts either a FASTA/FASTQ file or a stdin stream (-), enabling convenient indexing of compressed sequences in one or many files with a single step. Indexing a human genome takes a few seconds. Indexing uses 2-4x as much RAM as filtering. For indexing large collections approaching terabase scale—such as mammalian pangenomes—it may be practical to index genomes individually in parallel and later combine them using the deacon index union set operation, described below.

[!NOTE]

Indexing a 3 Gbp human genome takes ~5s and 8 GB of RAM with default parameters

Set operations

A differentiating feature of Deacon is the ease of combining, subtracting and intersecting minimizer indexes. For example, deacon index diffcan be used to subtract shared minimizers between target and host genomes when building custom indexes for host depletion.

  • Use deacon index union 1.idx 2.idx 3.idx… > 1+2+3.idx to succinctly combine two or more indexes.

  • Use deacon index diff 1.idx 2.idx > 1-2.idx to subtract minimizers in 2.idx from 1.idx. Useful for masking out shared minimizer content between e.g. target and host genomes.

    • deacon index diff also supports subtracting minimizers from an index using a fastx file or stream directly, e.g. deacon index diff 1.idx 2.fa.gz > 1-2.idx or zcat *.fa.gz | deacon index diff 1.idx - > 1-2.idx. This enables diffing with larger-than-memory sequence collections if desired.
    • Specifying w=1 subtracts every single k-mer in the second index/FASTX from the first.
  • Use deacon index intersect 1.idx 2.idx… > 1∩2.idx to find the intersection of minimizers in two or more indexes.

    • Indexes with w=1 are accepted in any position after the first, whatever the first index's w, reporting which minimizers of the first occur as exact k-mers in them.

Inspecting indexes

  • Use deacon index info 1.idx to display index information including minimizer k and w parameters, number of minimizers, and index format version.
  • Use deacon index dump 1.idx > 1.fa to dump a minimizer index to FASTA.

Command line reference

Filtering

$ deacon filter -h
Retain or deplete sequence records with sufficient minimizer hits to the index

Usage: deacon filter [OPTIONS] <INDEX> [INPUT] [INPUT2]

Arguments:
  <INDEX>   Path to minimizer index file
  [INPUT]   Optional path to fastx or binseq cbq file (or - for stdin) [default: -]
  [INPUT2]  Optional path to second paired fastx file

Options:
  -a, --abs-threshold <ABS_THRESHOLD>
          Minimum absolute number of distinct minimizer hits for a match [default: 2]
  -r, --rel-threshold <REL_THRESHOLD>
          Minimum proportion of distinct minimizer hits for a match (0.0-1.0) [default: 0.01]
  -p, --prefix-length <PREFIX_LENGTH>
          Search only the first N nucleotides per sequence (0 = entire sequence) [default: 0]
  -c, --complexity-threshold <COMPLEXITY_THRESHOLD>
          Ignore minimizer hits below this kdust complexity threshold (0.0-1.0)
  -d, --deplete
          Discard matching sequences (invert filtering behaviour)
  -R, --rename
          Replace sequence headers with incrementing numbers (deterministic with --ordered)
  -o, --output <OUTPUT>
          Path to output file (fastx to stdout by default; detects .gz, .zst, .xz, .cbq, .cba)
  -O, --output2 <OUTPUT2>
          Optional path to second paired output fastx file (detects .gz, .zst, .xz)
  -s, --summary <SUMMARY>
          Path to JSON summary output file
  -t, --threads <THREADS>
          Number of threads (0 = auto) [default: 8]
      --compression-threads <COMPRESSION_THREADS>
          Number of threads used for output compression (0 = auto) [default: 0]
      --compression-level <COMPRESSION_LEVEL>
          Output compression level (1-9 for gz & xz; 1-22 for zstd including cbq) [default: 2]
      --cbq-block-size <CBQ_BLOCK_SIZE>
          cbq output block size in MiB (or cbq input block size if higher) [default: 16]
      --discard-quality
          Emit fasta or quality-free cbq regardless of input format
      --interleaved
          Treat INPUT as interleaved paired records from single file or stdin
      --ordered
          Preserve input record ordering (deterministic, slightly slower)
      --check-pairs
          Validate paired record names (Illumina CASAVA or /1 /2 suffixes)
      --debug <PATH>
          Write per-record minimizer hits to TSV
  -q, --quiet
          Suppress progress reporting
  -h, --help
          Print help

Indexing

$ deacon index -h
Build, inspect, compose and fetch minimizer indexes

Usage: deacon index <COMMAND>

Commands:
  build      Index minimizers contained within a fastx file
  union      Combine multiple minimizer indexes (A ∪ B…)
  intersect  Intersect multiple minimizer indexes (A ∩ B…)
  diff       Subtract minimizers in one index from another (A - B)
  dump       Dump minimizer index to fasta
  filter     Discard minimizers below a complexity threshold
  info       Show index information
  freeze     Freeze an index into a binary fuse filter (BFF) index (k<=32)
  fetch      Fetch a pre-built index from remote storage
  help       Print this message or the help of the given subcommand(s)

Options:
  -h, --help  Print help
$ deacon index build -h
Index minimizers contained within a fastx file

Usage: deacon index build [OPTIONS] <INPUT>

Arguments:
  <INPUT>  Path to input fastx file (or - for stdin; supports gz, zst and xz compression)

Options:
  -k <KMER_LENGTH>         K-mer length used for indexing (k+w-1 must be <= 96 and odd) [default: 31]
  -w <WINDOW_SIZE>         Minimizer window size used for indexing [default: 15]
  -o, --output <OUTPUT>    Path to output file (stdout if not specified)
  -t, --threads <THREADS>  Number of execution threads (0 = auto) [default: 8]
  -q, --quiet              Suppress progress reporting
  -h, --help               Print help

Filtering summary statistics

Use -s summary.json to save detailed filtering statistics:

{
  "version": "deacon 0.18.0",
  "index": "panhuman-1.k31w15.idx",
  "input": "HG02334.100MB.fastq.gz",
  "input2": null,
  "output": "-",
  "output2": null,
  "k": 31,
  "w": 15,
  "abs_threshold": 2,
  "rel_threshold": 0.01,
  "prefix_length": 0,
  "deplete": true,
  "rename": false,
  "ordered": false,
  "check_pairs": false,
  "seqs_in": 37500,
  "seqs_out": 454,
  "seqs_out_proportion": 0.012106666666666667,
  "seqs_removed": 37046,
  "seqs_removed_proportion": 0.9878933333333333,
  "bp_in": 141474280,
  "bp_out": 227079,
  "bp_out_proportion": 0.001605090338682056,
  "bp_removed": 141247201,
  "bp_removed_proportion": 0.9983949096613179,
  "time": 3.354754749,
  "seqs_per_second": 83149,
  "bp_per_second": 313694198,
  "seqs_per_second_total": 11178,
  "bp_per_second_total": 42171273
}

Paired and interleaved counts cover both mates, so a pair contributes 2 to seqs_* and both its sequences to bp_*. time and the *_total rates include index loading; seqs_per_second and bp_per_second cover filtering only.

Server mode

From version 0.11.0, it is possible to eliminate index loading overhead at the start of each filter operation by preloading the index in the memory of a local server process. Subsequent filtering commands with --use-server are executed by the server process using a UNIX socket. Having started a server process, the index of the first filtering command it receives persists in memory for the life of that server process, enabling subsequent filter commands to be served rapidly without HashSet construction overhead.

# Start the server
deacon server start

# The first filter command loads the index as usual
deacon --use-server filter ref.idx reads.fq > /dev/null

# Subsequent filter commands use the existing index stored in memory
deacon --use-server filter ref.idx reads.fq -o filt.fq -s summary.json

# Stop the server
deacon --use-server server stop

Python bindings

Since 0.16.0, Deacon can also be used as a Python package distributed via PyPI. Like server mode, this enables multihreaded filtering with index reuse. Refer to the Python readme for more information.

Workflow manager integration

Nextflow (nf-core)

Galaxy

  • Tool Shed suite suite_deacon, installable into any Galaxy instance

Limitations

Deacon is benchmarked using both long reads and 2x150bp short reads. Classification accuracy deteriorates with very short reads however – if classifying old 50bp Illumina reads for example, consider using --abs-threshold 1 (-a 1).

Citation

biorXiv preprint

Bede Constantinides, John Lees, Derrick W Crook. "Deacon: fast sequence filtering and contaminant depletion" bioRxiv 2025.06.09.658732, https://doi.org/10.1101/2025.06.09.658732

Please also consider citing the SimdMinimizers paper:

Ragnar Groot Koerkamp, Igor Martayan. "SimdMinimizers: Computing random minimizers, fast" bioRxiv 2025.01.27.634998, https://doi.org/10.1101/2025.01.27.634998