datalad-fuse lets you read files of DataLad
datasets and git-annex repositories
without downloading them first: only the parts of files that are actually
read are fetched, via fsspec, from
the URLs that git-annex knows. Unlike datalad get, which downloads whole
files before you can use them, this pays off when you need only parts of
large files, e.g. a few arrays from NWB/HDF5 files or the headers of many
images. Use it through a FUSE mount, so that any program can open the files,
or directly from Python.
Documentation: https://datalad-fuse.readthedocs.io
python3 -m pip install datalad-fuse
git-annex is required, and FUSE
(libfuse 2 or 3) for mounting datasets (e.g. sudo apt-get install fuse3 on
Debian/Ubuntu). See the
installation instructions
for details.
Clone a dataset, here Dandiset 000582 from the DANDI Archive. This gets all file names, but none of the 1.86 GB of file content:
datalad clone https://github.com/dandisets/000582
On the command line, mount the dataset; datalad fusefs keeps running until
the dataset is unmounted:
mkdir mnt
datalad fusefs -d 000582 --foreground mnt
Then, in another terminal, use any tool on its files (here h5ls from the
HDF5 tools), and unmount when done:
h5ls mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
fusermount -u mnt
In Python, open files directly, without FUSE:
from contextlib import closing
import h5py
import pynwb
from datalad_fuse.fsspec import DatasetAdapter
# caching=False: keep fetched data in memory only
with closing(DatasetAdapter("000582", caching=False)) as dsa:
with dsa.open("sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb") as f:
with h5py.File(f, "r") as h5, pynwb.NWBHDF5IO(file=h5) as io:
print(io.read().units.to_dataframe())The tutorial continues this example, and the documentation covers the command line and Python interfaces in detail.