Motivation
Data reaches us in many shapes: 10x h5, MTX folders, loom (velocyto, SCope/SCENIC+), CellBender output, Parse split-pipe folders, pycistopic pickles. Today converting them means loading everything into memory with scanpy/anndata. adata already reads and writes the AnnData on-disk spec with plain h5py/zarr and streams everything, so it is a natural home for a streaming, dependency-light converter:
adata convert INPUT -o OUTPUT [--from FMT] [--to FMT] [--zarr-format 2|3]
[--chunk N] [--layout csr|csc] [--force] [format-specific options]
This epic tracks the shared design; each format has its own sub-issue with a detailed plan.
Naming
adata convert currently changes a matrix's dtype/layout/density. It will be renamed to adata recast (with a deprecation period) so that convert means format conversion — see #21.
Proposed design
Format detection. Sniff content first, extension second, --from overrides:
| Signature |
Format |
HDF5 with /matrix/features |
10x h5 v3 (and CellBender if /droplet_latents is also present) |
HDF5 with /<genome>/genes + /<genome>/indptr |
10x h5 v2 |
HDF5 with /matrix + /row_attrs + /col_attrs |
loom (SCope loom if MetaData attr / col_attrs/RegulonsAUC) |
directory with matrix.mtx[.gz] + barcodes.tsv[.gz] |
10x MTX folder |
directory with count_matrix.mtx + all_genes.csv |
Parse split-pipe |
.pkl |
pycistopic (explicit --from pycistopic required) |
.h5ad, .zarr |
AnnData |
The target format comes from the -o extension; --to overrides it (e.g. --to 10x-h5, --to 10x-mtx).
Architecture. A new package (e.g. src/adata/formats/foreign/) with one module per format and a small registry. Each reader implements:
sniff(path) -> bool
plan(path, opts) -> ConversionPlan — shapes, index dtypes, elements to be written, warnings. Runs before any output is created.
write(plan, dst: Store)
Writers (export direction) mirror this with plan_export / write_export.
Reuse, don't reinvent: storage.open_store / copy_dataset / copy_store_contents, the elements.write helpers (write_sparse, write_dense, write_string_array, write_categorical, write_dataframe_header, ensure_anndata_skeleton), transpose_sparse_streaming and sparsify from core/convert.py, and the temp-path-then-rename logic from commands/convert.py.
Conventions every converter must follow (same as the existing convert/subset/concat):
- stream in chunks (
--chunk); --in-memory only as an opt-in
- validate everything before creating the output; a refusal never leaves a half-written store
- write to a temp path, rename into place
- refuse lossy or inflating operations unless
--force
- finish with
ensure_anndata_skeleton; honour --zarr-format
- no anndata/scipy/pandas at runtime — pure h5py/zarr wherever the format allows; heavy dependencies only behind optional extras (currently only needed for pycistopic pickles)
Common options: --var-names gene_symbols|gene_ids (scanpy default: symbols), --make-unique (like var_names_make_unique), --dtype (hands off to the recast code path).
Mapping convention: counts → X; extra matrices → layers; per-cell/per-feature attributes → obs/var; multi-column attributes → obsm/varm; anything tool-specific that doesn't fit → uns/<tool>/….
Sub-issues
Suggested order: #21 + #22 (framework with one real user) → #23, #24 (they provide the building blocks) → #29, #30 (reuse #23/#24) → #26 → #27 → #25, #28.
Definition of done (every sub-issue)
- The converter streams; a test in
tests/test_performance.py bounds its operation counts.
- Fixtures are built with h5py in
tests/conftest.py, parametrised over h5ad/zarr2/zarr3 outputs.
- The output opens in anndata (
integration marker, uv run --with anndata), and where scanpy has a reader for the format, the result is compared with scanpy's (uv run --with scanpy).
docs/COMMANDS.md, the "Format support" table in docs/index.md, the README and CHANGELOG.md are updated; tests/test_docs_are_accurate.py passes.
Non-goals
No analysis; no R-native formats (.rds) in the first round — see #31.
Motivation
Data reaches us in many shapes: 10x h5, MTX folders, loom (velocyto, SCope/SCENIC+), CellBender output, Parse split-pipe folders, pycistopic pickles. Today converting them means loading everything into memory with scanpy/anndata.
adataalready reads and writes the AnnData on-disk spec with plain h5py/zarr and streams everything, so it is a natural home for a streaming, dependency-light converter:This epic tracks the shared design; each format has its own sub-issue with a detailed plan.
Naming
adata convertcurrently changes a matrix's dtype/layout/density. It will be renamed toadata recast(with a deprecation period) so thatconvertmeans format conversion — see #21.Proposed design
Format detection. Sniff content first, extension second,
--fromoverrides:/matrix/features/droplet_latentsis also present)/<genome>/genes+/<genome>/indptr/matrix+/row_attrs+/col_attrsMetaDataattr /col_attrs/RegulonsAUC)matrix.mtx[.gz]+barcodes.tsv[.gz]count_matrix.mtx+all_genes.csv.pkl--from pycistopicrequired).h5ad,.zarrThe target format comes from the
-oextension;--tooverrides it (e.g.--to 10x-h5,--to 10x-mtx).Architecture. A new package (e.g.
src/adata/formats/foreign/) with one module per format and a small registry. Each reader implements:sniff(path) -> boolplan(path, opts) -> ConversionPlan— shapes, index dtypes, elements to be written, warnings. Runs before any output is created.write(plan, dst: Store)Writers (export direction) mirror this with
plan_export/write_export.Reuse, don't reinvent:
storage.open_store/copy_dataset/copy_store_contents, theelements.writehelpers (write_sparse,write_dense,write_string_array,write_categorical,write_dataframe_header,ensure_anndata_skeleton),transpose_sparse_streamingandsparsifyfromcore/convert.py, and the temp-path-then-rename logic fromcommands/convert.py.Conventions every converter must follow (same as the existing
convert/subset/concat):--chunk);--in-memoryonly as an opt-in--forceensure_anndata_skeleton; honour--zarr-formatCommon options:
--var-names gene_symbols|gene_ids(scanpy default: symbols),--make-unique(likevar_names_make_unique),--dtype(hands off to the recast code path).Mapping convention: counts →
X; extra matrices →layers; per-cell/per-feature attributes →obs/var; multi-column attributes →obsm/varm; anything tool-specific that doesn't fit →uns/<tool>/….Sub-issues
converttorecast, and add theadata convertformat framework #21 Rename matrixconvert→recast; add the format-conversion framework*_feature_bc_matrix.h5), import and export #23 10x HDF5 (*_feature_bc_matrix.h5), import and exportread_hdf/read_csv/read_text)CistopicObjectpickles #28 pycistopicCistopicObjectpicklesremove-backgroundoutput h5 #29 CellBenderremove-backgroundh5adata convert#31 Discussion: other formats worth supportingSuggested order: #21 + #22 (framework with one real user) → #23, #24 (they provide the building blocks) → #29, #30 (reuse #23/#24) → #26 → #27 → #25, #28.
Definition of done (every sub-issue)
tests/test_performance.pybounds its operation counts.tests/conftest.py, parametrised over h5ad/zarr2/zarr3 outputs.integrationmarker,uv run --with anndata), and where scanpy has a reader for the format, the result is compared with scanpy's (uv run --with scanpy).docs/COMMANDS.md, the "Format support" table indocs/index.md, the README andCHANGELOG.mdare updated;tests/test_docs_are_accurate.pypasses.Non-goals
No analysis; no R-native formats (
.rds) in the first round — see #31.