Skip to content

Epic: format conversion — adata convert to and from h5ad/zarr #20

Description

@Claptar

Motivation

Data reaches us in many shapes: 10x h5, MTX folders, loom (velocyto, SCope/SCENIC+), CellBender output, Parse split-pipe folders, pycistopic pickles. Today converting them means loading everything into memory with scanpy/anndata. adata already reads and writes the AnnData on-disk spec with plain h5py/zarr and streams everything, so it is a natural home for a streaming, dependency-light converter:

adata convert INPUT -o OUTPUT [--from FMT] [--to FMT] [--zarr-format 2|3]
              [--chunk N] [--layout csr|csc] [--force] [format-specific options]

This epic tracks the shared design; each format has its own sub-issue with a detailed plan.

Naming

adata convert currently changes a matrix's dtype/layout/density. It will be renamed to adata recast (with a deprecation period) so that convert means format conversion — see #21.

Proposed design

Format detection. Sniff content first, extension second, --from overrides:

Signature Format
HDF5 with /matrix/features 10x h5 v3 (and CellBender if /droplet_latents is also present)
HDF5 with /<genome>/genes + /<genome>/indptr 10x h5 v2
HDF5 with /matrix + /row_attrs + /col_attrs loom (SCope loom if MetaData attr / col_attrs/RegulonsAUC)
directory with matrix.mtx[.gz] + barcodes.tsv[.gz] 10x MTX folder
directory with count_matrix.mtx + all_genes.csv Parse split-pipe
.pkl pycistopic (explicit --from pycistopic required)
.h5ad, .zarr AnnData

The target format comes from the -o extension; --to overrides it (e.g. --to 10x-h5, --to 10x-mtx).

Architecture. A new package (e.g. src/adata/formats/foreign/) with one module per format and a small registry. Each reader implements:

  • sniff(path) -> bool
  • plan(path, opts) -> ConversionPlan — shapes, index dtypes, elements to be written, warnings. Runs before any output is created.
  • write(plan, dst: Store)

Writers (export direction) mirror this with plan_export / write_export.

Reuse, don't reinvent: storage.open_store / copy_dataset / copy_store_contents, the elements.write helpers (write_sparse, write_dense, write_string_array, write_categorical, write_dataframe_header, ensure_anndata_skeleton), transpose_sparse_streaming and sparsify from core/convert.py, and the temp-path-then-rename logic from commands/convert.py.

Conventions every converter must follow (same as the existing convert/subset/concat):

  • stream in chunks (--chunk); --in-memory only as an opt-in
  • validate everything before creating the output; a refusal never leaves a half-written store
  • write to a temp path, rename into place
  • refuse lossy or inflating operations unless --force
  • finish with ensure_anndata_skeleton; honour --zarr-format
  • no anndata/scipy/pandas at runtime — pure h5py/zarr wherever the format allows; heavy dependencies only behind optional extras (currently only needed for pycistopic pickles)

Common options: --var-names gene_symbols|gene_ids (scanpy default: symbols), --make-unique (like var_names_make_unique), --dtype (hands off to the recast code path).

Mapping convention: counts → X; extra matrices → layers; per-cell/per-feature attributes → obs/var; multi-column attributes → obsm/varm; anything tool-specific that doesn't fit → uns/<tool>/….

Sub-issues

Suggested order: #21 + #22 (framework with one real user) → #23, #24 (they provide the building blocks) → #29, #30 (reuse #23/#24) → #26 → #27 → #25, #28.

Definition of done (every sub-issue)

  • The converter streams; a test in tests/test_performance.py bounds its operation counts.
  • Fixtures are built with h5py in tests/conftest.py, parametrised over h5ad/zarr2/zarr3 outputs.
  • The output opens in anndata (integration marker, uv run --with anndata), and where scanpy has a reader for the format, the result is compared with scanpy's (uv run --with scanpy).
  • docs/COMMANDS.md, the "Format support" table in docs/index.md, the README and CHANGELOG.md are updated; tests/test_docs_are_accurate.py passes.

Non-goals

No analysis; no R-native formats (.rds) in the first round — see #31.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions