Skip to content

Convert: MTX and 10x MTX folders, import and export (streaming) #24

Description

@Claptar

Part of #20. Depends on #21. Equivalent of scanpy.read_10x_mtx / read_mtx.

Import

  • 10x folders: matrix.mtx[.gz] + barcodes.tsv[.gz] + features.tsv[.gz] (or genes.tsv for Cell Ranger 2).
  • STARsolo Solo.out/Gene/{filtered,raw} (same layout; features.tsv without the feature-type column).
  • STARsolo velocyto output (Solo.out/Velocyto/raw/{spliced,unspliced,ambiguous}.mtx) → layers.
  • A bare .mtx file, with optional --obs-names-file / --var-names-file.
  • Orientation: 10x/STARsolo are genes × cells (transposed on read); --transpose / --no-transpose override.

Make it streaming. formats/sparse._read_mtx currently builds a Python list of every entry — unusable for large files. Replace with:

  1. pass 1: stream the file (gzip as a stream), count nnz per output row → indptr
  2. pass 2: scatter entries into preallocated on-disk indices/data, buffered in chunks
  3. sort indices within rows only if the file isn't already ordered

The same reader should back adata import sparse, which benefits too.

Export

h5ad → 10x v3 folder with gzipped files: extend formats/sparse.export_mtx and write barcodes.tsv.gz / features.tsv.gz from obs / var. --layer picks the matrix.

Tests

  • Fixtures for Cell Ranger 2, Cell Ranger 3, STARsolo (incl. velocyto) and a bare mtx.
  • Compare with scanpy.read_10x_mtx / read_mtx (integration marker).
  • Perf guard bounding reads/writes for the two-pass import.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions