Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ jobs:
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -e ".[dev,bdc,api]"
pip install -e ".[dev,bdc,netcdf,dissmodel]"

# Coverage XML is produced but not uploaded anywhere yet.
- name: Run tests with coverage
Expand All @@ -61,7 +61,8 @@ jobs:

test-minimal:
# Core install without optional extras: checks that tests depending on
# optional packages (fiona via disscube[bdc], fastapi via disscube[api])
# optional packages (fiona via disscube[bdc], h5py via disscube[netcdf], DisSModel via
# disscube[dissmodel])
# are skipped, not failed.
name: test (core install, no extras)
runs-on: ubuntu-latest
Expand Down
29 changes: 28 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,37 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]

### Changed
- **DisSModel is now optional.** DisSCube no longer depends on it: install
`disscube[dissmodel]` only to hand a cube to a model. `CubeClient.to_lucc_data()`
is replaced by `to_raster_backend()` (same result, needs the extra; raises an
`ImportError` that says how to install it). Nothing in the pipeline or the CLI
needs DisSModel any more, including exports. The package is described as
"Declarative spatial data cubes".
- `to_raster_backend()` / `to_dataset()` raise `ValueError` when a `period`
leaves no variable at all (the old backend came back empty).
- **Removed the experimental HTTP API** (`disscube.api` and the `api` extra). DisSCube is
a library and a CLI; a server belongs in its own package that depends on it.
- **Slimmed `examples/` and migrated real-data cases**: Real-world datasets, bundled GIS assets (~12 MB), and TerraME parity benchmarks were moved to the dedicated [LambdaGeo/disscube-recipes](https://github.com/LambdaGeo/disscube-recipes) repository (`cases/terrame_fill`, `cases/ilha_maranhao`, `cases/prodes_br163`). The core `examples/` directory now contains strictly lightweight, offline, self-contained examples with tests for both the Python API (`01_quickstart.py`, `02_vector_drivers.py`, `03_time_series.py`) and declarative TOML pipelines (`examples/pipelines/quickstart.toml`).

### Added
- `CubeClient.to_dataset()`: the cube as an `xarray.Dataset` — `(y, x)` static and
`(time, y, x)` temporal variables, CRS and transform via rioxarray — the
primary, dependency-free output.
- `CubeClient.export_netcdf()` (extra `disscube[netcdf]`): CF-1.8 netCDF with a
real `time` axis, per-variable attributes (`spec_hash`, `operator`, ...) and
`mask` kept as a variable.
- **Provenance in the exports.** The writer now stores `source_checksum` (the input the
slice came from) with each new variable, and `CubeClient.provenance()` lists, per time
slice, `spec_hash`, `content_hash`, `source_id` and `source_checksum`. netCDF files carry
it as the global JSON attribute `disscube_provenance` (and as plain attributes on
single-slice variables), plus `history`; GeoTIFF bands carry `SPEC_HASH`, `CONTENT_HASH`,
`SOURCE_CHECKSUM` and `SOURCE_ID` tags. Exports made by a pipeline also record
`pipeline_file`, `pipeline_checksum` and `pipeline_name`. Variables derived by earlier
versions lack `source_checksum` until they are derived again.
- `CubeClient.export_geotiff()`: one band per variable, and per year for temporal
ones (`<variable>_<year>`), with `VARIABLE`/`YEAR`/`SPEC_HASH` band tags.
Pipelines and `disscube run/export` write netCDF when the output ends in `.nc`
or `[export] format = "netcdf"`.
- **Declarative derivations.** A `Derivation` names a source, a target grid and
an operator; operators register themselves (`OPERATOR_REGISTRY`) and cover
zonal statistics (`mean`, `std`, `min`, `max`, `sum`, `majority`,
Expand All @@ -38,7 +66,6 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
and `export` (write the derived variables to a
multi-band GeoTIFF straight from the cube).
- `CubeClient.to_lucc_data()` hands a cube to DisSModel as a raster backend.
- An experimental HTTP API, available as the optional `disscube[api]` extra.
- Self-contained examples, the TerraME `Fill` correspondence with a
cell-by-cell parity test suite, and the MkDocs documentation site.
- `CHANGELOG.md`, `CODE_OF_CONDUCT.md`, `CITATION.cff`, and a publish
Expand Down
4 changes: 2 additions & 2 deletions CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@ authors:
- family-names: "Costa"
given-names: "Sérgio Souza"
orcid: "https://orcid.org/0000-0002-0232-4549"
title: "DisSCube: Declarative Spatial Layer for Dynamic Models"
title: "DisSCube: Declarative spatial data cubes"
abstract: "An open-source, declarative engine for constructing cellular spatial data cubes for dynamic modeling, spatial simulation, and environmental analysis."
version: 0.3.0
version: 0.4.0
date-released: "2026-09-29"
url: "https://github.com/DisSModel/disscube"
repository-code: "https://github.com/DisSModel/disscube"
Expand Down
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ pip install -e ".[dev]"
Tests that need optional packages are skipped, not failed, when the package is missing. To run all of them, install the extras too:

```bash
pip install -e ".[dev,bdc,api]"
pip install -e ".[dev,bdc,netcdf,dissmodel]"
```

### 4. Running the Test Suite
Expand Down
43 changes: 15 additions & 28 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,16 +100,23 @@ da = cube.load("forest_pct", grid_id="AC/5km")
print(da.shape) # (rows, cols)
```

### 6. Hand off to DisSModel
### 6. Get the cube out

```python
backend = cube.to_lucc_data(
["forest_pct", "dist_roads"],
grid_id="AC/5km",
period=("2015", "2020"),
)
ds = cube.to_dataset(["forest_pct", "dist_roads"], grid_id="AC/5km", period=("2015", "2020"))
# xarray.Dataset: (y, x) static and (time, y, x) temporal variables, CRS and transform via ds.rio

cube.export_geotiff(["forest_pct"], "forest.tif", grid_id="AC/5km") # one band per variable and year
cube.export_netcdf(["forest_pct"], "cube.nc", grid_id="AC/5km") # CF-1.8; pip install "disscube[netcdf]"
```

Exports carry their provenance: each GeoTIFF band and netCDF variable records the
`spec_hash`, the `content_hash` of the stored data and the `source_checksum` of the
input it came from (`cube.provenance("forest_pct")` lists them per year).

DisSCube does not need DisSModel. To hand a cube to a DisSModel model, install
`disscube[dissmodel]` and use `cube.to_raster_backend(...)`, which returns a `RasterBackend`.

## Pipeline files (TOML)

A whole data preparation — grid, sources, derived variables — can be declared
Expand Down Expand Up @@ -231,7 +238,6 @@ disscube/
├── pipeline/ Pipeline execution & planning (schema, runner) + internal stages
├── catalog/ CatalogStore (Protocol) + SQLite and JSON implementations
├── storage.py AssetStore (fsspec — local and S3)
├── api/ Experimental HTTP API (optional `api` extra)
├── cli.py `disscube validate` / `disscube run` / `disscube export`
├── sources/ Adapters that bring external data in as SpatialSources,
│ │ each with a checksum and a provenance.json sidecar
Expand Down Expand Up @@ -263,25 +269,6 @@ class WeightedMeanOperator(Operator):

The operator is registered automatically and accepted by `Derivation` / `SpatialDerivation` with no other change.

## Experimental: HTTP API

`disscube.api` exposes the catalog over HTTP for **remote orchestration**: registering grids and sources, triggering derivations and querying what has been derived. It is experimental and ships as an optional extra:

```bash
pip install -e ".[api]"
DISSCUBE_CATALOG=./catalog.db DISSCUBE_STORE=./data/ uvicorn disscube.api.app:app
```

| Endpoint | Purpose |
|---|---|
| `GET` / `POST /grids` | List / register `GridSpec`s |
| `GET` / `POST /sources` | List / register `SpatialSource`s |
| `POST /derive` | Run a `SpatialDerivation` (errors → HTTP 400) |
| `GET /catalog?grid=&role=` | List derived variables |
| `GET /variables/{id}` | Metadata of one derived variable, including its Zarr `asset_url` |

The API does **not** serve raster data. Models load derived variables in-process with `CubeClient.load()` / `CubeClient.to_lucc_data()`, reading the same Zarr store (local, or S3 via fsspec) that the API writes to. Interactive docs are available at `/docs` once the server is running.

## Known limitations

The limitations below are scope decisions for the current version, not bugs. They are documented so that users and reviewers understand what is implemented versus what is planned.
Expand Down Expand Up @@ -311,9 +298,9 @@ If you use DisSCube in your research, dynamic modeling, or spatial data pipeline
```bibtex
@software{costa_disscube_2026,
author = {Costa, S{\'e}rgio Souza},
title = {{DisSCube: Declarative Spatial Layer for Dynamic Models}},
title = {{DisSCube: Declarative spatial data cubes}},
year = {2026},
version = {0.3.0},
version = {0.4.0},
url = {https://github.com/DisSModel/disscube}
}
```
Expand Down
5 changes: 0 additions & 5 deletions disscube/api/__init__.py

This file was deleted.

113 changes: 0 additions & 113 deletions disscube/api/app.py

This file was deleted.

6 changes: 3 additions & 3 deletions disscube/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,14 +29,14 @@ def main(argv: list[str] | None = None) -> int:
p_run.add_argument("file")
p_run.add_argument("--workspace", help="output folder (default: the file's 'workspace', "
"else a folder named after the file)")
p_run.add_argument("--output", "-o", help="export derived variables to a multi-band GeoTIFF")
p_run.add_argument("--output", "-o", help="export derived variables (GeoTIFF; netCDF if the file ends in .nc)")
p_run.add_argument("--dry-run", action="store_true", help="simulate plan execution without downloading or computing")
p_run.add_argument("--json", action="store_true", help="output execution report as JSON")
p_run.add_argument("-v", "--verbose", action="store_true", help="log each step")

p_exp = sub.add_parser("export", help="export derived variables from an existing data cube to GeoTIFF")
p_exp = sub.add_parser("export", help="export derived variables from an existing data cube to GeoTIFF or netCDF")
p_exp.add_argument("file", help="pipeline TOML file")
p_exp.add_argument("--output", "-o", required=True, help="output GeoTIFF file path (e.g. data/cellspace.tif)")
p_exp.add_argument("--output", "-o", required=True, help="output file: GeoTIFF, or netCDF if it ends in .nc (e.g. data/cellspace.tif)")
p_exp.add_argument("--workspace", help="workspace folder (default: data/cube or from pipeline)")
p_exp.add_argument("--variables", nargs="*", help="specific variables to export (default: all derived)")
p_exp.add_argument("--json", action="store_true", help="output export result as JSON")
Expand Down
Loading
Loading