diff --git a/doc/tutorials/15.indexing-ctables.ipynb b/doc/tutorials/15.indexing-ctables.ipynb index 11474cc73..ece21c3b1 100644 --- a/doc/tutorials/15.indexing-ctables.ipynb +++ b/doc/tutorials/15.indexing-ctables.ipynb @@ -14,10 +14,11 @@ "\n", "1. Creating an index on a CTable column\n", "2. Querying with an index (automatic)\n", - "3. Stale detection and automatic scan fallback\n", - "4. Rebuilding and dropping indexes\n", - "5. Persistent tables: indexes survive close/reopen\n", - "6. Views and indexes\n" + "3. Automatic SUMMARY indexes\n", + "4. Stale detection and automatic scan fallback\n", + "5. Rebuilding and dropping indexes\n", + "6. Persistent tables: indexes survive close/reopen\n", + "7. Views and indexes\n" ] }, { @@ -32,6 +33,7 @@ }, { "cell_type": "code", + "execution_count": 1, "id": "b23746ca", "metadata": { "ExecuteTime": { @@ -39,8 +41,22 @@ "start_time": "2026-05-21T09:38:19.512234Z" } }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Table: 500 rows\n" + ] + } + ], "source": [ "import dataclasses\n", + "import shutil\n", + "import statistics\n", + "import tempfile\n", + "import time\n", + "from pathlib import Path\n", "\n", "import numpy as np\n", "\n", @@ -61,17 +77,7 @@ " t.append([i, 15.0 + rng.random() * 25, int(rng.integers(0, 4))])\n", "\n", "print(f\"Table: {N} rows\")" - ], - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Table: 500 rows\n" - ] - } - ], - "execution_count": 1 + ] }, { "cell_type": "markdown", @@ -94,6 +100,7 @@ }, { "cell_type": "code", + "execution_count": 2, "id": "d9411a66072ed12c", "metadata": { "ExecuteTime": { @@ -101,10 +108,6 @@ "start_time": "2026-05-21T09:38:20.518592Z" } }, - "source": [ - "t.add_computed_column(\"temperature_f\", \"temperature * 9 / 5 + 32\")\n", - "print(t.select([\"sensor_id\", \"temperature\", \"temperature_f\"]).head(3))" - ], "outputs": [ { "name": "stdout", @@ -119,7 +122,10 @@ ] } ], - "execution_count": 2 + "source": [ + "t.add_computed_column(\"temperature_f\", \"temperature * 9 / 5 + 32\")\n", + "print(t.select([\"sensor_id\", \"temperature\", \"temperature_f\"]).head(3))" + ] }, { "cell_type": "markdown", @@ -135,6 +141,7 @@ }, { "cell_type": "code", + "execution_count": 3, "id": "2ac1f281", "metadata": { "ExecuteTime": { @@ -142,13 +149,6 @@ "start_time": "2026-05-21T09:38:20.550183Z" } }, - "source": [ - "idx = t.create_index(\"sensor_id\")\n", - "print(idx)\n", - "print(\"stale?\", idx.stale)\n", - "print(\"storage stats:\", idx.storage_stats())\n", - "print(\"all indexes:\", t.indexes)" - ], "outputs": [ { "name": "stdout", @@ -161,7 +161,13 @@ ] } ], - "execution_count": 3 + "source": [ + "idx = t.create_index(\"sensor_id\")\n", + "print(idx)\n", + "print(\"stale?\", idx.stale)\n", + "print(\"storage stats:\", idx.storage_stats())\n", + "print(\"all indexes:\", t.indexes)" + ] }, { "cell_type": "markdown", @@ -176,6 +182,7 @@ }, { "cell_type": "code", + "execution_count": 4, "id": "dcc2dc87", "metadata": { "ExecuteTime": { @@ -183,22 +190,142 @@ "start_time": "2026-05-21T09:38:20.597339Z" } }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Rows sensor_id > 450: 49\n", + "sensor_ids: [451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499]\n" + ] + } + ], "source": [ "result = t.where(t[\"sensor_id\"] > 450)\n", "print(\"Rows sensor_id > 450:\", len(result))\n", "print(\"sensor_ids:\", sorted(int(v) for v in result[\"sensor_id\"][:]))" - ], + ] + }, + { + "cell_type": "markdown", + "id": "d9fa565cab114c0d", + "metadata": {}, + "source": [ + "## Automatic SUMMARY indexes\n", + "\n", + "A SUMMARY index stores the minimum and maximum value in each compressed block. `Column.min()`/`max()` (and `argmin()`/`argmax()` inside `group_by()`) then answer from those precomputed summaries instead of decompressing the column at all. \n", + "\n", + "The same SUMMARY indexes can also let a selective `where()` query skip whole blocks, but only when the column’s values are ordered or clustered enough that a block’s min/max range can exclude the predicate entirely. With independently random data every block spans nearly the full value range and there is nothing to skip — so the `min()/max()` speedup is the one you can always count on.\n", + "\n", + "SUMMARY indexes are built automatically for eligible numeric columns on the first close of a `CTable`. Pass `create_summary_index=False` in the `CTable` constructor to opt out.\n", + "\n", + "The following example creates two identical, time-ordered sensor datasets. Closing the first table builds SUMMARY indexes; the second disables them so we can compare the same query with a predicate that is well suited to benefit from the SUMMARY index." + ] + }, + { + "cell_type": "code", + "execution_count": 87, + "id": "71af0133c2264bda", + "metadata": { + "execution": { + "iopub.execute_input": "2026-09-30T16:04:25.421274Z", + "iopub.status.busy": "2026-09-30T16:04:25.420795Z", + "iopub.status.idle": "2026-09-30T16:04:25.616131Z", + "shell.execute_reply": "2026-09-30T16:04:25.615359Z" + } + }, "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ - "Rows sensor_id > 450: 49\n", - "sensor_ids: [451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499]\n" + "All 3,111 rows matched across both tables for predicate: timestamp >= 9980000 and temperature > 25\n", + "Mean runtime over 10 runs:\n", + "- Without SUMMARY: 23.518 ms\n", + "- With SUMMARY: 12.173 ms\n", + "- Speedup: 1.93x\n" ] } ], - "execution_count": 4 + "source": [ + "# Generate the dataset\n", + "\n", + "\n", + "@dataclasses.dataclass\n", + "class SensorReading:\n", + " timestamp: int = blosc2.field(blosc2.int64())\n", + " sensor_id: int = blosc2.field(blosc2.int32())\n", + " temperature: float = blosc2.field(blosc2.float32())\n", + "\n", + "\n", + "N_READINGS = 10_000_000\n", + "timestamps = np.arange(N_READINGS, dtype=np.int64)\n", + "benchmark_rng = np.random.default_rng(42)\n", + "readings = {\n", + " \"timestamp\": timestamps,\n", + " \"sensor_id\": (timestamps % 1_000).astype(np.int32),\n", + " \"temperature\": (20 + benchmark_rng.normal(0, 5, N_READINGS)).astype(np.float32),\n", + "}\n", + "\n", + "summary_tmpdir = Path(tempfile.mkdtemp())\n", + "summary_path = summary_tmpdir / \"readings-summary.b2d\"\n", + "no_summary_path = summary_tmpdir / \"readings-scan.b2d\"\n", + "\n", + "for path, create_summary_index in ((summary_path, True), (no_summary_path, False)):\n", + " with blosc2.CTable(\n", + " SensorReading,\n", + " urlpath=str(path),\n", + " mode=\"w\",\n", + " expected_size=N_READINGS,\n", + " create_summary_index=create_summary_index,\n", + " ) as readings_table:\n", + " readings_table.extend(readings)\n", + "\n", + "del readings, timestamps\n", + "\n", + "# Run the benchmark\n", + "\n", + "\n", + "def mean_ms(func, num_samples):\n", + " samples = []\n", + " for _ in range(num_samples):\n", + " start = time.perf_counter()\n", + " func()\n", + " samples.append((time.perf_counter() - start) * 1e3)\n", + " return statistics.mean(samples)\n", + "\n", + "\n", + "ts_cutoff = N_READINGS - 20_000\n", + "temp_cutoff = 25\n", + "num_samples = 10\n", + "with (\n", + " blosc2.open(str(summary_path), mode=\"r\") as summary_readings,\n", + " blosc2.open(str(no_summary_path), mode=\"r\") as no_summary_readings,\n", + "):\n", + "\n", + " def predicate(t):\n", + " return (t[\"timestamp\"] >= ts_cutoff) & (t[\"temperature\"] > temp_cutoff)\n", + "\n", + " summary_result = summary_readings.where(predicate(summary_readings))[\"timestamp\"][:]\n", + " no_summary_result = no_summary_readings.where(predicate(no_summary_readings))[\"timestamp\"][:]\n", + " np.testing.assert_array_equal(summary_result, no_summary_result)\n", + " print(\n", + " f\"All {len(summary_result):,} rows matched across both tables for predicate: timestamp >= {ts_cutoff} and temperature > {temp_cutoff}\"\n", + " )\n", + "\n", + " summary_ms = mean_ms(\n", + " lambda: summary_readings.where(predicate(summary_readings)), num_samples=num_samples\n", + " )\n", + " no_summary_ms = mean_ms(\n", + " lambda: no_summary_readings.where(predicate(no_summary_readings)), num_samples=num_samples\n", + " )\n", + " print(f\"Mean runtime over {num_samples} runs:\")\n", + " print(f\"- Without SUMMARY: {no_summary_ms:.3f} ms\")\n", + " print(f\"- With SUMMARY: {summary_ms:.3f} ms\")\n", + " print(f\"- Speedup: {no_summary_ms / summary_ms:.2f}x\")\n", + "\n", + "shutil.rmtree(summary_tmpdir, ignore_errors=True)" + ] }, { "cell_type": "markdown", @@ -214,6 +341,7 @@ }, { "cell_type": "code", + "execution_count": 5, "id": "b0132381", "metadata": { "ExecuteTime": { @@ -221,16 +349,6 @@ "start_time": "2026-05-21T09:38:20.618568Z" } }, - "source": [ - "t.append([9999, 30.0, 1]) # any mutation marks indexes stale\n", - "\n", - "idx = t.index(\"sensor_id\")\n", - "print(\"stale after append?\", idx.stale)\n", - "\n", - "# Query still works — scan fallback\n", - "result_stale = t.where(t[\"sensor_id\"] == 9999)\n", - "print(\"Found row:\", len(result_stale))" - ], "outputs": [ { "name": "stdout", @@ -241,7 +359,16 @@ ] } ], - "execution_count": 5 + "source": [ + "t.append([9999, 30.0, 1]) # any mutation marks indexes stale\n", + "\n", + "idx = t.index(\"sensor_id\")\n", + "print(\"stale after append?\", idx.stale)\n", + "\n", + "# Query still works — scan fallback\n", + "result_stale = t.where(t[\"sensor_id\"] == 9999)\n", + "print(\"Found row:\", len(result_stale))" + ] }, { "cell_type": "markdown", @@ -257,6 +384,7 @@ }, { "cell_type": "code", + "execution_count": 6, "id": "dc4d2897", "metadata": { "ExecuteTime": { @@ -264,13 +392,6 @@ "start_time": "2026-05-21T09:38:20.644706Z" } }, - "source": [ - "idx = t.rebuild_index(\"sensor_id\")\n", - "print(\"stale after rebuild?\", idx.stale)\n", - "\n", - "result_rebuilt = t.where(t[\"sensor_id\"] == 9999)\n", - "print(\"Found row via rebuilt index:\", len(result_rebuilt))" - ], "outputs": [ { "name": "stdout", @@ -281,7 +402,13 @@ ] } ], - "execution_count": 6 + "source": [ + "idx = t.rebuild_index(\"sensor_id\")\n", + "print(\"stale after rebuild?\", idx.stale)\n", + "\n", + "result_rebuilt = t.where(t[\"sensor_id\"] == 9999)\n", + "print(\"Found row via rebuilt index:\", len(result_rebuilt))" + ] }, { "cell_type": "markdown", @@ -295,6 +422,7 @@ }, { "cell_type": "code", + "execution_count": 7, "id": "e1583b4f", "metadata": { "ExecuteTime": { @@ -302,10 +430,6 @@ "start_time": "2026-05-21T09:38:20.703078Z" } }, - "source": [ - "t.drop_index(\"sensor_id\")\n", - "print(\"Indexes after drop:\", t.indexes)" - ], "outputs": [ { "name": "stdout", @@ -315,7 +439,10 @@ ] } ], - "execution_count": 7 + "source": [ + "t.drop_index(\"sensor_id\")\n", + "print(\"Indexes after drop:\", t.indexes)" + ] }, { "cell_type": "markdown", @@ -329,6 +456,7 @@ }, { "cell_type": "code", + "execution_count": 8, "id": "85d42133", "metadata": { "ExecuteTime": { @@ -336,11 +464,18 @@ "start_time": "2026-05-21T09:38:20.724202Z" } }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Created: Index(kind='bucket', col_name='sensor_id', name='__self__', stale=False)\n", + "Storage stats: (12583745, 7788, 1615.7864663585003)\n", + "Rows > 280 (before close): 19\n" + ] + } + ], "source": [ - "import shutil\n", - "import tempfile\n", - "from pathlib import Path\n", - "\n", "tmpdir = Path(tempfile.mkdtemp())\n", "path = str(tmpdir / \"sensors.b2d\")\n", "\n", @@ -359,22 +494,11 @@ "# Query before close\n", "r1 = pt.where(pt[\"sensor_id\"] > 280)\n", "print(\"Rows > 280 (before close):\", len(r1))" - ], - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Created: Index(kind='bucket', col_name='sensor_id', name='__self__', stale=False)\n", - "Storage stats: (12583745, 7788, 1615.7864663585003)\n", - "Rows > 280 (before close): 19\n" - ] - } - ], - "execution_count": 8 + ] }, { "cell_type": "code", + "execution_count": 9, "id": "149ddba5", "metadata": { "ExecuteTime": { @@ -382,6 +506,18 @@ "start_time": "2026-05-21T09:38:21.739759Z" } }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "Indexes after reopen: [Index(kind='bucket', col_name='sensor_id', name='__self__', stale=False)]\n", + "Storage stats after reopen: (12583745, 7788, 1615.7864663585003)\n", + "Rows > 280 (after reopen): 19\n", + "Results match ✓\n" + ] + } + ], "source": [ "# Close and reopen — catalog is preserved\n", "del pt\n", @@ -399,20 +535,7 @@ "print(\"Results match ✓\")\n", "\n", "shutil.rmtree(tmpdir, ignore_errors=True)" - ], - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "Indexes after reopen: [Index(kind='bucket', col_name='sensor_id', name='__self__', stale=False)]\n", - "Storage stats after reopen: (12583745, 7788, 1615.7864663585003)\n", - "Rows > 280 (after reopen): 19\n", - "Results match ✓\n" - ] - } - ], - "execution_count": 9 + ] }, { "cell_type": "markdown", @@ -427,6 +550,7 @@ }, { "cell_type": "code", + "execution_count": 10, "id": "83db418b", "metadata": { "ExecuteTime": { @@ -434,6 +558,17 @@ "start_time": "2026-05-21T09:38:21.777930Z" } }, + "outputs": [ + { + "name": "stdout", + "output_type": "stream", + "text": [ + "View type: CTable\n", + "create_index on view: Cannot create an index on a view.\n", + "drop_index on view: Cannot drop an index from a view.\n" + ] + } + ], "source": [ "t2 = blosc2.CTable(Measurement)\n", "for i in range(50):\n", @@ -452,19 +587,7 @@ " view.drop_index(\"sensor_id\")\n", "except ValueError as e:\n", " print(\"drop_index on view:\", e)" - ], - "outputs": [ - { - "name": "stdout", - "output_type": "stream", - "text": [ - "View type: CTable\n", - "create_index on view: Cannot create an index on a view.\n", - "drop_index on view: Cannot drop an index from a view.\n" - ] - } - ], - "execution_count": 10 + ] }, { "cell_type": "markdown", @@ -485,25 +608,13 @@ "\n", "Key behaviours:\n", "\n", + "- **SUMMARY indexes** store per-block min/max for eligible numeric columns and are automatically built on first close.\n", "- **Mutations** (`append`, `extend`, `setitem`, `assign`, `sort_by`, `compact`) mark indexes stale.\n", "- **Stale indexes** trigger automatic scan fallback — no user intervention needed.\n", "- **Persistent indexes** survive table close and reopen.\n", "- **OPSI indexes** are tunable iterative-ordering indexes for exact filtering; use `FULL` for completely sorted ordered reuse.\n", "- **Views** cannot own indexes; only root tables can.\n" ] - }, - { - "cell_type": "code", - "id": "363827fec805190a", - "metadata": { - "ExecuteTime": { - "end_time": "2026-05-21T09:38:21.910123Z", - "start_time": "2026-05-21T09:38:21.906401Z" - } - }, - "source": [], - "outputs": [], - "execution_count": 10 } ], "metadata": {