From dd79bea645a7861d778e477a456a5670cca9fa0f Mon Sep 17 00:00:00 2001 From: Ray Date: Wed, 30 Sep 2026 14:43:54 +0800 Subject: [PATCH 1/2] Remove docs/naming-rules.md The naming contract lives in tests/fixtures/naming-v1.json and pageindex/naming.py. --- docs/naming-rules.md | 63 -------------------------------------------- 1 file changed, 63 deletions(-) delete mode 100644 docs/naming-rules.md diff --git a/docs/naming-rules.md b/docs/naming-rules.md deleted file mode 100644 index c9f97d1a5..000000000 --- a/docs/naming-rules.md +++ /dev/null @@ -1,63 +0,0 @@ -# PageIndex naming rules v1 - -Chat, Compute and the Python SDK share this contract. The byte-identical -`naming-v1.json` fixtures run in Vitest and pytest; update all three copies -when changing the contract. Names preserve case. Existing duplicate-name -scopes and database comparison behavior are unchanged. - -## New names - -- Apply Unicode NFKC, then normalize quote variants as in Chat (`‘`, `’`, - `ʼ` to an apostrophe; `“`, `”` to a double quote). Collapse whitespace using - the JavaScript whitespace set and trim leading/trailing whitespace. - This order is idempotent, including characters that expand into quotes. -- Allow Unicode text, emoji (including joined emoji), and interior spaces. -- Disallow `/ \ : * ? " < > |`, C0/C1 controls including DEL, and Unicode - line/paragraph separators. Check controls before whitespace normalization. -- Disallow dot-only names, trailing periods, and case-insensitive Windows - device names CON, PRN, AUX, NUL, COM0–9 and LPT0–9, also with extensions. -- Before resolving duplicates, names including extensions and truncation - hashes must fit **180 UTF-8 bytes**. Automatically added short numeric - collision suffixes may exceed this budget by the suffix length. Do not - split a Unicode code point when truncating. - -## Upload allocation - -Replace disallowed characters with `_`, remove trailing periods, prefix -reserved device names with `_` (preserving the extension), and use `untitled` -when nothing remains. Replace unpaired surrogates with U+FFFD. - -Shorten overlong names with `_` plus the first eight hexadecimal characters -of the MD5 of the cleaned name. Preserve the extension where possible; if the -extension itself exceeds the budget, shorten it too. MD5 is only a stable -label here, not a security mechanism. A duplicate adds `_1`, `_2`, etc. -before the extension. Allocators may shorten again to stay within the -180-byte budget; a short numeric suffix exceeding that budget is also allowed. -Each backend keeps its existing collision-attempt limit. -The allocating backend returns the actual final name alongside its upload -URL or document ID. Chat and cloud SDK callers use that response unchanged. -The local SDK is its own backend and performs the same allocation locally. -The cloud SDK applies the same idempotent sanitization before multipart -encoding so header escaping cannot alter the name. The backend still -validates the name and allocates the final collision suffix. - -## Folder creation and rename - -Clients validate without rewriting the request. The backend normalizes -Unicode and spaces, then rejects invalid characters, reserved names -or excessive byte length with an actionable error. Never turn `/Research/` -into `Research`. Check duplicates using the normalized name before saving. -ZIP import reports invalid folder entries; it does not silently rename their -path components. Generated ZIP root folder names follow the upload rules. - -## Reading and rollout - -Read assigned names literally. Apply no new cleaning, truncation or case -folding to persisted names or to a final upload name submitted for processing. -At the processing boundary, still reject path separators, controls and `.`/`..` -so the name remains one path component. Existing names are not migrated. - -Deploy Compute first (through dev verification), then Chat and the SDK. -Compute's internal `/files/upload-url` response becomes -`{ "url": "...", "headers": {}, "name": "final-name.pdf" }` for both S3 and -Azure. Consumers must use `name`, not reconstruct it from the input or URL. From 2ce382ea838ba83fa77c0a988230c7219a4b6e72 Mon Sep 17 00:00:00 2001 From: Ray Date: Wed, 30 Sep 2026 15:04:31 +0800 Subject: [PATCH 2/2] Move example results next to the documents they index --- .../results/2023-annual-report-truncated_structure.json | 0 .../{documents => }/results/2023-annual-report_structure.json | 0 examples/{documents => }/results/PRML_structure.json | 0 .../Regulation Best Interest_Interpretive release_structure.json | 0 .../results/Regulation Best Interest_proposed rule_structure.json | 0 examples/{documents => }/results/earthmover_structure.json | 0 examples/{documents => }/results/four-lectures_structure.json | 0 examples/{documents => }/results/q1-fy25-earnings_structure.json | 0 8 files changed, 0 insertions(+), 0 deletions(-) rename examples/{documents => }/results/2023-annual-report-truncated_structure.json (100%) rename examples/{documents => }/results/2023-annual-report_structure.json (100%) rename examples/{documents => }/results/PRML_structure.json (100%) rename examples/{documents => }/results/Regulation Best Interest_Interpretive release_structure.json (100%) rename examples/{documents => }/results/Regulation Best Interest_proposed rule_structure.json (100%) rename examples/{documents => }/results/earthmover_structure.json (100%) rename examples/{documents => }/results/four-lectures_structure.json (100%) rename examples/{documents => }/results/q1-fy25-earnings_structure.json (100%) diff --git a/examples/documents/results/2023-annual-report-truncated_structure.json b/examples/results/2023-annual-report-truncated_structure.json similarity index 100% rename from examples/documents/results/2023-annual-report-truncated_structure.json rename to examples/results/2023-annual-report-truncated_structure.json diff --git a/examples/documents/results/2023-annual-report_structure.json b/examples/results/2023-annual-report_structure.json similarity index 100% rename from examples/documents/results/2023-annual-report_structure.json rename to examples/results/2023-annual-report_structure.json diff --git a/examples/documents/results/PRML_structure.json b/examples/results/PRML_structure.json similarity index 100% rename from examples/documents/results/PRML_structure.json rename to examples/results/PRML_structure.json diff --git a/examples/documents/results/Regulation Best Interest_Interpretive release_structure.json b/examples/results/Regulation Best Interest_Interpretive release_structure.json similarity index 100% rename from examples/documents/results/Regulation Best Interest_Interpretive release_structure.json rename to examples/results/Regulation Best Interest_Interpretive release_structure.json diff --git a/examples/documents/results/Regulation Best Interest_proposed rule_structure.json b/examples/results/Regulation Best Interest_proposed rule_structure.json similarity index 100% rename from examples/documents/results/Regulation Best Interest_proposed rule_structure.json rename to examples/results/Regulation Best Interest_proposed rule_structure.json diff --git a/examples/documents/results/earthmover_structure.json b/examples/results/earthmover_structure.json similarity index 100% rename from examples/documents/results/earthmover_structure.json rename to examples/results/earthmover_structure.json diff --git a/examples/documents/results/four-lectures_structure.json b/examples/results/four-lectures_structure.json similarity index 100% rename from examples/documents/results/four-lectures_structure.json rename to examples/results/four-lectures_structure.json diff --git a/examples/documents/results/q1-fy25-earnings_structure.json b/examples/results/q1-fy25-earnings_structure.json similarity index 100% rename from examples/documents/results/q1-fy25-earnings_structure.json rename to examples/results/q1-fy25-earnings_structure.json