diff --git a/02-collection-layout.md b/02-collection-layout.md index 8101f38..79589fb 100644 --- a/02-collection-layout.md +++ b/02-collection-layout.md @@ -148,8 +148,55 @@ This includes `..` traversal, symlink traversal where the implementation follows symlinks, and absolute paths supplied where a collection-relative path is required. -Implementations MAY reject platform-reserved filenames or characters when a -write operation targets a filesystem where those paths cannot be represented. +### Portable Paths + +A collection is read and written on Linux, macOS, and Windows file systems, +and inside applications such as Obsidian. A path that one platform reads +differently from another can escape the collection, reach a tool's private +or executable state, or fail to be written at all. Every tool MUST therefore +reject, for every record, resource, and file it creates, renames, moves, +replicates, or applies from another tool, a collection-relative path that: + +- is empty, or longer than 1,024 bytes of UTF-8; +- begins with `/` (an absolute or UNC path), or contains `\`; +- has an empty segment (`a//b`, a trailing `/`), a `.` or `..` segment, or a + segment longer than 255 bytes; +- contains any of `<`, `>`, `:`, `"`, `|`, `?`, `*` (`:` also covers drive + prefixes such as `C:x.md` and alternate data streams); +- contains a control character (U+0000–U+001F, U+007F–U+009F); +- contains a character that is invisible or that HFS+ ignores in names: + U+00AD, U+034F, U+115F, U+1160, U+17B4, U+17B5, U+180B–U+180F, + U+200B–U+200F, U+202A–U+202E, U+2060–U+206F, U+3164, U+FE00–U+FE0F, + U+FEFF, U+FFA0, U+FFF0–U+FFF8, U+1BCA0–U+1BCA3, U+1D173–U+1D17A, and + U+E0000–U+E0FFF (so `\u200C.git` cannot name `.git`); +- has a segment ending in `.` or a space, which Windows strips; +- has a segment that is a Windows device name, compared without ASCII case + on the part before its first `.` with trailing spaces removed: `CON`, + `PRN`, `AUX`, `NUL`, `CONIN$`, `CONOUT$`, `CLOCK$`, and `COM` or `LPT` + followed by one of `0`–`9`, `¹`, `²`, `³`; +- has a segment shaped like an NTFS 8.3 short name, which can alias another + name such as `.git`: one to six characters, `~`, a decimal number without a + leading zero, and optionally `.` and up to three characters (`GIT~1`, + `MDBASE~2.TXT`); +- has a segment beginning with `.`: hidden files and tool state such as + `.obsidian`, `.git`, `.vscode`, and `.mdbase` are never collection content + (see Record Discovery); +- has a segment equal to `node_modules`. + +The names `.mdbase` and `node_modules` compare without ASCII case, with U+017F +LATIN SMALL LETTER LONG S read as `s` and U+212A KELVIN SIGN as `k`: the only +characters whose case folding gives a letter of those names. This comparison +is fixed and does not depend on a Unicode version, so the rule gives the same +verdict in every release. + +A tool that reads a collection skips files at such paths as it skips +excluded paths. A write that would create one fails with `invalid_request` +(`details.reason` names the rule) before any write. + +**Provisional (rc.5).** The rule set comes from the +review of the first implementation. Earlier drafts let implementations +reject such paths optionally; replicating a collection across file systems +needs every tool to agree on them. ## Path Equivalence @@ -181,12 +228,26 @@ collides, whatever file system it runs on: Read reports a `path_collision` warning on every record of the group, with `details.paths` listing the group in code-point order. + **Provisional (rc.5).** A tool that keeps records in a log or replicates + them across file systems cannot hold two records with one path key. It + MAY resolve a discovered group instead, by the collision rule below: with + no other order, the path that is smaller in code-point order keeps it, and + each other record moves to its first free suffixed path. It reports the + moves as `record_renamed` and no `path_collision` warning remains. + Path globs (above) remain case-sensitive and match paths as written. **Provisional (rc.5).** Case folding uses the full mappings, so `ß` and `ss` collide. This flags more collisions than some file systems would, never fewer. +**Provisional (rc.5).** NFC and case folding use the data of **Unicode +17.0.0** (`UnicodeData.txt`, `CompositionExclusions.txt`, `CaseFolding.txt`). +Unicode's stability policies keep both stable for assigned characters, but a +character assigned in a later version can fold or compose differently once a +tool knows it. Pinning the version makes every tool compute the same path +keys. A later release names its own version. + ## Path Collisions When a new record would take a path whose path key is already in use, the diff --git a/03-records-and-frontmatter.md b/03-records-and-frontmatter.md index 3a82784..4ece361 100644 --- a/03-records-and-frontmatter.md +++ b/03-records-and-frontmatter.md @@ -210,3 +210,31 @@ data model before JSON Schema validation. Non-JSON YAML values such as NaN, Infinity, binary values, and timestamps with parser-specific objects MUST be handled by the mdbase YAML profile before schema validation or rejected with a clear diagnostic. + +### Scalar Resolution + +Plain (unquoted) scalars resolve with the YAML 1.2 core schema. A tool MUST +read a plain scalar as follows, and MUST NOT apply YAML 1.1 resolution: + +| Plain scalar | Value | +| --- | --- | +| empty, `~`, `null`, `Null`, `NULL` | null | +| `true`, `True`, `TRUE`, `false`, `False`, `FALSE` | boolean | +| `[-+]?[0-9]+`, `0o[0-7]+`, `0x[0-9a-fA-F]+` | integer | +| `[-+]?(\.[0-9]+\|[0-9]+(\.[0-9]*)?)([eE][-+]?[0-9]+)?` | number | +| anything else | string | + +So `yes`, `no`, `on`, and `off` are strings, `0777` is the integer 777, +`1_000` and `1:20` are strings, timestamps such as `2026-10-01` and +`2026-10-01T09:00:00Z` are strings, and `<<` is an ordinary key. Quoted and +block scalars are always strings. `.inf`, `.nan`, and decimal numbers whose +value is not a finite double are outside the JSON data model and are read as +the string as written. + +A writer that has no existing style to keep SHOULD quote a string that a YAML +1.1 parser would read as another type (for example `"yes"` or +`"2026-10-01"`), so tools that still use YAML 1.1 read the same value. + +**Provisional (rc.5).** This table resolves an ambiguity of earlier release +candidates, which asked for "a safe YAML parser" without naming a schema. It +matches Obsidian and the current engines. diff --git a/07-collection-semantics.md b/07-collection-semantics.md index 71ff94b..8db4560 100644 --- a/07-collection-semantics.md +++ b/07-collection-semantics.md @@ -333,6 +333,12 @@ fields. Values are converted to strings without expression evaluation: a string is used as written, and a number or boolean uses its JSON representation. A missing or null value produces `path_value_missing`. +**Provisional (rc.5).** The JSON representation of a number is the one RFC +8785 (JSON Canonicalization Scheme) specifies, which is ECMAScript's +`Number.prototype.toString`: an integer-valued number has no fraction +(`42.0` gives `42`), and exponents are written `1e+21` and `1e-7`. CEL's +`string(double)` uses the same text (Chapter 10). + A placeholder value always stays within one path component. A converted value is invalid, and the operation fails with `path_value_invalid` naming the field, when it: diff --git a/10-cel-profile.md b/10-cel-profile.md index 18bd5d3..15c8003 100644 --- a/10-cel-profile.md +++ b/10-cel-profile.md @@ -175,6 +175,13 @@ The durable runtime profile defines when each workflow expression is evaluated. ## Record Values +**Provisional (rc.5).** CEL maps are unordered, but comprehensions over a +map (`all`, `exists`, `exists_one`, `map`, `filter`) visit its keys, and +`map` and `filter` return lists in that order. Maps iterate in insertion +order: frontmatter maps in the order of their keys in the source, map +literals in the order written. `string(double)` writes the RFC 8785 number +text of Chapter 07 (`1e+21`, `0.3333333333333333`, and `0` for negative zero). + Frontmatter is converted to CEL values after the JSON data-model conversion in Chapter 06: @@ -291,7 +298,9 @@ profile adds two string methods: | `s.lower()` | `s` with every character mapped to lowercase | | `s.upper()` | `s` with every character mapped to uppercase | -Both use the Unicode default full case mappings without locale tailoring, so +Both use the Unicode default full case mappings without locale tailoring, +including the context-dependent final sigma rule, from the same Unicode +version as path keys (Chapter 02, provisionally 17.0.0), so `"Éclair".lower()` is `"éclair"` and `"Straße".upper()` is `"STRASSE"`. A case-insensitive search lowercases the text it searches: @@ -330,7 +339,31 @@ profile keeps that syntax and removes Unicode classes from it. - the flags `i`, `m`, `s`, and `x`, set as `(?flags)` or `(?flags:...)` A pattern that uses anything else is invalid. In particular, Unicode classes -such as `\p{L}` and `\pN`, backreferences, and look-around are invalid. +such as `\p{L}` and `\pN`, backreferences, and look-around are invalid, and +so are constructs that `regex-lite` accepts beyond this list: `\b{start}` and +the other `\b{...}` boundaries, `(?...)` (use `(?P...)`), the +`U` and `R` flags, escapes of letters or digits other than those listed (such +as `\a` and `\u0041`), nested bracket classes, and class set operations +(`&&`, `--`, `~~`). A `{` that does not begin a valid repetition, and a `]` or +`}` outside a class, are invalid too; escape them. `\<` and `\>` are +literal `<` and `>`, as the escape rule above says, although `regex-lite` +itself reads them as word boundaries. + +**Limits.** Every tool accepts and rejects the same patterns: + +- a pattern longer than 8,192 bytes is invalid; +- groups nested deeper than 64 are invalid; +- a repetition count above 1,000 in `{n}`, `{n,}`, or `{n,m}` is invalid, as + in RE2; +- a pattern whose **program size** exceeds 100,000 is invalid. The size is: 1 + for a literal, `.`, a bracket class, or a Perl class; 0 for an anchor or + boundary; the sum for a concatenation; the sum plus 1 per `|` for an + alternation; the inner size plus 1 for a group; twice the operand for `*`, + `+`, and `?`; and the operand times `max(n, m, 1) + 1` for `{n}`, `{n,}`, + and `{n,m}`. + +`regex-lite`'s own size limit depends on the platform's pointer width, so it +cannot be the portable limit. **Semantics.** Matching runs over Unicode scalar values, and a match is unanchored unless the pattern anchors it, as in RE2 and JSON Schema. Among diff --git a/12a-concurrent-edits.md b/12a-concurrent-edits.md index 5305494..f7dcc9b 100644 --- a/12a-concurrent-edits.md +++ b/12a-concurrent-edits.md @@ -128,6 +128,11 @@ The result is a **merged version** and a possibly empty list of **conflicts**. A conflict has a `kind`: `field` for one top-level frontmatter field, which it names in `field`; `frontmatter` for a whole frontmatter block; `body`; or `path`. It carries the base, first, and second values. + +Conflicts are listed in this order: field conflicts in key order (the first +version's keys in its order, then keys only the second version has, in its +order, then keys only the base has, in its order), then a `frontmatter` +conflict, then a `body` conflict, then a `path` conflict. Where a conflict exists, the merged version holds the first version's value. A tool that writes a merged version with conflicts MUST NOT discard the second version's conflicting values silently: how it keeps and surfaces them is @@ -159,7 +164,10 @@ standing for that key's state in each: 3. Otherwise, if `F` equals `B`, the result is `S`: only the second side changed the key. 4. Otherwise both sides changed the key differently, and the key's strategy - decides. + decides. When the matched types declare different strategies for the key + (a `type_conflict`, Chapter 05), the key uses `conflict`: a merge never + fails and never drops a value it cannot decide. Reading the merged record + still reports the `type_conflict`. | Strategy | Result when both sides changed the key differently | | --- | --- | @@ -223,10 +231,21 @@ standing for the three bodies: the result is `S`. 2. **Append-append.** If `F` and `S` both begin with all of `B`, so that both sides only appended text at the end, the result is `B`, then `F`'s - appended text, then `S`'s appended text. When `F`'s appended text is not - empty and does not end with a line terminator, a `\n` is inserted between - the two. Journals, logs, and checklists grow this way, and appending to - them concurrently is not a conflict. + appended text, then `S`'s appended text, joined as follows: + - when `B` is not empty and does not end with a line terminator, and both + appended texts begin with one, the line terminator at the start of + `S`'s text is dropped: `F`'s text already ended `B`'s last line; + - then, when `F`'s appended text is not empty and does not end with a line + terminator, and `S`'s remaining text does not begin with one, a line + terminator is inserted between the two, in the body's line-ending style: + the style of `B`'s first line terminator, or when `B` has none, of the + first line terminator in `F`'s and then `S`'s appended text, and `\n` + when there is none at all. A merge therefore never mixes `\n` and + `\r\n` in a body that used one style. + + So no empty line appears between the two appends. Journals, logs, and + checklists grow this way, and appending to them concurrently is not a + conflict. 3. Otherwise the bodies merge as a three-way line merge (diff3). The lines of `B` that are aligned with unchanged lines in both `F` and `S` divide the bodies into stable regions and changed chunks. For each changed chunk, the @@ -236,13 +255,29 @@ standing for the three bodies: When the body is in conflict, the merged body is `F` as a whole and the conflict carries the three bodies. The alignment of `B` with each side is a -longest common subsequence of lines. - -**Provisional (rc.5).** When several longest common subsequences exist, this -release candidate does not fix which one is used, so two implementations can -split some changed chunks differently. Implementations SHOULD use the Myers -difference algorithm. Conformance fixtures use bodies whose alignment is -unique. +longest common subsequence of lines, chosen as follows, so that every +implementation splits chunks the same way: + +1. Lines common to the start of both sequences are aligned, then lines + common to their end. +2. The remaining middle parts are split at the **middle snake** of Myers' + linear-space algorithm (E. Myers, "An O(ND) Difference Algorithm and Its + Variations", 1986, section 4b), in the formulation of diff-match-patch's + `bisect`: for each edit distance `d` the forward search runs before the + reverse search; on each diagonal a search continues from the neighbouring + diagonal with the larger furthest-reaching value, and from the lower + diagonal (a deletion) when they are equal; the split point is the forward + search's furthest-reaching point on the diagonal where the searches first + overlap. +3. Each part is aligned recursively with these rules. + +The executable model (`scripts/concurrent_edits_model.py`, `_lcs_pairs`) +is the reference for this procedure. + +**Provisional (rc.5).** Earlier drafts left the choice among several longest +common subsequences open. The procedure above is the one the reference model +and the first engine implement. A later release may name a simpler canonical +alignment if one proves as fast. **Provisional (rc.5).** Append-append applies only to appends at the end of the body. Two insertions at the same place inside the body remain a @@ -285,8 +320,12 @@ document records alike. A frontmatter source consists of **top-level entries** and the lines between them. An entry is a line that begins a top-level key at column 0, together -with every following line that belongs to that key's value. Blank lines and -comment lines at column 0 between entries are not part of any entry. +with every following line up to the last line of that key's value, and then +any directly following blank and indented comment lines up to the last +indented comment line. Lines inside the value belong to the entry even when +they are blank or column-0 comments, such as a comment between two `- ` items +of a block sequence at column 0. Blank lines and comment lines at column 0 +after the value are not part of any entry. 1. An entry whose key the write does not change MUST stay byte-identical, including its comments, quoting, indentation, and position. diff --git a/CHANGELOG.md b/CHANGELOG.md index 4449ccf..770daae 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -95,6 +95,23 @@ Bases, seed type upgrades, and saved-view identification. #### Clarified +- Errata from the first implementation (provisional; details in the release + notes): + - YAML 1.2 core scalar resolution; + - which lines belong to a frontmatter entry; + - the pinned Myers body alignment; + - a `type_conflict` during a merge merges as `conflict`; + - conflict order; + - append-append onto a base without a final line break, with separators in + the body's line-ending style; + - portable paths: a normative path-safety rule set (Chapter 02); + - regex `\<`/`\>`, `regex-lite`-only syntax and portable limits; + - Unicode 17.0.0 for path keys and case mappings; + - RFC 8785 number text; + - CEL map iteration order. + + The executable model follows each, and all `cel/cel-profile.yaml` tests now + have ids. - `settings.id_field` has no default; engines that resolve through `id` without configuration are non-conforming. It also serves as the move-detection identity hint. diff --git a/docs/releases/0.3.0-rc.5.md b/docs/releases/0.3.0-rc.5.md index e571d56..94c0413 100644 --- a/docs/releases/0.3.0-rc.5.md +++ b/docs/releases/0.3.0-rc.5.md @@ -143,6 +143,7 @@ An engine that conforms to rc.4 diverges on exactly these existing tests: | Test | rc.4 expectation | rc.5 expectation | | --- | --- | --- | | `links.duplicate_id_ambiguous` (`core/links-and-discovery.yaml`) | resolves to null | resolves to null **and** reports an `ambiguous_link` diagnostic with `details.candidates` | +| `membership.empty_list_persisted` (`core/optional-membership.yaml`, unnamed in rc.4) | the create fails: the v0.2 `write_empty_lists: false` omitted `[]` | the create succeeds: v0.3 has no such setting, so `[]` is written and counts as present | One existing test was renamed without changing its expectation: "collection links enforce validate_exists" (`core/core-collection.yaml`, a `validate` @@ -181,6 +182,44 @@ engine fails them even though rc.4 had no test for the case: | `regex.unicode_class_invalid`, `regex.cel_literal_invalid`, `regex.schema_unicode_class_invalid` | accepts `\p{L}` | | `regex.dot_astral` | a UTF-16 engine without the Unicode flag sees two code units | +## Errata from the first implementation + +Building the first rc.5 engine (mdbase-next `mdbn-core`) turned up points the +draft left ambiguous, or where the executable model disagreed with the prose. +The draft now fixes each one provisionally: + +| Topic | Now specified | Where | +| --- | --- | --- | +| YAML scalar schema | The YAML 1.2 core schema resolves plain scalars; `yes`, `on`, `0777`-as-octal, `1_000`, `1:20` and timestamps are strings or decimals. `.inf`/`.nan` stay strings. The model used YAML 1.1 (`yes` was `true`) and now uses the core schema. | 03 | +| Entry lines | Lines inside a value (including column-0 comments between `- ` items) belong to the entry, and so do indented comment lines directly after it. The model split such an entry at the comment. | 12A | +| Body alignment | The LCS is Myers' linear-space alignment (diff-match-patch `bisect`), so engines split chunks identically. The model used a dynamic-programming LCS. | 12A | +| Type conflict during a merge | The field merges with `conflict`; the merge never fails. The model raised an error. | 12A | +| Conflict order | Field conflicts in key order, then frontmatter, body, path. The model put the path conflict first. | 12A | +| Append onto an unterminated base | No empty line is introduced between the two appends, and an inserted separator uses the body's own line-ending style (CRLF in a CRLF body). | 12A | +| Regex syntax and limits | `\<` and `\>` are literals; `regex-lite`-only syntax is invalid; fixed portable limits on length, nesting, repetition counts and program size. | 10 | +| Unicode version | Path keys and `lower()`/`upper()` use Unicode 17.0.0. | 02, 10 | +| Portable paths | Every tool rejects paths that a platform reads differently: absolute and UNC paths, `\\`, `.`/`..`, `<>:"|?*`, control and HFS+-ignorable characters, trailing dots and spaces, Windows device names, NTFS 8.3 short names, dot-prefixed segments, and `node_modules`. Earlier drafts made this optional. | 02 | +| Number text | Path-pattern values and `string(double)` use RFC 8785 number text (`42.0` → `42`). | 07, 10 | +| CEL map iteration | Comprehensions over maps visit keys in insertion order. | 10 | + +Fixture fixes found by the first implementation: + +| Test | Problem | Fix | +| --- | --- | --- | +| `regex.schema_pattern_ascii` | validated `words/accented.md`, which matches no type (`regex.match_where_ascii`), so no schema applies | validates `words/mixed.md`, which matches the type and fails `pattern` | +| `collection.unique_collection_scope` | `tasks/active/a.md` also has a `slug` duplicate inside the `path_glob` scope | `issues_contain`; the suite README now defines exact versus `_contain` comparison | +| `membership.empty_list_persisted` (was unnamed, *changed*) | depended on `settings.write_empty_lists`, a v0.2 setting that v0.3 does not define | v0.3 writes `[]` and keeps it distinct from a missing key, so the create succeeds | +| `paths.discovered_collision` | contradicts engines that cannot hold two records with one path key (log-based engines void such a state) | Chapter 02 lets such a tool resolve the group by the collision rule; its adapter reports the test as skipped | + +New tests: `merge.column0_comment_inside_sequence`, +`merge.append_unterminated_base`, `merge.append_separator_crlf`, +`merge.body_alignment_pinned`, +`merge.yaml_core_schema_decimal`, `merge.yaml_core_schema_yes`, +`merge.declaration_type_conflict_in_merge`, `regex.angle_brackets_literal`, +`regex.repetition_limit_invalid`, `regex.angle_named_group_invalid`, +`paths.derive_float_integral`, `paths.derive_float_exponent`. The tests in +`cel/cel-profile.yaml` that lacked an `id` now have one. + ## Decided for rc.5 These choices were open during drafting and are now settled. @@ -236,12 +275,15 @@ declared stable. 7. **`union` edge cases.** A missing key or null reads as an empty list; a `tags` string reads as a one-item list; any other non-list is a conflict. An empty result where one side removed the key leaves the key missing. -8. **Body alignment.** rc.5 does not pin which longest common - subsequence aligns lines when several exist; implementations SHOULD use - Myers' algorithm, and the fixtures avoid ambiguous alignments. +8. **Body alignment.** The longest common subsequence is the one Myers' + linear-space algorithm finds, as Chapter 12A specifies step by step; + `merge.body_alignment_pinned` distinguishes it from other choices. 9. **Append-append scope.** Only appends at the end of the body union; two insertions at the same interior position are a conflict. When the first - append lacks a final line break, a `\n` is inserted. + append lacks a final line break, a line break in the body's own style is + inserted; when the base lacks + one and both appends start with a line break, the second's is dropped, + so no empty line appears. 10. **Paths and non-mapping frontmatter in a merge.** A path merges with the `conflict` strategy (one move wins over none; two moves conflict). Frontmatter that is not a mapping merges as one unit. diff --git a/scripts/concurrent_edits_model.py b/scripts/concurrent_edits_model.py index 3ef645e..3b0a19e 100644 --- a/scripts/concurrent_edits_model.py +++ b/scripts/concurrent_edits_model.py @@ -34,18 +34,54 @@ # --------------------------------------------------------------------- YAML -class _StringDateLoader(yaml.SafeLoader): - """Safe loader that keeps timestamps as strings (Chapter 03 YAML profile).""" +class _CoreSchemaLoader(yaml.SafeLoader): + """Safe loader with the YAML 1.2 core schema (Chapter 03 YAML profile). + Plain scalars resolve to null (`~`, `null`, `Null`, `NULL`, empty), + booleans (`true`/`True`/`TRUE`, `false`/`False`/`FALSE`), integers + (decimal, `0o` octal, `0x` hex), and decimal floats. Everything else, + including YAML 1.1 forms such as `yes`, `on`, `0777` as octal, `1_000`, + `1:20`, timestamps, and the `<<` merge key, is a string. `.inf` and `.nan` + are outside the JSON data model and stay strings. + """ -_StringDateLoader.yaml_implicit_resolvers = { - key: [(tag, regexp) for tag, regexp in resolvers if tag != "tag:yaml.org,2002:timestamp"] - for key, resolvers in yaml.SafeLoader.yaml_implicit_resolvers.items() -} + +_CoreSchemaLoader.yaml_implicit_resolvers = {} +for _tag, _pattern, _first in [ + ("tag:yaml.org,2002:null", r"^(?:~|null|Null|NULL|)$", ["~", "n", "N", ""]), + ("tag:yaml.org,2002:bool", r"^(?:true|True|TRUE|false|False|FALSE)$", list("tTfF")), + ("tag:yaml.org,2002:int", r"^(?:[-+]?[0-9]+|0o[0-7]+|0x[0-9a-fA-F]+)$", list("-+0123456789")), + ( + "tag:yaml.org,2002:float", + r"^[-+]?(?:\.[0-9]+|[0-9]+(?:\.[0-9]*)?)(?:[eE][-+]?[0-9]+)?$", + list("-+.0123456789"), + ), +]: + _CoreSchemaLoader.add_implicit_resolver(_tag, re.compile(_pattern), _first) + + +def _construct_core_int(loader: yaml.SafeLoader, node: yaml.Node) -> int: + text = loader.construct_scalar(node) + if text.startswith("0o"): + return int(text[2:], 8) + if text.startswith("0x"): + return int(text[2:], 16) + return int(text, 10) + + +def _construct_core_float(loader: yaml.SafeLoader, node: yaml.Node) -> Any: + text = loader.construct_scalar(node) + value = float(text) + # A decimal number too large for a finite double stays a string. + return text if value in (float("inf"), float("-inf")) else value + + +_CoreSchemaLoader.add_constructor("tag:yaml.org,2002:int", _construct_core_int) +_CoreSchemaLoader.add_constructor("tag:yaml.org,2002:float", _construct_core_float) def load_yaml_text(text: str) -> Any: - return yaml.load(text, Loader=_StringDateLoader) + return yaml.load(text, Loader=_CoreSchemaLoader) # ------------------------------------------------------------------ paths @@ -82,6 +118,36 @@ def allocate_path(requested: str, existing: list[str]) -> str: return suffixed(requested, n) +def es_number(value: float) -> str: + """RFC 8785 / ECMAScript Number.prototype.toString for a finite float.""" + if value == 0: + return "0" + # Shortest round-trip digits. + digits = repr(abs(value)) + if "e" in digits: + mantissa, _, exp = digits.partition("e") + else: + int_part, _, frac = digits.partition(".") + frac = frac.rstrip("0") + whole = (int_part + frac).lstrip("0") + lead_zeros = len(int_part + frac) - len((int_part + frac).lstrip("0")) + n = len(int_part) - lead_zeros + mantissa = whole[0] + ("." + whole[1:] if len(whole) > 1 else "") + exp = str(n - 1) + ds = mantissa.replace(".", "").rstrip("0") or "0" + n = int(exp) + 1 + k = len(ds) + sign = "-" if value < 0 else "" + if k <= n <= 21: + return sign + ds + "0" * (n - k) + if 0 < n <= 21: + return sign + ds[:n] + "." + ds[n:] + if -6 < n <= 0: + return sign + "0." + "0" * (-n) + ds + e = n - 1 + return sign + ds[0] + ("." + ds[1:] if k > 1 else "") + "e" + ("-" if e < 0 else "+") + str(abs(e)) + + class PathError(Exception): def __init__(self, code: str, field_name: str | None = None): super().__init__(code) @@ -106,6 +172,8 @@ def derive_path(pattern: str, frontmatter: dict[str, Any]) -> str: text = "true" if value else "false" elif isinstance(value, str): text = value + elif isinstance(value, float): + text = es_number(value) else: text = json.dumps(value) if text == "" or "/" in text or "\\" in text or "\0" in text or text.startswith("."): @@ -174,33 +242,61 @@ def _split_lines(text: str) -> list[str]: return text.splitlines(keepends=True) +def _is_key_line(line: str) -> bool: + return bool(line.strip()) and line[0] not in " \t#" and not _is_seq_line(line) + + +def _is_seq_line(line: str) -> bool: + return line.startswith("- ") or line.rstrip("\r\n") == "-" + + def _parse_items(fm_lines: list[str]) -> list[Any]: - """Split frontmatter lines into top-level entries and interstitial lines.""" + """Split frontmatter lines into top-level entries and interstitial lines. + + An entry is its key line, every line up to the last line of its value + (indented lines and `- ` items at column 0, with any blank or column-0 + comment lines among them), and then any blank and indented comment lines + up to the last indented comment line, stopping at a column-0 comment. + """ items: list[Any] = [] - current: Entry | None = None - pending_blank: list[str] = [] - for line in fm_lines: - if not line.strip(): - pending_blank.append(line) - continue - continuation = line[0] in " \t" or line.startswith("- ") or line.rstrip("\r\n") == "-" - if continuation and current is not None: - current.lines.extend(pending_blank) - pending_blank = [] - current.lines.append(line) - continue - items.extend(pending_blank) - pending_blank = [] - match = _ENTRY_KEY.match(line) - if match and not line.startswith("#"): - raw_key = match.group(1) - key = load_yaml_text(raw_key) if raw_key[0] in "\"'" else raw_key - current = Entry(key=str(key), lines=[line]) - items.append(current) - else: - current = None + i = 0 + n = len(fm_lines) + while i < n: + line = fm_lines[i] + match = _ENTRY_KEY.match(line) if _is_key_line(line) else None + if not match: items.append(line) - items.extend(pending_blank) + i += 1 + continue + raw_key = match.group(1) + key = load_yaml_text(raw_key) if raw_key[0] in "\"'" else raw_key + j = i + 1 + while j < n and not _is_key_line(fm_lines[j]): + j += 1 + # fm_lines[i+1:j] are candidates. The value ends at its last line that + # is not blank and not a comment. + end = i + 1 + for k in range(i + 1, j): + text = fm_lines[k] + stripped = text.strip() + if stripped and not stripped.startswith("#"): + end = k + 1 + # Then indented comment lines, with blank lines between them. + k = end + while k < j: + text = fm_lines[k] + stripped = text.strip() + if not stripped: + k += 1 + continue + if stripped.startswith("#") and text[0] in " \t": + end = k + 1 + k += 1 + continue + break + items.append(Entry(key=str(key), lines=list(fm_lines[i:end]))) + items.extend(fm_lines[end:j]) + i = j return items @@ -438,6 +534,15 @@ class MergeResult: conflicts: list[dict[str, Any]] +def _line_ending_style(*texts: str) -> str: + """The style of the first line terminator in the texts, in order.""" + for text in texts: + index = text.find("\n") + if index >= 0: + return "\r\n" if index > 0 and text[index - 1] == "\r" else "\n" + return "\n" + + def merge_body(base: str, first: str, second: str) -> str | None: if first == second or second == base: return first @@ -446,29 +551,111 @@ def merge_body(base: str, first: str, second: str) -> str | None: if first.startswith(base) and second.startswith(base): first_tail = first[len(base) :] second_tail = second[len(base) :] - separator = "\n" if first_tail and not first_tail.endswith("\n") else "" + if base and not base.endswith("\n"): + for eol in ("\r\n", "\n"): + if first_tail.startswith(("\r\n", "\n")) and second_tail.startswith(eol): + second_tail = second_tail[len(eol) :] + break + separator = ( + _line_ending_style(base, first_tail, second_tail) + if first_tail and not first_tail.endswith("\n") and not second_tail.startswith(("\n", "\r\n")) + else "" + ) return base + first_tail + separator + second_tail return diff3(_split_lines(base), _split_lines(first), _split_lines(second)) def _lcs_pairs(a: list[str], b: list[str]) -> list[tuple[int, int]]: + """The Chapter 12A alignment: Myers' linear-space algorithm. + + Common prefixes and suffixes are matched first; the rest splits at the + middle snake found by `_bisect` and each part is aligned the same way. + """ + out: list[tuple[int, int]] = [] + stack: list[Any] = [("range", 0, len(a), 0, len(b))] + while stack: + task = stack.pop() + if task[0] == "emit": + out.extend(task[1]) + continue + _, a0, a1, b0, b1 = task + while a0 < a1 and b0 < b1 and a[a0] == b[b0]: + out.append((a0, b0)) + a0 += 1 + b0 += 1 + suffix = [] + while a1 > a0 and b1 > b0 and a[a1 - 1] == b[b1 - 1]: + a1 -= 1 + b1 -= 1 + suffix.append((a1, b1)) + suffix.reverse() + stack.append(("emit", suffix)) + if a0 == a1 or b0 == b1: + continue + split = _bisect(a[a0:a1], b[b0:b1]) + if split is not None: + x, y = split + stack.append(("range", a0 + x, a1, b0 + y, b1)) + stack.append(("range", a0, a0 + x, b0, b0 + y)) + return out + + +def _bisect(a: list[str], b: list[str]) -> tuple[int, int] | None: + """The middle snake of a and b (which differ at both ends): a split point.""" n, m = len(a), len(b) - table = [[0] * (m + 1) for _ in range(n + 1)] - for i in range(n - 1, -1, -1): - for j in range(m - 1, -1, -1): - table[i][j] = table[i + 1][j + 1] + 1 if a[i] == b[j] else max(table[i + 1][j], table[i][j + 1]) - pairs = [] - i = j = 0 - while i < n and j < m: - if a[i] == b[j]: - pairs.append((i, j)) - i += 1 - j += 1 - elif table[i + 1][j] >= table[i][j + 1]: - i += 1 - else: - j += 1 - return pairs + max_d = (n + m + 1) // 2 + offset = max_d + size = 2 * max_d + 2 + v1 = [-1] * size + v2 = [-1] * size + v1[offset + 1] = 0 + v2[offset + 1] = 0 + delta = n - m + front = delta % 2 != 0 + k1start = k1end = k2start = k2end = 0 + for d in range(max_d): + for k1 in range(-d + k1start, d - k1end + 1, 2): + k1o = offset + k1 + if k1 == -d or (k1 != d and v1[k1o - 1] < v1[k1o + 1]): + x1 = v1[k1o + 1] + else: + x1 = v1[k1o - 1] + 1 + y1 = x1 - k1 + while x1 < n and y1 < m and a[x1] == b[y1]: + x1 += 1 + y1 += 1 + v1[k1o] = x1 + if x1 > n: + k1end += 2 + elif y1 > m: + k1start += 2 + elif front: + k2o = offset + delta - k1 + if 0 <= k2o < size and v2[k2o] != -1 and x1 >= n - v2[k2o]: + return x1, y1 + for k2 in range(-d + k2start, d - k2end + 1, 2): + k2o = offset + k2 + if k2 == -d or (k2 != d and v2[k2o - 1] < v2[k2o + 1]): + x2 = v2[k2o + 1] + else: + x2 = v2[k2o - 1] + 1 + y2 = x2 - k2 + while x2 < n and y2 < m and a[n - x2 - 1] == b[m - y2 - 1]: + x2 += 1 + y2 += 1 + v2[k2o] = x2 + if x2 > n: + k2end += 2 + elif y2 > m: + k2start += 2 + elif not front: + k1o = offset + delta - k2 + if 0 <= k1o < size and v1[k1o] != -1: + x1 = v1[k1o] + y1 = offset + x1 - k1o + if x1 >= n - x2: + return x1, y1 + return None def diff3(base: list[str], first: list[str], second: list[str]) -> str | None: @@ -520,17 +707,18 @@ def merge_records( else: content_shortcut = None - # Path merges with the conflict strategy. + # Path merges with the conflict strategy. Its conflict is reported last. + path_conflicts: list[dict[str, Any]] = [] if first_path == second_path or second_path == base_path: path = first_path elif first_path == base_path: path = second_path else: path = first_path - conflicts.append({"kind": "path", "base": base_path, "first": first_path, "second": second_path}) + path_conflicts.append({"kind": "path", "base": base_path, "first": first_path, "second": second_path}) if content_shortcut is not None: - return MergeResult(document=content_shortcut, path=path, conflicts=conflicts) + return MergeResult(document=content_shortcut, path=path, conflicts=path_conflicts) mappings = all(isinstance(d.frontmatter, dict) for d in (base, first, second)) result = parse_document(first_text) @@ -550,7 +738,11 @@ def merge_records( if states_equal(f, b): _take(result, second, key, eol) continue - kind = strategy(record_types, key) + try: + kind = strategy(record_types, key) + except ValueError: + # A type_conflict between declarations: the field conflicts. + kind = "conflict" if kind in ("max", "min"): if f is MISSING or s is MISSING: if f is MISSING: @@ -593,7 +785,7 @@ def merge_records( conflicts.append({"kind": "body"}) body = first.body result.body = body - return MergeResult(document=render(result), path=path, conflicts=conflicts) + return MergeResult(document=render(result), path=path, conflicts=conflicts + path_conflicts) def render_frontmatter(doc: Document) -> str: @@ -706,11 +898,21 @@ class InvalidPattern(Exception): (re.compile(r"\\[pP]"), "Unicode class"), (re.compile(r"\\[1-9]"), "backreference"), (re.compile(r"\(\?(=|!|<=|...)"), + (re.compile(r"\\b\{"), "\\b{...} boundary"), ] +_REPEAT = re.compile(r"(? bool: """Unanchored search with ASCII-only classes and case folding.""" + if len(pattern.encode("utf-8")) > 8192: + raise InvalidPattern("pattern longer than 8192 bytes") + for found in _REPEAT.finditer(pattern): + if int(found.group(1)) > 1000 or (found.group(2) and int(found.group(2)) > 1000): + raise InvalidPattern("repetition count above 1000") for forbidden, what in _FORBIDDEN: if forbidden.search(pattern): raise InvalidPattern(what) diff --git a/tests/v0.3/README.md b/tests/v0.3/README.md index af292ba..5ca2f56 100644 --- a/tests/v0.3/README.md +++ b/tests/v0.3/README.md @@ -67,6 +67,13 @@ conformance claim. `manifest.yaml` records its non-normative `coverage_targets`; verified claims must use `schemas/v0.3/conformance-claim.schema.json` and provide evidence for every claimed profile. +**Comparing issue and diagnostic lists.** `expect.issues` and +`expect.diagnostics` list *every* issue or diagnostic the operation reports, +compared as a set: same length, each expected entry matching a distinct +actual one, order ignored. An expected entry matches when every key it gives +matches; extra keys in the actual entry are ignored. `issues_contain` and +`diagnostics_contain` require only that each listed entry is present. + `input` and `expect` form an adapter-facing semantic assertion DSL. They are not the native API or wire shape. Adapters may normalize a language-specific API into this shape; an implementation's v0.3 operation surface must still use the @@ -165,7 +172,9 @@ are pure functions of their input and the group's `setup.types`, so `update` with `add` and `remove`, and `load_types` to return type-loading diagnostics. Tests that create two files differing only in case, such as `paths.discovered_collision`, need a case-sensitive file system; an adapter on -a case-insensitive one reports them as skipped. +a case-insensitive one reports them as skipped. An adapter for an engine that +resolves discovered collisions by the collision rule (Chapter 02) also reports +`paths.discovered_collision` as skipped. ### `rename` with `update_refs` (since rc.5) diff --git a/tests/v0.3/cel/cel-profile.yaml b/tests/v0.3/cel/cel-profile.yaml index 0c7f5bd..b13bf0b 100644 --- a/tests/v0.3/cel/cel-profile.yaml +++ b/tests/v0.3/cel/cel-profile.yaml @@ -119,6 +119,7 @@ groups: value: true - name: "explicit null is not replaced by read default" + id: cel.explicit_null_not_defaulted operation: evaluate_cel input: path: "tasks/null-status.md" @@ -294,6 +295,7 @@ groups: severity: warning - name: "file host binding exposes path and folder helpers" + id: cel.file_path_folder_helpers operation: evaluate_cel input: path: "tasks/open.md" @@ -303,6 +305,7 @@ groups: value: true - name: "file tag helper uses segment prefix semantics" + id: cel.file_tag_segment_prefix operation: evaluate_cel input: path: "tasks/open.md" @@ -312,6 +315,7 @@ groups: value: true - name: "file body can be filtered without being returned" + id: cel.query_body_filter operation: query input: where: 'file.body.contains("runtime")' @@ -374,6 +378,7 @@ groups: path: tasks/card-001.md tests: - name: "event payload binding works in workflow context" + id: cel.workflow_event_binding operation: evaluate_cel input: context: workflow @@ -408,6 +413,7 @@ groups: value: true - name: "step result binding works in workflow context" + id: cel.workflow_step_binding operation: evaluate_cel input: context: workflow @@ -417,6 +423,7 @@ groups: value: true - name: "workflow input expression object is evaluated but plain strings are not" + id: cel.workflow_input_expression operation: evaluate_workflow_input input: template: diff --git a/tests/v0.3/cel/regex-profile.yaml b/tests/v0.3/cel/regex-profile.yaml index 025efa2..c6bebc4 100644 --- a/tests/v0.3/cel/regex-profile.yaml +++ b/tests/v0.3/cel/regex-profile.yaml @@ -161,6 +161,38 @@ groups: text: abc expect: matches: true + - name: "\\< and \\> are literal angle brackets" + id: regex.angle_brackets_literal + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\$" + text: "" + expect: + matches: true + - name: a repetition count above 1000 is invalid + id: regex.repetition_limit_invalid + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "a{1001}" + text: a + expect: + error: + code: invalid_pattern + - name: "(?...) is invalid; named groups use (?P...)" + id: regex.angle_named_group_invalid + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "(?a)" + text: a + expect: + error: + code: invalid_pattern - name: "\\p{L} is invalid" id: regex.unicode_class_invalid since: "0.3.0-rc.5" @@ -227,6 +259,11 @@ groups: label: café code: café --- + words/mixed.md: | + --- + label: cafe + code: café + --- tests: - name: "CEL matches() uses ASCII-only \\w" id: regex.cel_matches_ascii @@ -257,7 +294,7 @@ groups: covers: [core_read.regex_profile] operation: validate input: - path: words/accented.md + path: words/mixed.md expect: valid: false issues: diff --git a/tests/v0.3/core/optional-membership.yaml b/tests/v0.3/core/optional-membership.yaml index 3eb8c77..0970583 100644 --- a/tests/v0.3/core/optional-membership.yaml +++ b/tests/v0.3/core/optional-membership.yaml @@ -139,13 +139,12 @@ groups: frontmatter: {title: Milk} expect: valid: false - - name: Serialization policy cannot erase membership + - name: An empty list is persisted and establishes membership setup: config: | spec_version: "0.3.0" settings: explicit_type_keys: [] - write_empty_lists: false default_validation: error types: task.md: | @@ -159,11 +158,13 @@ groups: value: {type: object} --- tests: - - name: omitted empty list cannot establish persisted membership + - name: an empty list is written and establishes membership + id: membership.empty_list_persisted + changed: 0.3.0-rc.5 operation: create input: type: task path: no.md frontmatter: {markers: []} expect: - valid: false + valid: true diff --git a/tests/v0.3/core/paths.yaml b/tests/v0.3/core/paths.yaml index 2a2eff1..84b6c8f 100644 --- a/tests/v0.3/core/paths.yaml +++ b/tests/v0.3/core/paths.yaml @@ -110,6 +110,28 @@ groups: title: 42 expect: path: tasks/42.md + - name: "an integral float uses the RFC 8785 number text" + id: paths.derive_float_integral + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: 42.0 + expect: + path: tasks/42.md + - name: "a large float uses the RFC 8785 exponent form" + id: paths.derive_float_exponent + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: 1.0e+21 + expect: + path: tasks/1e+21.md - name: a value containing a slash cannot create a folder id: paths.derive_slash since: 0.3.0-rc.5 diff --git a/tests/v0.3/core/validation-tiers.yaml b/tests/v0.3/core/validation-tiers.yaml index d3ec64c..9fe5769 100644 --- a/tests/v0.3/core/validation-tiers.yaml +++ b/tests/v0.3/core/validation-tiers.yaml @@ -246,7 +246,7 @@ groups: path: "tasks/active/a.md" expect: valid: false - issues: + issues_contain: - code: duplicate_value field: code details: { paths: ["notes/n.md"] } diff --git a/tests/v0.3/merge/merge.yaml b/tests/v0.3/merge/merge.yaml index 0e472d6..2e4aafa 100644 --- a/tests/v0.3/merge/merge.yaml +++ b/tests/v0.3/merge/merge.yaml @@ -1596,6 +1596,216 @@ groups: conflicts: - kind: path path: items/b.md + - name: "a column-0 comment inside a block sequence belongs to the entry" + id: merge.column0_comment_inside_sequence + since: 0.3.0-rc.5 + covers: [merge.merge_format_fidelity] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + reviewers: + - ann + # reviewed in round two + - bo + --- + Body. + first: | + --- + title: Renamed + score: 1 + reviewers: + - ann + # reviewed in round two + - bo + --- + Body. + second: | + --- + title: Item + score: 1 + reviewers: + - ann + # reviewed in round two + - bo + - cy + --- + Body. + expect: + document: | + --- + title: Renamed + score: 1 + reviewers: + - ann + # reviewed in round two + - bo + - cy + --- + Body. + conflicts: [] + - name: "append-append onto a base without a final line break adds no empty line" + id: merge.append_unterminated_base + since: 0.3.0-rc.5 + covers: [merge.append_append_union] + operation: merge_records + input: + path: items/a.md + base: |- + --- + title: Item + score: 1 + --- + a + first: | + --- + title: Item + score: 1 + --- + a + b + second: | + --- + title: Item + score: 1 + --- + a + c + expect: + document: | + --- + title: Item + score: 1 + --- + a + b + c + conflicts: [] + - name: "append-append in a CRLF body inserts a CRLF separator" + id: merge.append_separator_crlf + since: 0.3.0-rc.5 + covers: [merge.append_append_union] + operation: merge_records + input: + path: items/a.md + base: "---\r\ntitle: Item\r\nscore: 1\r\n---\r\na\r\n" + first: "---\r\ntitle: Item\r\nscore: 1\r\n---\r\na\r\nb" + second: "---\r\ntitle: Item\r\nscore: 1\r\n---\r\na\r\nc\r\n" + expect: + document: "---\r\ntitle: Item\r\nscore: 1\r\n---\r\na\r\nb\r\nc\r\n" + conflicts: [] + - name: "the body alignment is the pinned Myers alignment" + id: merge.body_alignment_pinned + since: 0.3.0-rc.5 + covers: [merge.body_three_way] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + - milk + - bread + - eggs + first: | + --- + title: Item + score: 1 + --- + - milk + - bread + second: | + --- + title: Item + score: 1 + --- + - eggs + - bread + expect: + document: | + --- + title: Item + score: 1 + --- + - eggs + - bread + conflicts: [] + - name: "plain scalars use the YAML 1.2 core schema: 0777 is decimal" + id: merge.yaml_core_schema_decimal + since: 0.3.0-rc.5 + covers: [merge.max_min_semantics] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 0777 + --- + Body. + second: | + --- + title: Item + score: 600 + --- + Body. + expect: + document: | + --- + title: Item + score: 0777 + --- + Body. + conflicts: [] + - name: "plain scalars use the YAML 1.2 core schema: yes is a string" + id: merge.yaml_core_schema_yes + since: 0.3.0-rc.5 + covers: [merge.field_three_way, merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 1 + flag: yes + --- + Body. + second: | + --- + title: Item + score: 1 + flag: true + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + flag: yes + --- + Body. + conflicts: + - kind: field + field: flag - name: frontmatter that is not a mapping merges as one unit id: merge.non_mapping_frontmatter since: 0.3.0-rc.5 @@ -1715,3 +1925,30 @@ groups: expect: strategies: note: union + - name: "a type_conflict between declarations merges the field as a conflict" + id: merge.declaration_type_conflict_in_merge + since: 0.3.0-rc.5 + covers: [collection_semantics.merge_declarations, merge.conflict_reporting] + operation: merge_records + input: + path: shared/x.md + base: | + --- + rank: 1 + --- + first: | + --- + rank: 2 + --- + second: | + --- + rank: 3 + --- + expect: + document: | + --- + rank: 2 + --- + conflicts: + - kind: field + field: rank