From 627b908fb94c130eb6b061e12256858705139053 Mon Sep 17 00:00:00 2001 From: callumalpass Date: Sun, 4 Oct 2026 14:29:32 +1100 Subject: [PATCH 1/5] Specify concurrent-edit semantics and reported validity for rc.5 Validity is reported, never guaranteed: checks fall into request and safety, single-record, and cross-record tiers, and cross-record checks never block a write made through an engine. - collection.merge declares conflict, max, min, or union per top-level field, with max for lifecycle now/today fields and union for tags and uniqueItems arrays by default. - Update accepts add and remove list operations. - collection.unique gains enforce: write | report (default report) and exact governed-record and comparison-set scope semantics. - Paths are equivalent under NFC plus full case folding, with a deterministic ' (n)' suffix for derived and concurrent collisions. - A time-dependent or cross-record match.expr loads with a nondeterministic_match warning; it becomes an error in 0.3.0 stable. - Lifecycle has no sequence provider; CEL is the one expression language and Obsidian Bases expressions are an adapter dialect. - New chapter 12A defines record identity without in-file IDs, move detection, the three-way record merge with append-append body union, and the writer format fidelity rule. - settings.id_field has no default, path-pattern values may not contain '/', and ambiguous IDs resolve to null with ambiguous_link. - Chapter 13 lists the changes since rc.4 and how engines report the tightenings. - The v0.3 type-file schema gains collection.merge and unique[].enforce, and the claim schema gains the merge profile. --- 00-overview.md | 39 ++- 01-concepts.md | 40 ++- 02-collection-layout.md | 63 +++++ 03-records-and-frontmatter.md | 20 +- 04-configuration.md | 88 ++++-- 05-type-files.md | 9 +- 06-json-schema-profile.md | 9 + 07-collection-semantics.md | 174 +++++++++++- 08-links.md | 42 ++- 09-lifecycle.md | 17 ++ 10-cel-profile.md | 30 +- 12-operations.md | 84 +++++- 12a-concurrent-edits.md | 303 +++++++++++++++++++++ 13-migrations-and-compatibility.md | 110 +++++++- 14-conformance.md | 72 ++++- README.md | 2 +- schemas/v0.3/config.schema.json | 3 +- schemas/v0.3/conformance-claim.schema.json | 5 + schemas/v0.3/type-file.schema.json | 21 ++ site/build.mjs | 3 +- 20 files changed, 1027 insertions(+), 107 deletions(-) create mode 100644 12a-concurrent-edits.md diff --git a/00-overview.md b/00-overview.md index a7df1c8..6a8f361 100644 --- a/00-overview.md +++ b/00-overview.md @@ -73,6 +73,16 @@ dependencies. **Files are the source of truth.** Tools read from and write to the filesystem. Indexes, caches, and derived databases can be rebuilt from collection state. +**Plain Markdown.** A collection never requires mdbase metadata in a user's +files. Record identity, revisions, merge state, and other engine bookkeeping +live outside records, so a file written by any editor is a complete record. + +**Validity is reported, not guaranteed.** Files are edited by tools that know +nothing about types. A conforming tool reads, indexes, and reports every record +whatever its validity. Engines may reject an invalid write made through them, +as feedback to the writer, but only for checks within that single record. +Chapter 04 defines the principle and its three tiers. + **Human-readable first.** Persistent collection data uses open text formats. A user with a text editor can read and modify every record and type file. @@ -243,13 +253,28 @@ projections, ordering, grouping, and summaries remain machine-readable, while the Markdown body documents the view for people. Optional presentation metadata can select a renderer without changing query results. -### Validation is progressive +### Validation is progressive and reported Files in a collection can remain untyped records. Types can be added incrementally, and validation severity is configurable as `off`, `warn`, or `error`. JSON Schema controls field shape and unknown-property handling. Collection rules add checks that depend on other records or paths. +Validity is a property reported when records are read. At level `error`, a +write made through an engine fails when the resulting record breaks one of +its own checks. Checks that span records, such as link existence and +uniqueness, are reported and never block a write, unless a uniqueness rule +explicitly opts into `enforce: write`. + +### Concurrent edits merge field by field + +When two edits to one record meet, for example an application update and an +edit made in a text editor, a tool that reconciles them uses the three-way +record merge of Chapter 12A. Different fields merge, timestamps take the +later value, set-like lists take the union, appends to the body are both kept, +and only real disagreements are conflicts. How a tool surfaces a conflict, and +whether it replicates collections at all, is outside this specification. + ### Links connect records across the collection Records can reference each other with wikilinks such as `[[alice]]`, Markdown @@ -296,6 +321,7 @@ keep those claims precise and independently testable. | [10-cel-profile.md](./10-cel-profile.md) | Portable expressions and host bindings | | [11-querying.md](./11-querying.md) | Filters, ordering, projection, and result envelopes | | [12-operations.md](./12-operations.md) | Read and write operations, concurrency, and diagnostics | +| [12a-concurrent-edits.md](./12a-concurrent-edits.md) | Record identity, move detection, three-way merge, and writer format fidelity | | [13-migrations-and-compatibility.md](./13-migrations-and-compatibility.md) | Migration from earlier versions and compatibility | | [14-conformance.md](./14-conformance.md) | Profiles, claims, fixtures, and runners | @@ -315,7 +341,9 @@ directly comparable without making one product's internal API normative. ## Versioning -This specification uses semantic versioning. The current version is **0.3.0**. +This specification uses semantic versioning. The current version is **0.3.0**, +in its fifth release candidate (`0.3.0-rc.5`). 0.3.0 is declared stable once +implementations pass the conformance suite. Collections declare their specification version with `spec_version` in `mdbase.yaml`. Tools declare the profiles and versions they implement. @@ -341,7 +369,12 @@ through version requirements. The keywords `MUST`, `MUST NOT`, `SHOULD`, `SHOULD NOT`, and `MAY` are to be interpreted as described in RFC 2119. -Draft notes use ordinary prose and are non-normative. +Draft notes use ordinary prose and are non-normative. A paragraph that begins +with **Provisional (rc.5).** records a choice made where the design input +left a detail open. It is normative in the release candidate and may change +before 0.3.0 is declared stable. The +[0.3.0-rc.5 release notes](./docs/releases/0.3.0-rc.5.md) list every such +choice. ## License diff --git a/01-concepts.md b/01-concepts.md index df1e09e..dead79e 100644 --- a/01-concepts.md +++ b/01-concepts.md @@ -22,6 +22,12 @@ or data contract file, and not another reserved collection file. A record has: The persisted frontmatter object is the raw record value. Effective read values may additionally include `collection.read_defaults`. +A record is identified by its path. A record carries no required mdbase +metadata: no ID, revision, or merge state is ever required in the file. +Implementations MAY keep internal record identities outside the collection's +records and follow a record across moves with the move detection of Chapter +12A. + ## Type A type is a Markdown file, usually under `_types/`, whose frontmatter has @@ -62,13 +68,26 @@ record matches multiple types, it is valid only if it validates against every matched type's JSON Schema and every matched type's mdbase collection validators. +Membership should be a deterministic function of the record's path, its +persisted frontmatter, and the type registry, independent of the current time +and of other records. Chapter 07 defines the diagnostic for match rules that +break this. + +## Validity + +Validity is a property reported when a record is read, never a guarantee. +Any tool can write any bytes to a file, so a collection may always contain +invalid records. Conforming tools read, index, and query them, and report +their issues. Chapter 04 defines which checks may reject a write made through +an engine. + ## Collection Semantics Collection semantics are rules that require knowledge of the file tree or runtime context. Examples: - link parsing and target resolution -- cross-file uniqueness +- cross-record uniqueness - effective read defaults - path generation - display metadata @@ -86,14 +105,23 @@ simple transforms. Lifecycle policy is deterministic operation behavior within Core Write. It runs from type policy during the active mutation. +## Merge + +A merge combines two concurrent edits of one record against their common base +version. Each top-level frontmatter field has a merge strategy, declared in +`collection.merge` or derived from the type, and the body merges line by line. +Chapter 12A defines the merge function. When and where merges happen is an +implementation concern. + ## Expression -Portable v0.3 expressions use the mdbase CEL profile. Expressions appear in -queries, projections, runtime conditions, workflow input templates, and optional -lifecycle guards. +mdbase has one expression language: the mdbase CEL profile. Expressions appear +in `match.expr`, queries, projections, lifecycle guards, runtime conditions, +and workflow input templates. -`match.where` uses the standalone structured predicate language defined in -Chapter 07. +`match.where` is a structured predicate written as YAML data, not an +expression language; Chapter 07 defines it. Other expression syntaxes, such +as Obsidian Bases formulas, are adapter dialects (Chapter 10). ## View diff --git a/02-collection-layout.md b/02-collection-layout.md index 5ca7a1a..8101f38 100644 --- a/02-collection-layout.md +++ b/02-collection-layout.md @@ -150,3 +150,66 @@ required. Implementations MAY reject platform-reserved filenames or characters when a write operation targets a filesystem where those paths cannot be represented. + +## Path Equivalence + +Two collection paths name the same record path when their **path keys** are +equal. The path key of a path is computed as follows: + +1. normalize the path to Unicode Normalization Form C (NFC) +2. apply Unicode default case folding (the full `C` and `F` mappings of + `CaseFolding.txt`, without locale tailoring) +3. normalize the result to NFC again + +`Notes/Café.md` written with a precomposed `é` and `notes/CAFE\u0301.md` +written with a combining accent have the same path key. So do `Straße.md` and +`STRASSE.md`. Path keys never appear in results; paths are always reported as +written. + +Path equivalence exists because macOS and Windows file systems treat such +paths as one file, while Linux does not. Every tool therefore agrees on what +collides, whatever file system it runs on: + +- A write MUST NOT create a record whose path key equals the path key of a + different existing record. The collision rule below decides what happens + instead. +- A rename whose source and target have the same path key, such as + `tasks/todo.md` to `tasks/Todo.md`, changes only the spelling of the path + and is not a collision. +- When record discovery finds several files with one path key, which can + happen on a case-sensitive file system, each file is still a record. Core + Read reports a `path_collision` warning on every record of the group, with + `details.paths` listing the group in code-point order. + +Path globs (above) remain case-sensitive and match paths as written. + +**Provisional (rc.5).** Case folding uses the full mappings, so `ß` and +`ss` collide. This flags more collisions than some file systems would, never +fewer. + +## Path Collisions + +When a new record would take a path whose path key is already in use, the +outcome depends on where the path came from: + +| Path source | Outcome | +| --- | --- | +| an explicit path supplied by the caller of create or rename | the operation fails with `path_conflict` before any write | +| a path derived from `collection.path.pattern` (Chapter 07) | the record receives the first free suffixed path | +| two records that already hold equivalent paths, for example after concurrent creates or an engine copying records onto another file system | the earlier-ordered record keeps the path; each later one receives the first free suffixed path | + +A suffixed path inserts ` (n)`, a space and a decimal integer in parentheses, +before the final extension of the last path component: `tasks/Call Bob.md` +becomes `tasks/Call Bob (2).md`. Candidates are tried with `n = 2, 3, 4, …`, +and the first candidate whose path key is unused is chosen. Existing suffixes +are not parsed: the next candidate for `Call Bob (2).md` is +`Call Bob (2) (2).md`. + +The ordering of records is supplied by whatever applies the rule, for example +the order in which an engine confirmed two creates. When no order exists +between the records, the record whose path as written is smaller in Unicode +code-point order is earlier. Every tool that applies the rule to the same +records in the same order computes the same paths. + +A suffixed path is an ordinary path. Applying the rule never edits a record's +frontmatter or body, and the record keeps whatever identity its tool tracks. diff --git a/03-records-and-frontmatter.md b/03-records-and-frontmatter.md index a4c92ce..3a82784 100644 --- a/03-records-and-frontmatter.md +++ b/03-records-and-frontmatter.md @@ -171,14 +171,19 @@ explicitly maps it to ordinary fields. ## Serialization -Write-capable tools SHOULD preserve unrelated body text and line ending style. +Write-capable tools MUST preserve unrelated body text and SHOULD preserve the +line ending style. A YAML document record serializes as its frontmatter mapping alone. A create or update that supplies a non-empty body for a YAML document record fails with -`invalid_request` before any write. Structured writes re-emit the mapping and -need not preserve comments, key order, or quoting style; a whole-document -`document` replacement (Chapter 12) is written exactly as supplied. Tools that -edit files another application owns SHOULD use whole-document replacement. +`invalid_request` before any write. A whole-document `document` replacement +(Chapter 12) is written exactly as supplied. + +A write that changes some frontmatter keys of an existing record follows the +format fidelity rule of Chapter 12A: it re-emits only the changed top-level +entries, keeps every other entry byte-identical, including comments, quoting, +blank lines, and order, and keeps a changed entry's collection style. The rule +applies to Markdown records and YAML document records alike. When serializing frontmatter, tools MUST: @@ -186,8 +191,9 @@ When serializing frontmatter, tools MUST: - omit keys that are missing - quote empty strings -Tools SHOULD preserve array and object structure and SHOULD produce -deterministic key ordering when an operation rewrites a generated file. +Tools SHOULD produce deterministic key ordering when an operation writes a new +record or rewrites a generated file. New keys added to an existing record are +appended after its existing entries. A null value in written frontmatter always means explicit null. Removing a key is a distinct operation; Chapter 12 defines how an update requests it. diff --git a/04-configuration.md b/04-configuration.md index d54ab8a..eb7d190 100644 --- a/04-configuration.md +++ b/04-configuration.md @@ -44,12 +44,19 @@ Pre-1.0 draft versions MAY be accepted by explicit compatibility setting. | `settings.record_extensions` | list of strings | `[md]` | record file extensions without dot | | `settings.validation` | string | `error` | default validation level: `off`, `warn`, or `error` | | `settings.explicit_type_keys` | list of strings | `[type, types]` | frontmatter keys used for explicit type declarations | -| `settings.id_field` | string | none | field used for ID-based wikilink resolution; when absent, wikilinks resolve by path and filename only | +| `settings.id_field` | string | none | field used for ID-based wikilink resolution and as a move-detection identity hint; when absent, wikilinks resolve by path and filename only | | `settings.exclude` | list of globs | `[]` | paths excluded in addition to the built-in exclusions in Chapter 02 | `settings.explicit_type_keys` replaces the default key list. An empty list makes all type membership inferred. +`settings.id_field` has no default. When the key is absent, a tool MUST NOT +resolve wikilinks by any frontmatter field, including a field named `id`, and +MUST NOT treat any field as a move-detection identity hint. A collection that +wants ID-based resolution names the field explicitly, as the recommended +configuration above does. The configured field is ordinary frontmatter: mdbase +never requires it, writes it, or reserves it. + `settings.timezone`, when present, MUST be an IANA timezone identifier. `UTC` is the canonical identifier for Coordinated Universal Time. Numeric offsets and ambient aliases such as `local` are invalid because they do not name a durable @@ -63,31 +70,70 @@ reserved control-file folders and are excluded from ordinary record discovery. Unknown config keys MUST produce a warning while normal config loading continues. An explicit strict-config mode MAY reject them. +## Validation Principle + +Validity is a property reported when records are read, never a guarantee. +Files are edited by tools that know nothing about types, and a merge of two +individually valid edits can produce an invalid record. A conforming tool +MUST read, index, query, and report every record whatever its validity. An +engine MUST NOT refuse to take in a file because it is invalid, and MUST NOT +modify an invalid file to make it valid unless a caller asks for that write. + +Engines MAY reject an invalid write made through them, as feedback to the +writer, but only for request and safety checks and for checks within the +single record being written. Checks are divided into three tiers: + +| Tier | Checks | Effect on a write made through an engine | +| --- | --- | --- | +| request and safety | invalid requests, path escapes and unsafe paths, `path_conflict` for an explicit path, configuration and type-file errors, `type_conflict`, `type_membership_changed`, lifecycle failures, `if_revision` failures, expression compilation errors, and `collection.unique` rules with `enforce: write` | always rejected, at every validation level | +| single-record | JSON Schema failures, `format_invalid`, non-mapping frontmatter, and data contract view failures | follows the validation level below | +| cross-record | `collection.unique` rules with `enforce: report`, `link_not_found`, link `target_type` mismatches, `ambiguous_link`, and `path_collision` | reported, never rejected | + +A single-record check depends only on the record's own path, frontmatter, +body, and the type registry. A cross-record check also depends on other +records; such a check never blocks a write, because another tool, another +device, or a later edit can change the other records at any time. The only +exception is a uniqueness rule that explicitly opts into `enforce: write` +(Chapter 07). + +Writes that do not go through an engine's write operations are never rejected. +That includes edits made by other tools, files that appear through a file +system or synchronization tool, and the results of the merge in Chapter 12A. +Their issues are reported when the record is read or validated. + ## Validation Levels -`settings.validation` controls how record validation issues affect operations. -Record validation issues are JSON Schema failures, `format_invalid`, collection -validator failures such as uniqueness and `validate_exists`, non-mapping -frontmatter, and data contract view failures. +`settings.validation` controls how single-record and cross-record validation +issues are reported, and whether single-record issues reject a write made +through an engine. | Level | Record validation | Reads and queries | Create, update, rename, batch | | --- | --- | --- | --- | | `off` | not performed | return records without validation diagnostics | write without record validation | | `warn` | performed | return records with `warning` diagnostics | write and report `warning` diagnostics | -| `error` | performed | return records with `error` diagnostics | fail before writing when the resulting record has any issue | +| `error` | performed | return records with `error` diagnostics | fail before writing when the resulting record has a single-record issue; write and report cross-record issues as `warning` diagnostics | At every level a read returns the record, including an invalid one, and a query evaluates every candidate. Queries do not report per-record validation issues unless the caller requests them. -Validation levels never relax request and safety checks: invalid requests, path -escapes, type-file and configuration errors, `type_conflict`, -`type_membership_changed`, lifecycle failures, concurrency conflicts, and -expression compilation errors are always errors. +A successful write reports any cross-record issues of the written record as +`warning` diagnostics at every level except `off`, so a write that succeeded +always returns `valid: true`. A read or `validate` of the same record reports +them with the level's severity. + +Validation levels never relax the request and safety tier: those checks are +always errors. An explicit `validate` operation always performs record validation. At level `off` it reports issues as `warning`. +**Provisional (rc.5).** Data contract view failures are classed as +single-record, although a contract view built from collection projections can +depend on other records or on the current time. Such projections are rare in +contract mappings; a later release candidate may move the affected failures to the +cross-record tier. + ## Runtime Host Config Durable-runtime enablement, worker identity, storage, transport binding, and @@ -103,16 +149,18 @@ has one resolution model and no contract mode. ## Expressions -Portable v0.3 expressions are CEL. No config key is required to opt into CEL. - -Tools MAY support non-portable UI expression dialects. Portable stored v0.3 -files MUST use the mdbase CEL profile unless a feature declares a different -extension namespace. - -View records use CEL for portable filters, projections, selections, and custom -summaries. Compatibility tools MAY read another view or expression format, but -alternate source and round-trip metadata belong under an `x-*` extension and do -not change the meaning of the portable CEL fields. +mdbase has one expression language: the mdbase CEL profile (Chapter 10). No +config key is required to opt into CEL. Every expression in a type file, +`mdbase.yaml`, a query, a view record, a lifecycle policy, or a workflow is +CEL. + +Other expression syntaxes, such as Obsidian Bases filters and formulas, are +adapter dialects. An adapter dialect appears only in a source format that an +adapter owns, such as an `obsidian.base` record, or under an `x-*` extension +whose owner defines its semantics. A dialect never affects type membership, +validation, lifecycle, merge, or the meaning of a portable query. Tools MAY +translate a dialect to CEL, and MAY offer a dialect in a user interface, as +long as what they store in portable members is CEL. ## Version Compatibility diff --git a/05-type-files.md b/05-type-files.md index b92ead2..7b66684 100644 --- a/05-type-files.md +++ b/05-type-files.md @@ -143,7 +143,7 @@ The following top-level sections are defined by v0.3: | Section | Purpose | | --- | --- | | `match` | select records for inferred type membership | -| `collection` | define Markdown-aware collection semantics | +| `collection` | define Markdown-aware collection semantics, including merge strategies | | `lifecycle` | assign managed values during mutations | | `implements` | declare exact, schema-validated data contract implementations | @@ -223,11 +223,12 @@ only when all of those validations pass. Collection behavior composes as follows: - uniqueness rules are additive and are each evaluated in the type that - declared them + declared them, with that rule's own `enforce` mode - identical read defaults, link rules, path policies, lifecycle assignments, - and projections coalesce + projections, and merge strategies coalesce - different values for the same read-default field, link selector, path policy, - lifecycle event and field, or projection name produce `type_conflict` + lifecycle event and field, projection name, or `collection.merge` field + produce `type_conflict` - display metadata remains associated with its declaring type; a flattened display uses the first explicit type or first canonical inferred type diff --git a/06-json-schema-profile.md b/06-json-schema-profile.md index b7aa8af..aaf667f 100644 --- a/06-json-schema-profile.md +++ b/06-json-schema-profile.md @@ -107,6 +107,15 @@ Effective read defaults belong in `collection.read_defaults`. Dynamic defaults such as `now`, `uuid`, `ulid`, and `slugify` belong in `lifecycle`. +## Merge Annotations + +Merge strategies are mdbase collection semantics and are declared in +`collection.merge` (Chapter 07), not as JSON Schema keywords. A JSON Schema +`uniqueItems: true` on a top-level array property selects the `union` merge +strategy by default, as Chapter 07 defines. Implementation-specific keywords +such as `x-merge` inside `schema.value` are annotations with no portable +meaning. + ## Format JSON Schema `format` is an annotation in the base dialect. The mdbase v0.3 diff --git a/07-collection-semantics.md b/07-collection-semantics.md index e1f537b..7475242 100644 --- a/07-collection-semantics.md +++ b/07-collection-semantics.md @@ -15,12 +15,14 @@ collection: read_defaults: status: open unique: - - field: id + - field: code scope: type links: assignee: target_type: person validate_exists: true + merge: + completedDate: max ``` ## Field References @@ -146,6 +148,25 @@ The expression combines with the other members of `match` using AND. It receives the matching context defined in Chapter 10 and MUST evaluate to boolean true for the type to match. +Type membership drives validation, lifecycle, path policy, and merge, so every +tool and every replay MUST compute the same membership for the same record. +`match.expr` SHOULD therefore be deterministic: it SHOULD NOT call `now()` or +`today()`, read `file.mtime` or `file.ctime`, or use a helper that reads +other records, such as `asFile()`, `file.backlinks`, or `file.hasLink()`. +Queries, projections, and lifecycle guards MAY use these functions. + +A tool that loads a type whose `match.expr` uses one of them MUST report a +`warning` diagnostic with code `nondeterministic_match`, the type name, and +`details.binding` naming the function or field. The type still loads and the +expression is evaluated as before, with `now()` and `today()` reading the +operation's captured instant. Membership of such a type can differ between +tools and over time, so tools that merge or replay edits cannot rely on it. + +In 0.3.0 stable such a `match.expr` is an error: the type is invalid and +reports `nondeterministic_match` with severity `error`. The warning in the +release candidates gives authors a window to move time-dependent or +cross-record logic into queries or collection projections. + Implementations compile `match.expr` when loading the type. Parse and type errors invalidate the type definition. A per-record evaluation error reports a diagnostic for that record and expression and yields a non-match. False and null @@ -211,24 +232,90 @@ collection: defines link parsing and resolution. `/relations` applies the same item-wise rule when the exactly selected value is an array. -## Cross-File Uniqueness +## Uniqueness -`collection.unique` declares collection-level uniqueness: +`collection.unique` declares that a field's values must differ between +records: ```yaml collection: unique: - - field: id + - field: code scope: type + enforce: write + - field: slug + scope: path_glob + path_glob: "docs/**" ``` +| Member | Required | Meaning | +| --- | --- | --- | +| `field` | yes | field reference (Chapter 07); with `[]`, every selected item takes part | +| `scope` | no, default `type` | which records the governed records must differ from | +| `path_glob` | when `scope` is `path_glob` | a portable glob (Chapter 02) | +| `enforce` | no, default `report` | `report` or `write` | + +A rule **governs** every record that matches its declaring type. With +`scope: path_glob`, it governs only those whose path also matches +`path_glob`. The scope selects the **comparison set**: + | Scope | Comparison set | | --- | --- | -| `collection` | every record in the collection | -| `type` | every record matching the declaring type | -| `path_glob` | every record under the configured path glob | - -Missing and null values are exempt from uniqueness comparison. +| `type` | every record that matches the declaring type | +| `collection` | every record in the collection, whatever its types | +| `path_glob` | every record whose path matches `path_glob`, whatever its types | + +A governed record violates the rule when another record in the comparison set +holds an equal value for the field. Values are compared as follows: + +- the persisted (raw) value is compared; read defaults and projections never + take part +- missing and null values are exempt +- values are equal under deep JSON equality, where numbers are equal when + their numeric values are equal; there is no coercion between types, so `1` + and `"1"` differ, and strings are compared exactly, without case folding or + normalization +- when the field reference selects several values, each value takes part; + repeated items within one record's own array are not a uniqueness violation + (JSON Schema `uniqueItems` covers that) + +Each violation is a `duplicate_value` issue on the governed record, with the +rule's field and `details.paths` listing the other records that hold the value +in code-point order. + +### Enforcement + +`enforce` declares whether a rule is checked when a write is made through an +engine: + +- `report` (the default): violations are cross-record issues (Chapter 04). + They are reported and never block a write. +- `write`: a create, update, rename, or batch item made through an engine + fails with `duplicate_value` before writing when it would give a governed + record a value that another record in the comparison set already holds, or + would add to the comparison set a record whose value a governed record + already holds. This belongs to the request and safety tier, so it applies + at every validation level. + +`enforce: write` is the one cross-record check that can reject a write. It +requires that concurrent writes are checked in a single order: when two writes +would claim the same value, the first one in that order succeeds and the later +one fails. How an implementation orders writes, for example through a single +writer or a coordination service, is outside this specification. An +implementation that cannot order writes that way for a collection MUST reject +writes to fields covered by an `enforce: write` rule rather than accept them +unchecked. Collections that need unique identifiers without coordination +SHOULD generate them with lifecycle `ulid` or `uuid`. + +An `enforce: write` rule rejects only writes that introduce a duplicate. A +write that leaves a record's covered values unchanged succeeds even when the +record already violates the rule, for example after an external edit. Edits +made by other tools and the results of a merge (Chapter 12A) are never +rejected; their violations are reported. + +**Provisional (rc.5).** An `enforce: write` rule applies at every +validation level, including `off`, because the author opted into it +explicitly. ## Path Policy @@ -241,12 +328,71 @@ collection: ``` The portable grammar uses `{field}` placeholders for top-level frontmatter -fields. Values are converted to strings without expression evaluation. A -missing or null value produces `path_value_missing`. +fields. Values are converted to strings without expression evaluation: a +string is used as written, and a number or boolean uses its JSON +representation. A missing or null value produces `path_value_missing`. + +A placeholder value always stays within one path component. A converted +value is invalid, and the operation fails with `path_value_invalid` naming the +field, when it: + +- is an array or an object +- is empty +- contains `/`, `\`, or a NUL character +- begins with `.` + +A title such as `Q3/Q4 plan` therefore cannot create a subfolder. Generated +paths MUST remain inside the collection root and MUST satisfy the path safety +rules of Chapter 02. When a derived path's path key is already in use, the +record receives the first free suffixed path defined in Chapter 02; the +operation does not fail. Runtime-owned path logic and richer template +languages belong under an `x-*` extension. -Generated paths MUST remain inside the collection root. Placeholder values that -produce `/`, `\\`, `.`, or `..` path components are invalid. Runtime-owned path -logic and richer template languages belong under an `x-*` extension. +## Merge Strategies + +`collection.merge` declares how each top-level frontmatter field combines when +two concurrent edits of a record are merged (Chapter 12A): + +```yaml +collection: + merge: + completedDate: max + status: conflict + reviewers: union +``` + +Each key names one top-level frontmatter field, as a field path with one +segment and no `[]`, or a JSON Pointer with one token. Each value is a +strategy: + +| Strategy | When both sides changed the field to different values | +| --- | --- | +| `conflict` | the field is in conflict | +| `max` | the greater of the two values | +| `min` | the lesser of the two values | +| `union` | an observed-remove set union of the two lists | + +Chapter 12A defines each strategy exactly. A strategy matters only when both +sides changed the same field differently. A field changed on one side always +takes that side's value, and a field changed identically on both sides takes +the shared value. + +A field without a declaration uses its default strategy. The first rule that +applies wins: + +1. `max` when a matched type's lifecycle assigns the field with `{ now: true }` + or `{ today: true }`, so that concurrent modification timestamps never + conflict +2. `union` when the field is named `tags`, or when a matched type's schema + declares the field's top-level property with `uniqueItems: true` +3. `conflict` otherwise + +A declaration always replaces the default, so `dateModified: conflict` turns +the timestamp default off. Applications that maintain timestamps in their own +code, rather than through lifecycle, declare `max` explicitly. + +Merge strategies describe concurrent edits only. They never change +validation, reads, queries, or the result of a single write. ## Display Metadata diff --git a/08-links.md b/08-links.md index a988800..b0dc2cd 100644 --- a/08-links.md +++ b/08-links.md @@ -42,24 +42,38 @@ General rules: target against record filenames with or without their record extension. - When `settings.id_field` is configured, a simple wikilink first tries ID-based resolution against that field and falls back to filename - resolution when no record has that ID. + resolution only when no record has that ID. When `settings.id_field` is + absent, no ID-based resolution happens, whatever fields records contain. + +A record has an ID when its persisted `id_field` value is a non-empty string. +ID-based resolution compares the wikilink target with record IDs exactly, +without case folding or normalization. After normalization, a link that escapes the collection root is invalid. ## Ambiguity -If multiple records have the same configured ID, ID-based resolution is -ambiguous and MUST fail without falling back to filename resolution. +If several records have the configured ID, ID-based resolution is ambiguous. +The link resolves to null, and the tool MUST NOT fall back to filename +resolution. Duplicate IDs are not a record validation issue of the records +that hold them. -If filename resolution finds multiple candidates, tools SHOULD apply stable -tiebreakers: +If filename resolution finds multiple candidates, tools MUST apply these +tiebreakers in order: 1. same directory as referring file 2. shortest collection path -3. alphabetical path +3. smallest path in Unicode code-point order + +Filename candidates whose paths are equivalent under Chapter 02 path keys are +ambiguous with each other even after the tiebreakers. If ambiguity remains, +the link resolves to null. -If ambiguity remains, resolution returns null and reports an ambiguous link -warning. +An ambiguous link reports an `ambiguous_link` cross-record issue on the +referring record, with `details.candidates` listing the candidate paths in +code-point order. An ambiguous link creates no backlink, and for +`validate_exists` it counts as unresolved but reports `ambiguous_link` rather +than `link_not_found`. ## Target Constraints @@ -74,11 +88,15 @@ collection: ``` When `validate_exists` is true, an unresolved link is a `link_not_found` -record validation issue whose severity follows the validation level in -Chapter 04. +record validation issue. + +When `target_type` is present, a resolved target that does not match the +target type is a `link_target_type_mismatch` record validation issue. -When `target_type` is present, a resolved target is valid only if it matches the -target type. +Both are cross-record checks (Chapter 04): their severity follows the +validation level, and they are reported but never block a write. A record +can lose its link target at any time through an edit, delete, or rename of +another record. ## Body Links diff --git a/09-lifecycle.md b/09-lifecycle.md index ce9bf58..b820188 100644 --- a/09-lifecycle.md +++ b/09-lifecycle.md @@ -65,6 +65,14 @@ Core lifecycle providers: operation observes the same time. The operation timezone follows the same precedence as the query timezone in Chapter 11. +Lifecycle has no counter or sequence provider. A dense sequence such as +"largest existing value plus one" depends on every other record, so two tools +or devices creating records concurrently allocate the same number unless every +create is coordinated. Collections use `ulid` or `uuid` for identifiers. +An implementation MAY offer a sequence provider under an `x-*` extension; such +a provider is not portable, and its allocation and coordination are +implementation behavior. + `slugify` lowercases the value, transliterates it to ASCII where a transliteration exists, replaces each run of other characters with `-`, and trims leading and trailing `-`. A missing, null, or non-string source value @@ -95,6 +103,15 @@ raises an evaluation error fails the operation with `lifecycle_expression_error`. A guard that evaluates to anything other than boolean `true` skips its action. +## Merge Defaults + +A top-level field that a matched type's lifecycle assigns with `{ now: true }` +or `{ today: true }` merges with the `max` strategy by default (Chapter 07), +so two concurrent updates that both refresh `dateModified` keep the later +value instead of conflicting. A merge does not run lifecycle: lifecycle values +in a merged record come from the two sides through their merge strategies +(Chapter 12A). + ## Validation Order For create and update: diff --git a/10-cel-profile.md b/10-cel-profile.md index 92ee9c7..9ceb051 100644 --- a/10-cel-profile.md +++ b/10-cel-profile.md @@ -2,9 +2,12 @@ ## Purpose -Portable v0.3 expressions use the -[Common Expression Language](https://github.com/google/cel-spec) (CEL). This -chapter defines the values an mdbase host binds, the mdbase host functions and +mdbase has one expression language: the +[Common Expression Language](https://github.com/google/cel-spec) (CEL) with the +host bindings of this profile. Matching, lifecycle guards, collection +projections, queries, views, and workflows all use it, so one parser, one set +of semantics, and one conformance suite serve every context. This chapter +defines the values an mdbase host binds, the mdbase host functions and types, limits, and how each embedding context handles evaluation errors. mdbase does not change CEL's language semantics. Operators, macros, error @@ -51,11 +54,18 @@ CEL appears in: - lifecycle guards - workflow variables, conditions, inputs, iteration, and run policy -`match.where` uses the structured predicate language from Chapter 07. +`match.where` is a structured predicate written as YAML data (Chapter 07). It +has no expression syntax. -Tools may translate another user-interface expression language to CEL before -writing portable records. A stored alternate dialect uses an `x-*` extension -whose owner defines its semantics. +Other expression syntaxes are **adapter dialects**. The Obsidian Bases +filter and formula language is one: it is evaluated only by the +[Obsidian Bases adapter](./adapters/obsidian-bases.md) for records that +implement the `obsidian.base` contract. An adapter dialect is stored only in a +source format that its adapter owns or under an `x-*` extension whose owner +defines its semantics. A dialect expression never decides type membership, +validation, lifecycle, merge, or the meaning of a portable query. Tools MAY +translate a dialect to CEL before writing portable members, and MAY translate +CEL back for a user interface. ## Evaluation Contexts @@ -128,6 +138,12 @@ file.inFolder("tasks") && has(raw.tags) && tags.exists(t, t == "task") Read defaults and projections enter after matching and are absent from this context. +Matching should be deterministic (Chapter 07). Using `now()`, `today()`, +`file.mtime`, `file.ctime`, or a helper that reads another record, such as +`asFile()`, `file.backlinks`, or `file.hasLink()`, in this context reports a +`nondeterministic_match` warning when the type loads. In 0.3.0 stable it is +an error that invalidates the type. + ### Lifecycle Context Lifecycle expressions evaluate against the current write draft. Top-level diff --git a/12-operations.md b/12-operations.md index 8fceda3..f6f0d43 100644 --- a/12-operations.md +++ b/12-operations.md @@ -55,7 +55,7 @@ Pipeline: 4. apply lifecycle `on_create` 5. verify type membership did not change as a lifecycle side effect 6. validate JSON Schema -7. run collection validators +7. run collection validators, applying the validation tiers of Chapter 04 8. choose or validate path policy; the final path's extension fixes the record's format (Chapter 03), and a body is rejected when that format is a YAML document @@ -63,6 +63,11 @@ Pipeline: 10. update derived indexes 11. emit watch/runtime events after state is consistent +An explicit `path` whose path key (Chapter 02) equals that of an existing +record fails with `path_conflict`. A path derived from `collection.path.pattern` +never fails for that reason: it receives the first free suffixed path from +Chapter 02, and the result reports the final path. + Static JSON Schema defaults MAY be used by editor and create interfaces. Validation-time mutation occurs when the create operation explicitly copies a default into the draft. @@ -99,13 +104,14 @@ Update modifies an existing record. Pipeline: 1. read existing raw frontmatter -2. apply the requested `patch` and `unset` +2. apply the requested `patch`, `unset`, `add`, and `remove` 3. re-match and freeze types when type-affecting fields or path changed 4. apply lifecycle `on_update` 5. verify type membership did not change as a lifecycle side effect 6. validate JSON Schema -7. run collection validators -8. write frontmatter and preserve body +7. run collection validators, applying the validation tiers of Chapter 04 +8. write the changed frontmatter entries with format fidelity (Chapter 12A) + and preserve the body 9. update derived indexes 10. emit watch/runtime events @@ -115,6 +121,10 @@ A structured update accepts: replacing any existing value; a null value persists an explicit null - `unset`: a list of field references, as defined in Chapter 07, whose keys are removed from persisted frontmatter +- `add`: an object mapping top-level field names to lists of items to add to + that field's list +- `remove`: an object mapping top-level field names to lists of items to + remove from that field's list - `body`: optional replacement Markdown body; a non-empty body for a YAML document record (Chapter 03) is `invalid_request` @@ -124,6 +134,44 @@ replaces, or uses an `unset` reference that selects an array item, is invalid and produces `invalid_request` before any write. When `unset` removes the last key of a nested object, the now-empty object remains. +### List operations + +`add` and `remove` change a list field relative to its current value instead +of replacing it, so "add tag `x`" from one writer and "add tag `y`" from +another both take effect: + +```yaml +path: tasks/a.md +add: + tags: [urgent] +remove: + tags: [someday] +``` + +They apply to the record's current persisted value when the update is +applied, after `patch` and `unset`: + +- `add` appends each listed item that is not already present, in request + order. A missing or null field is treated as an empty list, so `add` creates + the field. +- `remove` removes every item equal to a listed item. Removing from a missing + or null field, or removing an item that is absent, is not an error and + leaves the field unchanged. +- Items are compared with the deep JSON equality used for uniqueness in + Chapter 07. +- When the existing value is a string and the field is `tags`, it is treated + as a one-item list, matching how `file.tags` reads it. Any other non-list + existing value is `invalid_request`. + +A request is `invalid_request` before any write when it names a field in +`add` or `remove` that it also names in `patch` or `unset`, when the same +item appears in both `add` and `remove` for one field, or when an `add` or +`remove` value is not a list. List operations work on any list field; they do +not require `uniqueItems`. A list emptied by `remove` stays as an empty list. + +**Provisional (rc.5).** `add` and `remove` address top-level fields +only, like `patch`, because the merge unit is the top-level field. + As an alternative to a frontmatter patch and body replacement, Update accepts `document` containing the complete candidate source in the record's format: Markdown source for a Markdown record, or the whole YAML document for a YAML @@ -152,7 +200,10 @@ Delete is a Core Write operation and MAY emit an event for workflow runtimes. Rename moves a record within the collection. It accepts `from` and `to` collection-relative paths, optional `if_revision`, and optional `update_refs`. -Tools MUST reject target paths that escape the collection root. +Tools MUST reject target paths that escape the collection root. A target whose +path key equals that of a different existing record fails with +`path_conflict`. A target whose path key equals the source's own path key only +changes the path's spelling, such as its case, and is not a conflict. If reference updating is enabled, link updates SHOULD preserve link style, alias, and anchor where possible. ID-based links SHOULD not be rewritten if the @@ -338,14 +389,27 @@ delete returns `path` and `deleted: true`. Write-capable tools SHOULD detect external modification between read and write using mtime, content hash, version token, or platform-specific file identity. -On conflict, tools MUST preserve the current file and report a concurrency -diagnostic. - Successful reads MUST return a stable `revision` token derived from the raw file state. Write operations MUST accept an optional `if_revision` token and fail with `concurrent_modification` when it no longer matches. The token format is implementation-defined and opaque to callers. +`if_revision` is an opt-in compare-and-swap for callers that need +read-modify-write semantics, such as a counter or a workflow state +transition. Implementations MUST NOT supply `if_revision` on a caller's +behalf. + +A write without `if_revision` describes its own change: the keys it sets or +removes, its list operations, and its body edit. When the record changed after +the caller read it, an implementation either applies that change to the +current record, as the update pipeline above does, or reconciles the two +versions with the three-way record merge of Chapter 12A. It MUST NOT discard +the other change silently. A conflict found by the merge is reported; how it +is held, shown, and resolved is implementation behavior. + +On a failed `if_revision` check, tools MUST preserve the current file and +report a concurrency diagnostic. + ## Operation Result Envelope Every operation returns a mapping with: @@ -376,6 +440,8 @@ runtime state, or revisions. ## Events -After a successful mutation, tools MAY emit watch/runtime events. Events MUST be +After a successful mutation, tools MAY emit watch/runtime events. A tool that +reports moves of files edited outside it uses the move detection of Chapter +12A. Events MUST be delivered after the derived read/query state is consistent. Watch consumers and workflow runtimes may subscribe to the same stream. diff --git a/12a-concurrent-edits.md b/12a-concurrent-edits.md new file mode 100644 index 0000000..f49cb52 --- /dev/null +++ b/12a-concurrent-edits.md @@ -0,0 +1,303 @@ +# 12A. Concurrent Edits + +## Scope + +A collection is edited by many tools at once: an application writing through +an engine, a person in a text editor, a synchronization tool, an agent. This +chapter defines the data semantics that let those edits meet without losing +either one: + +- how a record keeps its identity without any ID in its file, and how a moved + file is recognized as the same record +- the three-way record merge, which combines two concurrent versions of a + record against their common base +- the format fidelity rule that every writer follows, so that a write changes + only the bytes it means to change + +This specification does not define when a tool merges, how edits travel +between devices, logs, sequencers, ordering services, conflict envelopes, how +a conflict is held or shown to a person, or what happens when one side deletes +a record that the other side edited. Those are implementation concerns. A tool +that never reconciles concurrent edits, such as a command-line validator, does +not need this chapter's merge function and does not claim the `merge` +profile. + +## Record Identity + +A record is identified by its collection-relative path. Record files carry no +required mdbase metadata: + +- An implementation MAY keep an internal identity for each record, for example + to follow it across moves or to key its history. +- Such identities, revisions, merge bases, and other bookkeeping MUST be kept + outside record files, for example in `.mdbase/` or the implementation's own + storage. +- An implementation MUST NOT write an identity into a record's frontmatter or + body, and MUST NOT require one to read, validate, query, or write a record. + +A collection may still store its own identifiers in frontmatter, such as a +TaskNotes `id` field. When `settings.id_field` names such a field, it is used +for ID-based link resolution (Chapter 08) and as an identity hint for move +detection below. It remains ordinary user data. + +## Move Detection + +Tools that watch a collection see moves made by other tools as a file +disappearing at one path and a file appearing at another. A tool that reports +moves, follows a record's internal identity across them, or keeps merge bases +across them MUST pair disappearances and appearances with the rules below. + +### Observation window + +A disappearance and an appearance can pair when the tool observes both within +one **observation window**. A window MUST span at least 5 seconds, so that a +deletion and a creation that a file watcher reports in separate batches still +pair. A scan that compares the collection with the state a tool last recorded, +for example on start-up after the tool was not running, treats everything it +finds as one window. + +For each disappeared record the tool uses the last bytes and file identity it +observed. For each appeared record it uses the current bytes and file +identity. A **file identity** is a platform identifier that survives a rename, +such as a device and inode number or a Windows file ID. A tool that cannot +observe file identities skips the rule that uses them. + +### Pairing rules + +A disappeared record `D` and an appeared record `A` are a **candidate pair** +under the first of these rules that holds: + +| Rank | Rule | +| --- | --- | +| 0 | `settings.id_field` is configured and both `D` and `A` hold the same non-empty string value for it | +| 1 | `D` and `A` have identical bytes | +| 2 | `D` and `A` have the same file identity and a similarity of at least 0.5 | +| 3 | `D` and `A` have the same file name (`file.name`) and a similarity of at least 0.8 | + +When `settings.id_field` is configured and both `D` and `A` hold non-empty +string values for it that differ, they are never a candidate pair, whatever +the other rules say. The identity hint wins whenever both sides have one. + +The **similarity** of two files is the Jaccard index of their line sets: each +file's lines are trimmed of leading and trailing whitespace, empty lines are +discarded, and the remaining distinct lines form the set. The similarity is +the size of the intersection divided by the size of the union. Two files with +empty line sets have similarity 1. + +Candidate pairs are chosen greedily in this order: + +1. lower rank first +2. higher similarity first +3. smaller `D` path in Unicode code-point order first +4. smaller `A` path in Unicode code-point order first + +A chosen pair removes both its records from further pairing. Each chosen pair +is one record that moved from `D`'s path to `A`'s path, and possibly also +changed. Every disappearance left unpaired is a deletion and every appearance +left unpaired is a creation. + +Unrelated files never pair: a file deleted at one path and a different file +created at another, with different bytes, file identity, and name, are a +deletion and a creation. + +A tool reports a detected move as `record_renamed` (Chapter 14), followed by +`record_modified` when the bytes also changed. Internal identity, history, and +merge bases follow the record to its new path. + +**Provisional (rc.5).** The 5-second minimum window and the 0.5 and 0.8 +thresholds come from the feasibility prototype's measurements and may be +tuned before release. + +## Three-Way Record Merge + +### Inputs and result + +The merge combines two concurrent versions of one record. Its inputs are: + +- the **base**: the version both edits started from +- the **first** and **second** versions: the two edited versions, in an order + supplied by the caller of the merge, such as the order in which an engine + confirmed them +- the type registry + +Each version consists of a path, a persisted frontmatter mapping with its +source text, and a body. Merge strategies (Chapter 07) come from the types +that the first version matches at its path. + +The result is a **merged version** and a possibly empty list of +**conflicts**. A conflict has a `kind`: `field` for one top-level frontmatter +field, which it names in `field`; `frontmatter` for a whole frontmatter +block; `body`; or `path`. It carries the base, first, and second values. +Where a conflict exists, the merged version holds the first version's value. +A tool that writes a merged version with conflicts MUST NOT discard the second +version's conflicting values silently: how it keeps and surfaces them is +implementation behavior. + +The merge is a pure function: the same inputs always produce the same result. +It reads no clock and no other record. When one edited version has the same +bytes as the base, the merged version is the other edited version, byte for +byte; when the two edited versions have the same bytes, the merged version is +that version. + +### Equality + +Two frontmatter values are equal under deep JSON equality, where numbers are +equal when their numeric values are equal, strings are compared exactly, and +there is no coercion between types. A missing key is a state of its own: it is +equal only to another missing key, and never equal to null. + +### Frontmatter + +The frontmatter merges one top-level key at a time. For each key present in +the base, the first version, or the second version, with `B`, `F`, and `S` +standing for that key's state in each: + +1. If `F` equals `S`, the result is `F`. Both sides made the same change, or + neither changed the key. +2. Otherwise, if `S` equals `B`, the result is `F`: only the first side + changed the key. +3. Otherwise, if `F` equals `B`, the result is `S`: only the second side + changed the key. +4. Otherwise both sides changed the key differently, and the key's strategy + decides. + +| Strategy | Result when both sides changed the key differently | +| --- | --- | +| `conflict` | a conflict of kind `field` on the key | +| `max` | the greater of `F` and `S` | +| `min` | the lesser of `F` and `S` | +| `union` | the observed-remove union of `F` and `S` | + +**`max` and `min`.** Two values are ordered as follows: + +- two numbers by numeric value +- two strings that are both RFC 3339 date-times with an offset, as instants +- two strings that are both RFC 3339 `full-date` values, in calendar order +- any other two strings, in Unicode code-point order + +A present value is greater than a missing key under `max`, and also preferred +over a missing key under `min`, so a value survives a concurrent removal. When +`F` and `S` are incomparable, for example a number and a string, or null and a +value, the key is in conflict. When they are equal under the ordering but not +equal values, such as one instant written with two offsets, the result is `F`. + +**`union`.** Each of `B`, `F`, and `S` is read as a list: a list as itself, a +missing key or null as the empty list, and, for the `tags` field only, a +string as a one-item list. If any of them is another value, the key is in +conflict. Otherwise the result is: + +1. every item of `F` that is not an item of `B` removed by `S` (an item of `B` + absent from `S`), in `F`'s order +2. followed by every item of `S` that is not in `B` and not already in the + result, in `S`'s order + +Items are compared with the equality above. Additions from both sides are +kept, and an item removed on either side stays removed unless the other side +added it again. When the result is empty and `F` or `S` is a missing key, the +key is missing from the merged version; otherwise the result is a list. + +A key whose value came from one side unchanged takes that side's source text +under the format fidelity rule below. A value computed by `union`, or a `max` +or `min` result written in a different form, is re-emitted. + +The merged frontmatter source starts from the first version's source. Comment +and blank lines between entries therefore come from the first version. + +**Provisional (rc.5).** Comment or blank lines between entries that only +the second version changed are not carried into the merged version, unless the +first version has the same bytes as the base. A later release candidate may merge them as +text. + +When the frontmatter of any of the three versions is not a mapping, the whole +frontmatter is one unit: the result follows rules 1 to 3 above on its source +text, and otherwise there is one conflict of kind `frontmatter`. + +### Body + +The body merges line by line. A line is a run of characters ending with a line +terminator (`\n` or `\r\n`), or the final run of characters without one. Lines +are compared exactly, including their terminators. With `B`, `F`, and `S` +standing for the three bodies: + +1. If `F` equals `S`, or `S` equals `B`, the result is `F`. If `F` equals `B`, + the result is `S`. +2. **Append-append.** If `F` and `S` both begin with all of `B`, so that both + sides only appended text at the end, the result is `B`, then `F`'s + appended text, then `S`'s appended text. When `F`'s appended text is not + empty and does not end with a line terminator, a `\n` is inserted between + the two. Journals, logs, and checklists grow this way, and appending to + them concurrently is not a conflict. +3. Otherwise the bodies merge as a three-way line merge (diff3). The lines of + `B` that are aligned with unchanged lines in both `F` and `S` divide the + bodies into stable regions and changed chunks. For each changed chunk, the + result takes the side that changed it; where both sides changed a chunk to + the same lines, those lines; and where both sides changed a chunk + differently, the body is in conflict. + +When the body is in conflict, the merged body is `F` as a whole and the +conflict carries the three bodies. The alignment of `B` with each side is a +longest common subsequence of lines. + +**Provisional (rc.5).** When several longest common subsequences exist, this +release candidate does not fix which one is used, so two implementations can +split some changed chunks differently. Implementations SHOULD use the Myers +difference algorithm. Conformance fixtures use bodies whose alignment is +unique. + +**Provisional (rc.5).** Append-append applies only to appends at the end +of the body. Two insertions at the same place inside the body remain a +conflict. + +### Path + +When the first and second versions have different paths, the path merges like +a frontmatter key with the `conflict` strategy: one side's move wins over an +unchanged path, and two different moves are a conflict of kind `path`. The merged +path then goes through the path collision rule of Chapter 02. + +### Validity and lifecycle + +The merge never checks validity. A merged version is validated when it is +read, like any other record, and its issues are reported (Chapter 04). It is +never rejected, and it is never turned into a conflict because it is invalid, +even when both sides were valid on their own: for example, two edits that +each keep a record within an `if`/`then` schema can together leave it outside. + +Membership of the merged version is recomputed from its path and frontmatter +and may differ from the base. The merge does not run lifecycle; managed values +such as `dateModified` combine through their strategies. + +## Writer Format Fidelity + +Every tool that writes changes to the frontmatter of an existing record, +through an update, a lifecycle assignment, a merge, or a reference update +after a rename, MUST follow this rule. It applies to Markdown records and YAML +document records alike. + +A frontmatter source consists of **top-level entries** and the lines between +them. An entry is a line that begins a top-level key at column 0, together +with every following line that belongs to that key's value. Blank lines and +comment lines at column 0 between entries are not part of any entry. + +1. An entry whose key the write does not change MUST stay byte-identical, + including its comments, quoting, indentation, and position. +2. Lines that are not part of any entry MUST be kept, except that removing an + entry removes its own lines only. +3. A changed entry is re-emitted in place. When the previous value was a flow + collection, such as `tags: [a, b]`, the new value MUST be written as a flow + collection, and a block collection MUST stay in block style. Writers + SHOULD keep a scalar's quoting style and a trailing comment on the entry's + first line when the new value can be written that way. +4. A new key is appended after the last entry. +5. A value that a merge takes unchanged from one side MUST be copied verbatim + from that side's source text for the entry. +6. The byte-order mark, the frontmatter delimiters, the line ending style, and + the body, unless the write changes the body, MUST stay as they were. + +A whole-document `document` replacement (Chapter 12) is written exactly as +supplied and is outside this rule. + +The rule keeps unrelated bytes stable, so files diff cleanly in Git and a +write never manufactures a change on a line it did not mean to touch. Byte +changes beyond the edited keys would also turn into needless conflicts in the +next merge. diff --git a/13-migrations-and-compatibility.md b/13-migrations-and-compatibility.md index 95d80bc..f08ca23 100644 --- a/13-migrations-and-compatibility.md +++ b/13-migrations-and-compatibility.md @@ -2,9 +2,87 @@ ## Migration Philosophy -v0.3 uses the source model defined by Chapters 01–12. Migration translates -v0.2.x collections into that model and reports features that need adapter-owned -handling. +This chapter has two parts. The first lists what changed for v0.3 collections +and engines between the fourth and fifth release candidates. The second +migrates v0.2.x collections into the v0.3 source model. + +Migration never rewrites a record. + +## Changes Since rc.4 + +0.3.0-rc.5 adds the concurrent-edit semantics of Chapter 12A and relaxes +several write-time checks. Collections keep `spec_version: "0.3.0"` and need +no edit. Record files, type files, contracts, type packs, view records, and +queries that were valid under rc.4 stay valid. + +Most changes are relaxations: a write that rc.4 rejected now succeeds and is +reported. A few rules are tightened, and each tightening is reported as a +diagnostic rather than by rejecting a collection, so that no collection that +worked under rc.4 stops loading. + +### Behavior changes + +| Area | rc.4 | rc.5 | To keep the rc.4 behavior | +| --- | --- | --- | --- | +| `collection.unique` | at level `error`, a write that creates a duplicate fails | duplicates are reported and never block a write (`enforce: report` is the default) | add `enforce: write` to the rule | +| `unique.scope` | listed, but its comparison set was underspecified | a rule governs records of its type; `scope` selects the records they must differ from; raw values compare without coercion; the default scope is `type` | nothing; check reports for records in the now exact scope | +| link `validate_exists` and `target_type` | at level `error`, a write with a broken link fails | reported, never blocking | no equivalent; cross-record checks never block | +| a successful write with cross-record issues | not applicable | returns `valid: true` with `warning` diagnostics | nothing | +| concurrent edits | whole-record revision checks | tools that reconcile edits merge field by field (Chapter 12A); `if_revision` is opt-in and never added on a caller's behalf | pass `if_revision` explicitly | +| lifecycle `now` and `today` fields, `tags`, `uniqueItems` arrays | no merge semantics | merge as `max`, `union`, and `union` by default | declare `conflict` in `collection.merge` | +| update | `patch`, `unset`, `body`, `document` | adds `add` and `remove` list operations | nothing | +| derived path already taken | `path_conflict` | the first free suffixed path | supply an explicit path to get an error instead | +| structured writes to Markdown frontmatter | only array and object structure SHOULD be preserved | MUST re-emit only changed entries | nothing | +| structured writes to YAML document records | could re-emit the whole mapping | follow writer format fidelity | nothing | +| filename link tiebreakers | SHOULD | MUST, ending in code-point order | nothing | + +### Tightenings reported as diagnostics + +| Rule | How an engine reports a collection that relies on the rc.4 behavior | +| --- | --- | +| `match.expr` should not call `now()` or `today()`, read `file.mtime` or `file.ctime`, or follow links (Chapter 07) | the type still loads and its expression is evaluated as before; the engine reports a `warning` with code `nondeterministic_match`, the type name, and `details.binding`. In 0.3.0 stable this becomes an error that invalidates the type, so authors must move such logic before then | +| a `collection.path.pattern` value may not contain `/` or `\`, begin with `.`, or be empty (Chapter 07) | existing records are unaffected, whatever their paths; only a create that would derive such a path fails, with `path_value_invalid` naming the field. rc.4 already called such values invalid, but engines created folders from them | +| paths equal under case folding and NFC name one record path (Chapter 02) | existing files that collide stay records and each reports a `warning` with code `path_collision` and `details.paths`; an explicit create or rename onto an equivalent path fails with `path_conflict` | +| an ambiguous configured ID does not fall back to filename resolution (Chapter 08) | the link resolves to null and the referring record reports an `ambiguous_link` warning with `details.candidates`. rc.4 already required this; engines that fell back are non-conforming | +| `settings.id_field` has no default (Chapter 04) | rc.4 already required this; an engine that resolved through `id` without configuration is non-conforming. A collection that relies on ID resolution adds `id_field: id` | + +### What needs no change + +- `spec_version`, which stays `0.3.0` +- record files, frontmatter, and bodies +- JSON Schemas, read defaults, links, projections, and display metadata +- lifecycle policies, apart from the merge defaults above +- CEL queries, collection projections, view records, and lifecycle guards, + which may still use `now()` and `today()` +- data contracts and their digests, type packs, and `mdbase.lock.yaml` +- the event/action interoperability profile 0.1 and the durable runtime + profile 0.2 + +### Engines + +An engine moving from rc.4 to rc.5: + +- stops rejecting writes for cross-record issues, other than + `enforce: write` uniqueness rules, and reports those issues as warnings on + successful writes +- implements `add` and `remove`, path keys and the collision rule, + `path_value_invalid`, and the writer format fidelity rule +- stops supplying `if_revision` on a caller's behalf +- implements the merge of Chapter 12A if it reconciles concurrent edits, and + claims the `merge` profile +- implements move detection if it reports renames of externally moved files +- reports `nondeterministic_match`, `path_collision`, and `ambiguous_link` +- removes any `id` default for `settings.id_field` + +Engines that conform to rc.4 keep claiming `0.3.0-rc.4` until they pass the +rc.5 suite. The [rc.5 release notes](./docs/releases/0.3.0-rc.5.md) list every +conformance test whose expected outcome changed, which is where an rc.4 engine +diverges. + +## From v0.2 To v0.3 + +The rest of this chapter migrates v0.2.x collections into the v0.3 source +model. Migration tooling should produce: @@ -17,7 +95,7 @@ Migration analyzes the complete collection. Before a write, tooling MUST validate every existing record against the proposed target types and include incompatible records in the report. -## Configuration +### Configuration Configuration migration preserves which files are records and how they validate and resolve: @@ -45,7 +123,7 @@ validate and resolve: v0.3 meaning. Migration moves them under `x-legacy-v0.2` in the configuration and migrates strictness into each type's `additionalProperties`. -## Type Mapping +### Type Mapping | v0.2.x feature | v0.3 destination | | --- | --- | @@ -71,7 +149,7 @@ validate and resolve: | `display_name_key` | `collection.display.name_field` | | `extends` | JSON Schema `$ref`/`allOf` or explicit duplication | -## Defaults +### Defaults Migration should distinguish: @@ -83,7 +161,7 @@ For current mdbase field defaults, the safest migration is to emit both JSON Schema `default` and `collection.read_defaults` for static scalar defaults, with a report explaining the difference. -## Generated Fields +### Generated Fields Generated fields migrate to lifecycle. v0.2 generated values apply only when the field is missing, while lifecycle `set` always assigns, so migration adds a @@ -97,8 +175,16 @@ field is missing, while lifecycle `set` always assigns, so migration adds a | `ulid` | guarded `on_create` action setting `{ ulid: true }` | | `slugify from field` | guarded `on_create` action setting `{ slugify: source }` | | `from field` without a transform | guarded `on_create` action setting `{ copy: source }` | +| `sequence` | no destination; reported as unsupported | + +`sequence` assigned "largest existing value plus one". It has no portable +destination because concurrent creates on different devices allocate the same +number unless every create is coordinated (Chapter 09). Migration reports each +sequence field and recommends a guarded `ulid` or `uuid` action. Tools that +must keep allocating numbers do so through an implementation provider under +an `x-*` extension. -## Computed Fields +### Computed Fields Computed fields migrate to one of: @@ -108,7 +194,7 @@ Computed fields migrate to one of: - unsupported note when the computation has side effects or depends on non-portable functions -## Expressions +### Expressions Current mdbase expressions migrate to CEL where possible. @@ -116,7 +202,7 @@ Tool-specific expression dialects are adapter concerns. Tools may translate them to CEL for portable storage and translate them back for user interfaces or exports. -## Runtime Workflows +### Runtime Workflows Current generated-field and tool-conforming behavior that causes mutation should be reviewed as lifecycle or workflow behavior. @@ -126,7 +212,7 @@ Generated IDs and timestamps usually become lifecycle. Cross-record behavior, agent work, approval flows, external APIs, and scheduled checks become workflows and action/event contracts. -## Data Contract Implementations +### Data Contract Implementations Portable application interfaces migrate to a local data contract plus a type-local `implements` entry. @@ -150,7 +236,7 @@ contracts merely because they contain keys named `contract` or `version`. Private annotations with no portable contract meaning remain under a namespaced `x-*` section. They SHOULD NOT become JSON Schema custom keywords. -## Version Detection +### Version Detection A v0.2.x type file usually has `name` and `fields` without `kind: mdbase.type`. diff --git a/14-conformance.md b/14-conformance.md index 56a55c5..594f34a 100644 --- a/14-conformance.md +++ b/14-conformance.md @@ -16,6 +16,7 @@ queries, writes, runtime preflight, workflow execution, and watching. | Query | evaluate contextual CEL filters, projections, grouping, summaries, and query envelopes | | Links | parse, resolve, validate, and traverse links | | Core Write | create, update, delete, rename, and batch records | +| Merge | merge concurrent versions of a record with declared field strategies | | Type Packs | assess and transactionally apply managed type packs | | Lifecycle | apply standard managed-field policy during writes | | Event/Action Interoperability | exchange CloudEvents and admitted action invocations through independently claimable roles | @@ -34,6 +35,7 @@ Normative profile IDs and dependencies are: | `cel_query` | `collection_semantics`, `cel` | | `links` | `collection_semantics`, `cel` | | `core_write` | `collection_semantics` | +| `merge` | `core_write` | | `type_packs` | `data_contracts`, `core_write` | | `lifecycle` | `core_write`, `cel` | | `event_action_interop/0.1` | none | @@ -68,6 +70,14 @@ The canonical claim schema enforces profile dependencies. Claim verification tools SHOULD reject evidence produced for a different implementation artifact or specification version and SHOULD report stale evidence. +### Conformance Suites + +The shared fixtures live in `tests/v0.3/`, and `tests/v0.3/README.md` defines +their format. That includes the `merge_records` form, which states a merge case +as a base version, two edited versions, and the type declarations, with the +exact merged bytes and conflicts as the expected result. Each test added in a +release candidate after rc.4 records it with `since: 0.3.0-rc.5`. + ### Portable Interoperability Testbed The spec-owned interoperability testbed in `testbed/v0.1/` supplements the @@ -138,8 +148,10 @@ frontmatter selector. Implementations MAY add fields under `x-*`. The v0.3 core codes include `invalid_request`, `duplicate_batch_path`, `unsupported_profile`, `unsupported_feature`, `invalid_frontmatter`, `expression_compile_error`, `expression_evaluation_error`, -`projection_shadowed`, `link_not_found`, `type_conflict`, -`type_membership_changed`, `path_value_missing`, `schema_ref_forbidden`, +`projection_shadowed`, `link_not_found`, `link_target_type_mismatch`, +`ambiguous_link`, `duplicate_value`, `path_conflict`, `path_collision`, +`type_conflict`, `type_membership_changed`, `path_value_missing`, +`path_value_invalid`, `nondeterministic_match`, `schema_ref_forbidden`, `schema_ref_unresolved`, `schema_ref_cycle`, `format_invalid`, `lifecycle_expression_error`, `concurrent_modification`, `invalid_query`, `context_not_found`, `context_required`, `context_type_mismatch`, @@ -171,6 +183,9 @@ Core Read implementations MUST: - reject a type that requires an unsupported optional profile with `unsupported_profile` - report diagnostics in the canonical machine-readable shape +- read, index, and report invalid records rather than skip them +- compute path keys and report `path_collision` for discovered records with + equivalent paths ## Data Contracts Requirements @@ -208,7 +223,11 @@ Collection Semantics implementations MUST: - apply `collection.read_defaults` to effective reads - preserve missing, null, raw, and effective distinctions -- validate every `collection.unique` rule in its declared scope +- validate every `collection.unique` rule over its governed records and + comparison set, comparing raw values without coercion +- report cross-record issues without blocking writes, except for + `enforce: write` uniqueness rules +- load `collection.merge` declarations and derive default merge strategies - validate portable `collection.path` policies for write-capable tools - expose display metadata as advisory values - compose compatible behavior from multiple matched types @@ -243,6 +262,8 @@ CEL Match implementations MUST: - combine it with other members of `match` using AND - match only a boolean true result - report per-record evaluation errors and treat that candidate as a non-match +- report `nondeterministic_match` for a `match.expr` that uses a + time-dependent or cross-record binding ## Query Requirements @@ -328,9 +349,12 @@ Links implementations MUST: - parse wikilinks, Markdown links, and bare path link values - resolve collection-relative and file-relative paths safely -- enforce `collection.links.target_type` and `validate_exists` +- report `collection.links.target_type` and `validate_exists` issues as + cross-record issues - resolve simple wikilinks by filename, using ID resolution only when `settings.id_field` is configured +- resolve ambiguous IDs to null without filename fallback, and report + `ambiguous_link` - expose `file.links`, `file.embeds`, `file.tags`, and `file.backlinks` - provide the CEL link helpers from Chapter 10 - bound `asFile()` traversal @@ -339,10 +363,18 @@ Links implementations MUST: Core Write implementations MUST: -- validate a complete draft before writing -- preserve unrelated Markdown body content where possible +- validate a complete draft before writing, rejecting only request and + safety failures and, at level `error`, single-record issues +- preserve unrelated Markdown body content +- follow the writer format fidelity rule of Chapter 12A +- apply `add` and `remove` list operations to the current value - reject paths that escape the collection root -- enforce `if_revision` and report common concurrency conflicts +- reject an explicit path that collides under path equivalence with + `path_conflict`, give a colliding derived path the first free suffix, and + reject invalid path-pattern values with `path_value_invalid` +- enforce `enforce: write` uniqueness rules in a single write order +- enforce `if_revision` when the caller supplies it, and never supply it on a + caller's behalf - persist null patch values as explicit null and remove `unset` keys - apply the validation level to record validation issues - execute batches atomically by default, support `allow_partial` and @@ -350,6 +382,24 @@ Core Write implementations MUST: - return the canonical operation envelope and final record revision - update derived state before reporting a successful mutation +## Merge Requirements + +Merge implementations MUST: + +- merge frontmatter one top-level key at a time with the three-way rules of + Chapter 12A +- apply declared and default `conflict`, `max`, `min`, and `union` + strategies exactly +- treat end-of-body append-append as a union, earlier-ordered side first +- merge other body changes line by line and report overlapping changes as a + `body` conflict +- keep the first version's value at every conflict and report the conflict + with its base, first, and second values +- never reject a merged version or turn it into a conflict because it is + invalid +- copy values taken unchanged from one side verbatim from that side's source + text + ## Lifecycle Requirements Lifecycle implementations MUST: @@ -406,9 +456,9 @@ An implementation that also claims `data_contracts` reports `contract_changed` with the contract file `path` after the contract registry and type implementations have been re-resolved. Implementations MAY also report `schema_changed`, `view_changed`, and `lock_changed`, each with a `path`, for -referenced local schema files, saved-view sources, and `mdbase.lock.yaml`. A rename may be -reported as `record_deleted` followed by `record_created` when the host cannot -establish file identity. +referenced local schema files, saved-view sources, and `mdbase.lock.yaml`. A rename that move +detection (Chapter 12A) does not pair is reported as `record_deleted` followed +by `record_created`. `record_created` and `record_modified` include current effective frontmatter. `record_renamed` includes `path`, `previous_path`, and current effective @@ -429,3 +479,5 @@ Watch implementations MUST: - coalesce duplicate host notifications for one logical change and report the final observed state - isolate listener failures so later notifications continue to be delivered +- pair disappearances and appearances with the move detection of Chapter 12A + when reporting renames of files changed outside the implementation diff --git a/README.md b/README.md index b9a92e5..6353fc1 100644 --- a/README.md +++ b/README.md @@ -10,7 +10,7 @@ update files safely. The durable data stays in ordinary Markdown files. Collections remain readable in a text editor, reviewable in Git, and usable across conforming tools. -The current specification is **v0.3.0**. +The current specification is **v0.3.0** (release candidate 5). [Read the overview](./00-overview.md) · [Visit mdbase.dev](https://mdbase.dev) · diff --git a/schemas/v0.3/config.schema.json b/schemas/v0.3/config.schema.json index a0208a2..838cfef 100644 --- a/schemas/v0.3/config.schema.json +++ b/schemas/v0.3/config.schema.json @@ -55,7 +55,8 @@ }, "id_field": { "type": "string", - "minLength": 1 + "minLength": 1, + "description": "Field used for ID-based wikilink resolution and as a move-detection identity hint. No default." }, "exclude": { "type": "array", diff --git a/schemas/v0.3/conformance-claim.schema.json b/schemas/v0.3/conformance-claim.schema.json index 442940d..f66cf4e 100644 --- a/schemas/v0.3/conformance-claim.schema.json +++ b/schemas/v0.3/conformance-claim.schema.json @@ -203,6 +203,10 @@ "if": { "properties": { "profiles": { "contains": { "const": "core_write" } } } }, "then": { "properties": { "profiles": { "contains": { "const": "collection_semantics" } } } } }, + { + "if": { "properties": { "profiles": { "contains": { "const": "merge" } } } }, + "then": { "properties": { "profiles": { "contains": { "const": "core_write" } } } } + }, { "if": { "properties": { "profiles": { "contains": { "const": "lifecycle" } } } }, "then": { @@ -257,6 +261,7 @@ "cel_query", "links", "core_write", + "merge", "type_packs", "lifecycle", "event_action_interop/0.1", diff --git a/schemas/v0.3/type-file.schema.json b/schemas/v0.3/type-file.schema.json index dc0ca0e..7e9f928 100644 --- a/schemas/v0.3/type-file.schema.json +++ b/schemas/v0.3/type-file.schema.json @@ -233,6 +233,9 @@ }, "projections": { "$ref": "#/$defs/projections" + }, + "merge": { + "$ref": "#/$defs/mergeStrategies" } }, "patternProperties": { @@ -316,6 +319,10 @@ "path_glob": { "type": "string", "minLength": 1 + }, + "enforce": { + "enum": ["report", "write"], + "default": "report" } }, "allOf": [ @@ -359,6 +366,20 @@ }, "additionalProperties": false }, + "mergeStrategies": { + "description": "Per-field merge strategies for top-level frontmatter fields (Chapter 07).", + "type": "object", + "minProperties": 1, + "propertyNames": { + "anyOf": [ + { "$ref": "#/$defs/fieldName" }, + { "pattern": "^/([^/~]|~0|~1)+$" } + ] + }, + "additionalProperties": { + "enum": ["conflict", "max", "min", "union"] + } + }, "projections": { "type": "object", "propertyNames": { diff --git a/site/build.mjs b/site/build.mjs index c0ea2ee..cd6d3ba 100644 --- a/site/build.mjs +++ b/site/build.mjs @@ -86,6 +86,7 @@ const SPEC_FILES = [ { file: '10-cel-profile.md', num: '10', title: 'CEL Profile', id: 'section-10' }, { file: '11-querying.md', num: '11', title: 'Querying', id: 'section-11' }, { file: '12-operations.md', num: '12', title: 'Operations', id: 'section-12' }, + { file: '12a-concurrent-edits.md', num: '12A', title: 'Concurrent Edits', id: 'section-12a' }, { file: '13-migrations-and-compatibility.md', num: '13', title: 'Migrations & Compatibility', id: 'section-13' }, { file: '14-conformance.md', num: '14', title: 'Conformance', id: 'section-14' }, { file: 'interop/0.1.md', num: 'I', title: 'Event & Action Interoperability', id: 'companion-interop', group: 'Companion Profiles' }, @@ -244,7 +245,7 @@ function build() { entries: SPEC_FILES, output: 'spec.html', title: 'Specification', - version: 'v0.3.0', + version: 'v0.3.0-rc.5', switchLink: 'v0.2 archive', }); buildSpec({ From 93687bf219bf2f89bba0c80d8deed4c44a4236b9 Mon Sep 17 00:00:00 2001 From: callumalpass Date: Sun, 4 Oct 2026 14:29:57 +1100 Subject: [PATCH 2/5] Add the regex profile and body_edits to rc.5 Regex: CEL matches(), JSON Schema pattern, and match.where matches share one flavor, the mdbase regex profile defined in chapter 10: RE2 syntax with ASCII-only \d, \w, \s, \b and case folding over Unicode scalar values, as in regex-lite. Unicode classes, backreferences, and look-around are invalid_pattern and make a type invalid. JSON Schema pattern following the profile instead of ECMA-262 is provisional. Body edits: update accepts body_edits, ranges whose offsets count Unicode scalar values of a base body identified by a SHA-256 body_base digest, with optional body_base_text. Edits apply directly to an unchanged body and otherwise rebase with the chapter 12A body merge; conflicts are found per line and fail with body_conflict, and a missing base fails with body_base_unavailable. They are exclusive with body and document and need no live-collaboration support. --- 06-json-schema-profile.md | 16 +++++++ 07-collection-semantics.md | 7 +-- 10-cel-profile.md | 64 +++++++++++++++++++++++++++ 12-operations.md | 71 +++++++++++++++++++++++++++++- 12a-concurrent-edits.md | 9 ++++ 13-migrations-and-compatibility.md | 6 ++- 14-conformance.md | 10 ++++- 7 files changed, 177 insertions(+), 6 deletions(-) diff --git a/06-json-schema-profile.md b/06-json-schema-profile.md index aaf667f..2586e74 100644 --- a/06-json-schema-profile.md +++ b/06-json-schema-profile.md @@ -60,6 +60,22 @@ Core v0.3 JSON Schema support includes: Tools MAY support more of JSON Schema 2020-12. The required profile defines the portable baseline for Core Read conformance. +## Pattern Dialect + +`pattern`, and `patternProperties` where a tool supports it, use the mdbase +regex profile defined in Chapter 10: RE2 syntax with ASCII-only `\d`, `\w`, +`\s`, `\b`, and case-insensitive matching, over Unicode scalar values and +unanchored. A pattern outside the profile, such as one using `\p{L}`, a +backreference, or look-around, makes the type file invalid with an `invalid_pattern` +diagnostic. + +**Provisional (rc.5).** JSON Schema 2020-12 recommends ECMA-262 regular +expressions. The profile departs from it so that one engine serves schema +validation, matching, and CEL. The two agree on `\d`, `\w`, and `\b` for +patterns without the ECMA-262 `u` flag; they differ on `\s`, which ECMA-262 +defines over Unicode white space, and on look-around and backreferences, +which ECMA-262 supports. + ## Discriminated Unions v0.3 uses ordinary JSON Schema constructs for discriminated unions. diff --git a/07-collection-semantics.md b/07-collection-semantics.md index 7475242..71ff94b 100644 --- a/07-collection-semantics.md +++ b/07-collection-semantics.md @@ -128,9 +128,10 @@ null. Ordering evaluates to false for incomparable values. `neq` also evaluates to false for a missing field. A predicate with an operand of the wrong type evaluates to false. -`matches` uses the same portable regular-expression subset as JSON Schema -`pattern`: Unicode-aware matching without backreferences or look-around. An -unsupported or invalid pattern is a type-file diagnostic. +`matches` uses the mdbase regex profile (Chapter 10), the same flavor as JSON +Schema `pattern` and CEL `matches()`: RE2 syntax with ASCII-only classes and +case folding. An unsupported or invalid pattern makes the type file invalid +with an `invalid_pattern` diagnostic. ### CEL Matching diff --git a/10-cel-profile.md b/10-cel-profile.md index 9ceb051..18bd5d3 100644 --- a/10-cel-profile.md +++ b/10-cel-profile.md @@ -37,6 +37,7 @@ Hosts MUST support: selection syntax, `optional.of`, `optional.none`, `hasValue()`, `value()`, `or()`, and `orValue()` - the mdbase host functions and bindings defined in this chapter +- the mdbase regex profile below for `matches()` mdbase adds functions only under names that the CEL standard library does not define, so hosts never need to overload a standard function or operator. @@ -298,6 +299,69 @@ case-insensitive search lowercases the text it searches: file.body.lower().contains("mission body") ``` +## Regular Expressions + +mdbase has one regular-expression flavor, the **mdbase regex profile**. It +applies to CEL `matches()`, to JSON Schema `pattern` (Chapter 06), and to the +`matches` operator of `match.where` (Chapter 07). Every tool evaluates a +pattern the same way on every platform, which matters because a tool that +replays or re-verifies a write must reach the same result as the tool that +made it. + +The profile is RE2 syntax with ASCII-only character classes, the semantics of +the Rust `regex-lite` crate. CEL specifies RE2 syntax for `matches()`; the +profile keeps that syntax and removes Unicode classes from it. + +**Syntax.** A pattern may use: + +- literal characters; `\` before any ASCII punctuation character, such as + `\.` or `\\`; and the escapes `\n`, `\t`, `\r`, `\f`, `\v`, `\xHH`, and + `\x{H...}` for any Unicode scalar value +- `.`, which matches any one Unicode scalar value except `\n` (any scalar + value with the `s` flag) +- bracket classes such as `[a-z]`, `[^0-9]`, and ASCII class names such as + `[[:alpha:]]` +- the Perl classes `\d`, `\w`, `\s`, and their negations `\D`, `\W`, `\S` +- the anchors `^`, `$`, `\A`, `\z`, and the word boundaries `\b` and `\B` +- groups `(...)`, non-capturing groups `(?:...)`, and named groups + `(?P...)` +- alternation `|`, and the repetitions `*`, `+`, `?`, `{n}`, `{n,}`, + `{n,m}`, each optionally followed by `?` for lazy matching +- the flags `i`, `m`, `s`, and `x`, set as `(?flags)` or `(?flags:...)` + +A pattern that uses anything else is invalid. In particular, Unicode classes +such as `\p{L}` and `\pN`, backreferences, and look-around are invalid. + +**Semantics.** Matching runs over Unicode scalar values, and a match is +unanchored unless the pattern anchors it, as in RE2 and JSON Schema. Among +matches at the same position, the leftmost-first (Perl) alternative wins. +Classes and case folding are ASCII only: + +| Construct | Matches | +| --- | --- | +| `\d` | `[0-9]` | +| `\w` | `[0-9A-Za-z_]` | +| `\s` | `[\t\n\v\f\r ]` | +| `\b` | a boundary between a `\w` character and a non-`\w` character or the text edge | +| `(?i)` | ASCII letters case-insensitively; every other scalar value only as itself | + +So `"é".matches("^\\w$")` is false, `"café".matches("caf\\b")` is true, +`"١".matches("^\\d$")` is false, and `"É".matches("(?i)é")` is false. A +Unicode-aware case-insensitive search lowercases its text first, for example +`title.lower().matches("^éclair")`, and a pattern that needs non-ASCII letters +lists them, as in `[a-zà-ÿ]`. + +**Errors.** An invalid pattern written as a CEL string literal is an +`expression_compile_error` when the expression compiles. A pattern computed +during evaluation that turns out invalid raises an evaluation error. An +invalid JSON Schema `pattern` or `match.where` pattern makes its type file +invalid. + +**Provisional (rc.5).** The profile follows `regex-lite` because it keeps the +WebAssembly runtime about 740 KB smaller than a Unicode-table engine and can +run identically everywhere. Patterns that depend on Unicode classes behave +differently from rc.4 engines that used a Unicode-aware engine. + ## File And Link Helpers Core Read supplies file metadata and: diff --git a/12-operations.md b/12-operations.md index f6f0d43..ef003bc 100644 --- a/12-operations.md +++ b/12-operations.md @@ -104,7 +104,8 @@ Update modifies an existing record. Pipeline: 1. read existing raw frontmatter -2. apply the requested `patch`, `unset`, `add`, and `remove` +2. apply the requested `patch`, `unset`, `add`, and `remove`, and the body + replacement or `body_edits` 3. re-match and freeze types when type-affecting fields or path changed 4. apply lifecycle `on_update` 5. verify type membership did not change as a lifecycle side effect @@ -127,6 +128,8 @@ A structured update accepts: remove from that field's list - `body`: optional replacement Markdown body; a non-empty body for a YAML document record (Chapter 03) is `invalid_request` +- `body_edits` with `body_base`: optional text edits to the body against a + known base body, as an alternative to `body`; see below Unsetting a key that is already missing is not an error. A request that names the same field in `patch` and `unset`, names a field inside a key that `patch` @@ -172,6 +175,72 @@ not require `uniqueItems`. A list emptied by `remove` stays as an empty list. **Provisional (rc.5).** `add` and `remove` address top-level fields only, like `patch`, because the merge unit is the top-level field. +### Body edits + +`body_edits` changes parts of the body instead of replacing all of it. It is +what an editor buffer naturally produces, and it describes the writer's change +precisely enough to combine with a concurrent edit: + +```yaml +path: notes/meeting.md +body_base: sha256:2c26b46b68ffc68ff99b453c1d30413413422d706483bfa0f98a5e886266e7ae +body_edits: + - { start: 0, end: 5, text: "Agenda" } + - { start: 42, end: 42, text: "\n- follow up with Bo" } +``` + +- `body_base` identifies the body the edits were made against: `sha256:` + followed by the lowercase hexadecimal SHA-256 digest of the body's exact + UTF-8 bytes, as defined in Chapter 03 (everything after the closing + frontmatter delimiter, or the whole file without frontmatter). +- `body_edits` is a list of edits. Each replaces the text from offset `start` + up to, but not including, offset `end` of the base body with `text`. + `start == end` is an insertion and an empty `text` is a deletion. +- Offsets count Unicode scalar values from the start of the base body, so + every offset falls between two characters whatever the encoding. +- Every offset refers to the base body, not to the body after earlier edits. + Edits are sorted by `start` and do not overlap: each edit's `end` is at most + the next edit's `start`, and two insertions at the same offset are one edit. +- `body_base_text` optionally carries the complete base body. When present, + its digest MUST equal `body_base`. + +The update applies the edits as follows: + +1. If the record's current body has the digest `body_base`, the edits are + applied to it directly. +2. Otherwise the implementation obtains the base body, from `body_base_text` + or from bodies it has retained, and rebases the edits onto the current + body with the three-way body merge of Chapter 12A, where the current body + is the first version and the base with the edits applied is the second. + A merged body is written. A body conflict fails the update with + `concurrent_modification`, `details.reason: body_conflict`, and the + conflict; nothing is written. +3. If the base body cannot be obtained, the update fails with + `concurrent_modification` and `details.reason: body_base_unavailable`. + The caller can retry with `body_base_text` or against the current body. + +A request is `invalid_request` before any write when it combines `body_edits` +with `body` or `document`, has `body_edits` without `body_base`, has an +offset outside the base body, or has edits out of order or overlapping. Body +edits for a YAML document record are `invalid_request`, because such a record +has no body. Body edits combine with `patch`, `unset`, `add`, and `remove` in +one update, and lifecycle runs once for the whole update. + +Body edits need no live-collaboration machinery. An implementation that only +ever applies them directly, case 1 above, conforms, as long as it reports the +other cases as specified. + +Conflicts are found at line granularity, because they come from the +line-based body merge of Chapter 12A: two edits to different words of one line +conflict. A later release may detect conflicts at a finer granularity; that +is a compatible refinement, because it only turns some conflicts into merges. + +**Provisional (rc.5).** Offsets are Unicode scalar values, the unit that +regular expressions and `.` match over. Byte offsets could fall inside a +character, and UTF-16 code units, which JavaScript editors use, are +specific to one platform; adapters convert at their boundary. Retaining old +bodies is optional. + As an alternative to a frontmatter patch and body replacement, Update accepts `document` containing the complete candidate source in the record's format: Markdown source for a Markdown record, or the whole YAML document for a YAML diff --git a/12a-concurrent-edits.md b/12a-concurrent-edits.md index f49cb52..5305494 100644 --- a/12a-concurrent-edits.md +++ b/12a-concurrent-edits.md @@ -248,6 +248,15 @@ unique. of the body. Two insertions at the same place inside the body remain a conflict. +### Body edits + +An update with `body_edits` (Chapter 12) is a body merge whose second version +is built from the request: the base body with the edits applied. The current +body is the first version. Append-append applies as usual, so an edit that +only appends to the end of the base combines with a concurrent append, the +current text first. An edit range that overlaps a concurrent change to the +same lines is a body conflict. + ### Path When the first and second versions have different paths, the path merges like diff --git a/13-migrations-and-compatibility.md b/13-migrations-and-compatibility.md index f08ca23..8b6b099 100644 --- a/13-migrations-and-compatibility.md +++ b/13-migrations-and-compatibility.md @@ -30,7 +30,8 @@ worked under rc.4 stops loading. | a successful write with cross-record issues | not applicable | returns `valid: true` with `warning` diagnostics | nothing | | concurrent edits | whole-record revision checks | tools that reconcile edits merge field by field (Chapter 12A); `if_revision` is opt-in and never added on a caller's behalf | pass `if_revision` explicitly | | lifecycle `now` and `today` fields, `tags`, `uniqueItems` arrays | no merge semantics | merge as `max`, `union`, and `union` by default | declare `conflict` in `collection.merge` | -| update | `patch`, `unset`, `body`, `document` | adds `add` and `remove` list operations | nothing | +| update | `patch`, `unset`, `body`, `document` | adds `add` and `remove` list operations, and `body_edits` against a `body_base` digest | nothing | +| regular expressions in CEL `matches()`, JSON Schema `pattern`, and `match.where` | engine-dependent; Unicode-aware classes were implied | one profile everywhere: RE2 syntax with ASCII-only `\d`, `\w`, `\s`, `\b`, and case folding (Chapter 10); `\p{...}` is invalid | none; list non-ASCII characters explicitly, or lowercase text before matching | | derived path already taken | `path_conflict` | the first free suffixed path | supply an explicit path to get an error instead | | structured writes to Markdown frontmatter | only array and object structure SHOULD be preserved | MUST re-emit only changed entries | nothing | | structured writes to YAML document records | could re-emit the whole mapping | follow writer format fidelity | nothing | @@ -44,6 +45,7 @@ worked under rc.4 stops loading. | a `collection.path.pattern` value may not contain `/` or `\`, begin with `.`, or be empty (Chapter 07) | existing records are unaffected, whatever their paths; only a create that would derive such a path fails, with `path_value_invalid` naming the field. rc.4 already called such values invalid, but engines created folders from them | | paths equal under case folding and NFC name one record path (Chapter 02) | existing files that collide stay records and each reports a `warning` with code `path_collision` and `details.paths`; an explicit create or rename onto an equivalent path fails with `path_conflict` | | an ambiguous configured ID does not fall back to filename resolution (Chapter 08) | the link resolves to null and the referring record reports an `ambiguous_link` warning with `details.candidates`. rc.4 already required this; engines that fell back are non-conforming | +| a pattern may not use Unicode classes such as `\p{L}`, backreferences, or look-around (Chapter 10) | a type whose JSON Schema `pattern` or `match.where` pattern does so is invalid with `invalid_pattern`; a CEL literal pattern is an `expression_compile_error`. Patterns that only use `\w`, `\d`, `\s`, `\b`, or `(?i)` stay valid but match only ASCII in those constructs, so an rc.4 engine with Unicode classes diverges on non-ASCII text | | `settings.id_field` has no default (Chapter 04) | rc.4 already required this; an engine that resolved through `id` without configuration is non-conforming. A collection that relies on ID resolution adds `id_field: id` | ### What needs no change @@ -65,6 +67,8 @@ An engine moving from rc.4 to rc.5: - stops rejecting writes for cross-record issues, other than `enforce: write` uniqueness rules, and reports those issues as warnings on successful writes +- evaluates every pattern with the mdbase regex profile +- implements `body_edits`, at least the direct case - implements `add` and `remove`, path keys and the collision rule, `path_value_invalid`, and the writer format fidelity rule - stops supplying `if_revision` on a caller's behalf diff --git a/14-conformance.md b/14-conformance.md index 594f34a..5fca9a0 100644 --- a/14-conformance.md +++ b/14-conformance.md @@ -151,7 +151,8 @@ The v0.3 core codes include `invalid_request`, `duplicate_batch_path`, `projection_shadowed`, `link_not_found`, `link_target_type_mismatch`, `ambiguous_link`, `duplicate_value`, `path_conflict`, `path_collision`, `type_conflict`, `type_membership_changed`, `path_value_missing`, -`path_value_invalid`, `nondeterministic_match`, `schema_ref_forbidden`, +`path_value_invalid`, `nondeterministic_match`, `invalid_pattern`, +`schema_ref_forbidden`, `schema_ref_unresolved`, `schema_ref_cycle`, `format_invalid`, `lifecycle_expression_error`, `concurrent_modification`, `invalid_query`, `context_not_found`, `context_required`, `context_type_mismatch`, @@ -180,6 +181,8 @@ Core Read implementations MUST: - validate embedded JSON Schema against the v0.3 profile - select explicit types and evaluate structured inferred match rules - validate raw frontmatter independently against every matched schema +- evaluate JSON Schema `pattern` and `match.where` `matches` with the mdbase + regex profile - reject a type that requires an unsupported optional profile with `unsupported_profile` - report diagnostics in the canonical machine-readable shape @@ -247,6 +250,8 @@ CEL implementations MUST: timezone behavior - provide the `lower()` and `upper()` text helpers with Unicode default case mappings +- evaluate `matches()` with the mdbase regex profile: RE2 syntax with + ASCII-only classes and case folding - enforce and report expression, evaluation, and traversal limits - distinguish compilation diagnostics from evaluation diagnostics @@ -368,6 +373,9 @@ Core Write implementations MUST: - preserve unrelated Markdown body content - follow the writer format fidelity rule of Chapter 12A - apply `add` and `remove` list operations to the current value +- apply `body_edits` directly when the body matches `body_base`, rebase them + with the Chapter 12A body merge when the base body is available, and report + `body_conflict` and `body_base_unavailable` otherwise - reject paths that escape the collection root - reject an explicit path that collides under path equivalence with `path_conflict`, give a colliding derived path the first free suffix, and From b4fd6b8c4f1a9b1bbf3618348e94d06473a64661 Mon Sep 17 00:00:00 2001 From: callumalpass Date: Sun, 4 Oct 2026 14:30:31 +1100 Subject: [PATCH 3/5] Add the rc.5 conformance cases to the v0.3 suite New tests carry since: 0.3.0-rc.5 and changed tests carry changed: 0.3.0-rc.5. The manifest marks the rc.5 requirements and adds the merge profile and the merge and watch fixture sets. - merge/merge.yaml converts the 13 mdbase-next prototype merge fixtures into a merge_records format (base + first + second + type declarations -> exact merged bytes and conflicts) and adds 24 merge and strategy cases; merge/body-edits.yaml covers applying and rebasing body_edits. - core/paths.yaml, watch/move-detection.yaml, and cel/regex-profile.yaml cover path keys and suffixes, move pairing, and ASCII-versus-Unicode regex classes. - Adapter-target suites cover validation tiers, uniqueness modes and scope, non-blocking link checks, list operations, format fidelity, path collisions, body_edits through update, nondeterministic_match, and the id_field and ambiguous-ID rules. - links.duplicate_id_ambiguous now also expects ambiguous_link. - scripts/concurrent_edits_model.py is an executable model of the pure functions; check_v03_tests.py runs the 95 pure fixtures against it, validates the since/changed markers, and checks that the claim schema lists the manifest's profiles. --- .github/workflows/ci.yml | 2 +- scripts/check_v03_tests.py | 32 +- scripts/concurrent_edits_model.py | 885 ++++++++++ tests/v0.3/README.md | 139 +- tests/v0.3/cel/match-determinism.yaml | 121 ++ tests/v0.3/cel/regex-profile.yaml | 314 ++++ tests/v0.3/core/body-edits-update.yaml | 156 ++ tests/v0.3/core/core-collection.yaml | 2 +- tests/v0.3/core/link-ambiguity.yaml | 84 + tests/v0.3/core/links-and-discovery.yaml | 7 +- tests/v0.3/core/list-ops-and-fidelity.yaml | 374 +++++ tests/v0.3/core/paths.yaml | 215 +++ tests/v0.3/core/validation-tiers.yaml | 449 +++++ tests/v0.3/manifest.yaml | 52 + tests/v0.3/merge/body-edits.yaml | 304 ++++ tests/v0.3/merge/merge.yaml | 1717 ++++++++++++++++++++ tests/v0.3/watch/move-detection.yaml | 433 +++++ 17 files changed, 5281 insertions(+), 5 deletions(-) create mode 100644 scripts/concurrent_edits_model.py create mode 100644 tests/v0.3/cel/match-determinism.yaml create mode 100644 tests/v0.3/cel/regex-profile.yaml create mode 100644 tests/v0.3/core/body-edits-update.yaml create mode 100644 tests/v0.3/core/link-ambiguity.yaml create mode 100644 tests/v0.3/core/list-ops-and-fidelity.yaml create mode 100644 tests/v0.3/core/paths.yaml create mode 100644 tests/v0.3/core/validation-tiers.yaml create mode 100644 tests/v0.3/merge/body-edits.yaml create mode 100644 tests/v0.3/merge/merge.yaml create mode 100644 tests/v0.3/watch/move-detection.yaml diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 12d0b48..2c209c6 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -18,7 +18,7 @@ jobs: run: python -m pip install PyYAML==6.0.3 jsonschema==4.26.0 - name: Legacy conformance level guardrail run: python scripts/check_test_levels.py - - name: Validate v0.3 fixture manifest and coverage + - name: Validate v0.3 fixture manifest, coverage, and executable fixtures run: python scripts/check_v03_tests.py - name: Verify migration fixture run: python scripts/prototype_tasknotes_v03_migration.py --check-fixture diff --git a/scripts/check_v03_tests.py b/scripts/check_v03_tests.py index 3a6648b..c4334c0 100755 --- a/scripts/check_v03_tests.py +++ b/scripts/check_v03_tests.py @@ -11,6 +11,9 @@ - YAML documents parse and optionally validate against schemas - simple YAML pointer presence checks - the TaskNotes migration prototype can satisfy fixture report assertions +- merge, path, and move-detection fixtures (0.3.0-rc.5) match the executable + model in scripts/concurrent_edits_model.py +- the conformance claim schema lists exactly the manifest's profiles Adapter-target tests for core collection behavior, lifecycle, CEL, and runtime execution are shape-checked but not executed here. @@ -40,6 +43,9 @@ sys.exit(1) +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import concurrent_edits_model # noqa: E402 + REPO_ROOT = Path(__file__).resolve().parent.parent TEST_ROOT = REPO_ROOT / "tests" / "v0.3" @@ -59,7 +65,9 @@ "data_contract_digest", "data_contract_implementation_digest", "data_contract_registry_validate", -} +} | concurrent_edits_model.FIXTURE_OPERATIONS + +RELEASE_MARKER = re.compile(r"^0\.3\.0-rc\.[0-9]+$") CONFORMANCE_PROFILES = { "core_read", @@ -70,6 +78,7 @@ "cel_query", "links", "core_write", + "merge", "type_packs", "lifecycle", "event_action_interop/0.1", @@ -139,6 +148,8 @@ def main() -> int: f"{manifest_path}: coverage_complete profile {profile_id} has uncovered requirements: {missing}" ) + check_claim_schema_profiles(errors) + if errors: print("\n".join(errors), file=sys.stderr) print(f"v0.3 suite check failed: {len(errors)} error(s), {executed} executable test(s), {skipped} adapter-target test(s)", file=sys.stderr) @@ -148,6 +159,16 @@ def main() -> int: return 0 +def check_claim_schema_profiles(errors: list[str]) -> None: + claim = load_json("schemas/v0.3/conformance-claim.schema.json") + listed = set(claim["$defs"]["profile"]["enum"]) + if listed != CONFORMANCE_PROFILES: + errors.append( + "schemas/v0.3/conformance-claim.schema.json profiles differ from the suite: " + f"missing {sorted(CONFORMANCE_PROFILES - listed)}, extra {sorted(listed - CONFORMANCE_PROFILES)}" + ) + + def validate_manifest( path: Path, errors: list[str] ) -> tuple[dict[str, dict[str, Any]], set[str], dict[str, set[str]]]: @@ -278,6 +299,11 @@ def validate_suite_shape( for key in ["name", "operation", "input", "expect"]: if key not in test: errors.append(f"{path}: test {test_index} in group {group.get('name', group_index)!r} missing {key}") + for marker in ("since", "changed"): + if marker in test and not RELEASE_MARKER.match(str(test[marker])): + errors.append( + f"{path}: test {test.get('id', test_index)!r} {marker} must name a 0.3.0 release candidate" + ) covers = test.get("covers") if covers is None: continue @@ -336,6 +362,10 @@ def run_executable_test(test: dict[str, Any], setup: dict[str, Any] | None = Non expect = test.get("expect") or {} setup = setup or {} + if operation in concurrent_edits_model.FIXTURE_OPERATIONS: + concurrent_edits_model.run_fixture(test, setup) + return + if operation == "json_schema_meta_validate": for path in expand_paths(input_data.get("paths", [])): Draft202012Validator.check_schema(load_json(path)) diff --git a/scripts/concurrent_edits_model.py b/scripts/concurrent_edits_model.py new file mode 100644 index 0000000..3ef645e --- /dev/null +++ b/scripts/concurrent_edits_model.py @@ -0,0 +1,885 @@ +#!/usr/bin/env python3 +"""Executable model of the pure concurrent-edit functions in Chapters 02, 07, and 12A. + +This is not an mdbase implementation. It models the functions that the +0.3.0-rc.5 conformance fixtures can check without a collection engine: + +- path keys, path equivalence, and the collision suffix rule (Chapter 02) +- path-pattern derivation and placeholder validation (Chapter 07) +- default and declared merge strategies (Chapter 07) +- the three-way record merge and writer format fidelity (Chapter 12A) +- move detection pairing (Chapter 12A) +- the mdbase regex profile (Chapter 10), approximated with Python `re` in + ASCII mode for the constructs the fixtures use +- body edits and their rebase onto a changed body (Chapters 12 and 12A) + +`scripts/check_v03_tests.py` runs the corresponding fixtures against it +through `run_fixture`. +The model exists so that the fixtures are self-consistent and so that the +normative text has a second, independent reading. Where the model and the +specification disagree, the specification wins. +""" + +from __future__ import annotations + +import json +import re +import unicodedata +from dataclasses import dataclass, field +from typing import Any + +import yaml + + +# --------------------------------------------------------------------- YAML + + +class _StringDateLoader(yaml.SafeLoader): + """Safe loader that keeps timestamps as strings (Chapter 03 YAML profile).""" + + +_StringDateLoader.yaml_implicit_resolvers = { + key: [(tag, regexp) for tag, regexp in resolvers if tag != "tag:yaml.org,2002:timestamp"] + for key, resolvers in yaml.SafeLoader.yaml_implicit_resolvers.items() +} + + +def load_yaml_text(text: str) -> Any: + return yaml.load(text, Loader=_StringDateLoader) + + +# ------------------------------------------------------------------ paths + + +def path_key(path: str) -> str: + """Chapter 02: NFC, full default case folding, NFC.""" + return unicodedata.normalize("NFC", unicodedata.normalize("NFC", path).casefold()) + + +def equivalence_groups(paths: list[str]) -> list[list[str]]: + groups: dict[str, list[str]] = {} + for path in paths: + groups.setdefault(path_key(path), []).append(path) + return sorted(sorted(group) for group in groups.values() if len(group) > 1) + + +def suffixed(path: str, n: int) -> str: + folder, _, name = path.rpartition("/") + prefix = f"{folder}/" if folder else "" + stem, dot, ext = name.rpartition(".") + if not dot: + return f"{prefix}{name} ({n})" + return f"{prefix}{stem} ({n}).{ext}" + + +def allocate_path(requested: str, existing: list[str]) -> str: + used = {path_key(path) for path in existing} + if path_key(requested) not in used: + return requested + n = 2 + while path_key(suffixed(requested, n)) in used: + n += 1 + return suffixed(requested, n) + + +class PathError(Exception): + def __init__(self, code: str, field_name: str | None = None): + super().__init__(code) + self.code = code + self.field = field_name + + +def derive_path(pattern: str, frontmatter: dict[str, Any]) -> str: + out = [] + rest = pattern + while "{" in rest: + open_index = rest.index("{") + close_index = rest.index("}", open_index) + out.append(rest[:open_index]) + name = rest[open_index + 1 : close_index] + value = frontmatter.get(name) + if value is None: + raise PathError("path_value_missing", name) + if isinstance(value, (list, dict)): + raise PathError("path_value_invalid", name) + if isinstance(value, bool): + text = "true" if value else "false" + elif isinstance(value, str): + text = value + else: + text = json.dumps(value) + if text == "" or "/" in text or "\\" in text or "\0" in text or text.startswith("."): + raise PathError("path_value_invalid", name) + out.append(text) + rest = rest[close_index + 1 :] + out.append(rest) + return "".join(out) + + +# --------------------------------------------------------------- equality + + +def values_equal(a: Any, b: Any) -> bool: + if isinstance(a, bool) or isinstance(b, bool): + return type(a) is type(b) and a == b + if isinstance(a, (int, float)) and isinstance(b, (int, float)): + return a == b + if isinstance(a, list) and isinstance(b, list): + return len(a) == len(b) and all(values_equal(x, y) for x, y in zip(a, b)) + if isinstance(a, dict) and isinstance(b, dict): + return a.keys() == b.keys() and all(values_equal(a[k], b[k]) for k in a) + return type(a) is type(b) and a == b + + +MISSING = object() + + +def states_equal(a: Any, b: Any) -> bool: + if a is MISSING or b is MISSING: + return a is b + return values_equal(a, b) + + +# ------------------------------------------------------------ documents + + +@dataclass +class Entry: + key: str + lines: list[str] + + +@dataclass +class Document: + source: str + has_frontmatter: bool + open_delim: str = "" + close_delim: str = "" + # Items are Entry or a raw interstitial line. + items: list[Any] = field(default_factory=list) + frontmatter: Any = None + body: str = "" + + def entry(self, key: str) -> Entry | None: + for item in self.items: + if isinstance(item, Entry) and item.key == key: + return item + return None + + +_ENTRY_KEY = re.compile(r"""^("(?:[^"\\]|\\.)*"|'(?:[^']|'')*'|[^\s#:'"][^:]*?)\s*:(?=\s|$)""") + + +def _split_lines(text: str) -> list[str]: + return text.splitlines(keepends=True) + + +def _parse_items(fm_lines: list[str]) -> list[Any]: + """Split frontmatter lines into top-level entries and interstitial lines.""" + items: list[Any] = [] + current: Entry | None = None + pending_blank: list[str] = [] + for line in fm_lines: + if not line.strip(): + pending_blank.append(line) + continue + continuation = line[0] in " \t" or line.startswith("- ") or line.rstrip("\r\n") == "-" + if continuation and current is not None: + current.lines.extend(pending_blank) + pending_blank = [] + current.lines.append(line) + continue + items.extend(pending_blank) + pending_blank = [] + match = _ENTRY_KEY.match(line) + if match and not line.startswith("#"): + raw_key = match.group(1) + key = load_yaml_text(raw_key) if raw_key[0] in "\"'" else raw_key + current = Entry(key=str(key), lines=[line]) + items.append(current) + else: + current = None + items.append(line) + items.extend(pending_blank) + return items + + +def parse_document(source: str) -> Document: + lines = _split_lines(source) + if not lines or lines[0].rstrip("\r\n") != "---" or not lines[0].endswith("\n"): + return Document(source=source, has_frontmatter=False, frontmatter={}, body=source) + for index in range(1, len(lines)): + if lines[index].rstrip("\r\n") == "---": + fm_lines = lines[1:index] + parsed = load_yaml_text("".join(fm_lines)) + return Document( + source=source, + has_frontmatter=True, + open_delim=lines[0], + close_delim=lines[index], + items=_parse_items(fm_lines), + frontmatter={} if parsed is None else parsed, + body="".join(lines[index + 1 :]), + ) + return Document(source=source, has_frontmatter=False, frontmatter={}, body=source) + + +def _line_ending(doc: Document) -> str: + return "\r\n" if "\r\n" in doc.source else "\n" + + +_PLAIN_SCALAR = re.compile(r"^[A-Za-z0-9_][A-Za-z0-9_./ -]*$") +_YAML_SPECIAL = {"true", "false", "null", "yes", "no", "on", "off", "~", ""} + + +def emit_scalar(value: Any) -> str: + if isinstance(value, str): + if _PLAIN_SCALAR.match(value) and value.lower() not in _YAML_SPECIAL and not value.endswith(" "): + if not re.match(r"^[0-9.+-]", value): + return value + return json.dumps(value, ensure_ascii=False) + return json.dumps(value, ensure_ascii=False) + + +def emit_entry(key: str, value: Any, style_from: Entry | None, eol: str) -> list[str]: + first = style_from.lines[0] if style_from else "" + comment = "" + if style_from and len(style_from.lines) == 1: + found = re.search(r"(\s+#.*)$", first.rstrip("\r\n")) + if found: + comment = found.group(1) + flow = style_from is None or len(style_from.lines) == 1 + if isinstance(value, list): + if flow or not value: + return [f"{key}: [{', '.join(emit_scalar(v) for v in value)}]{comment}{eol}"] + indent = " " + if style_from and len(style_from.lines) > 1: + indent = re.match(r"^(\s*)", style_from.lines[1]).group(1) or "" + return [f"{key}:{eol}"] + [f"{indent}- {emit_scalar(v)}{eol}" for v in value] + return [f"{key}: {emit_scalar(value)}{comment}{eol}"] + + +def render(doc: Document) -> str: + if not doc.has_frontmatter: + return doc.body + out = [doc.open_delim] + for item in doc.items: + out.extend(item.lines if isinstance(item, Entry) else [item]) + out.append(doc.close_delim) + out.append(doc.body) + return "".join(out) + + +# --------------------------------------------------------------- types + + +@dataclass +class TypeDef: + name: str + path_globs: list[str] + merge: dict[str, str] + lifecycle_time_fields: set[str] + unique_items_fields: set[str] + + +def glob_to_regex(glob: str) -> re.Pattern[str]: + parts = glob.split("/") + regex = "" + for index, part in enumerate(parts): + last = index == len(parts) - 1 + if part == "**": + regex += "(?:.*/)?" if not last else ".*" + continue + piece = "" + i = 0 + while i < len(part): + char = part[i] + if char == "*": + piece += "[^/]*" + elif char == "?": + piece += "[^/]" + elif char == "[": + end = part.index("]", i) + body = part[i + 1 : end] + if body.startswith("!"): + body = "^" + body[1:] + piece += f"[{body}]" + i = end + else: + piece += re.escape(char) + i += 1 + regex += piece + ("" if last else "/") + return re.compile(f"^{regex}$") + + +def load_type(text: str) -> TypeDef: + doc = parse_document(text) + fm = doc.frontmatter + match = fm.get("match") or {} + globs = match.get("path_glob") or [] + if isinstance(globs, str): + globs = [globs] + collection = fm.get("collection") or {} + lifecycle = fm.get("lifecycle") or {} + time_fields: set[str] = set() + for event in ("on_create", "on_update"): + actions = lifecycle.get(event) or [] + if isinstance(actions, dict): + actions = [actions] + for action in actions: + for target, provider in (action.get("set") or {}).items(): + if isinstance(provider, dict) and (provider.get("now") is True or provider.get("today") is True): + if "." not in target and "[" not in target and not target.startswith("/"): + time_fields.add(target) + unique_items: set[str] = set() + schema = (fm.get("schema") or {}).get("value") or {} + for prop, spec in (schema.get("properties") or {}).items(): + if isinstance(spec, dict) and spec.get("uniqueItems") is True: + unique_items.add(prop) + return TypeDef( + name=fm["name"], + path_globs=globs, + merge=dict(collection.get("merge") or {}), + lifecycle_time_fields=time_fields, + unique_items_fields=unique_items, + ) + + +def matched_types(types: list[TypeDef], path: str, frontmatter: Any) -> list[TypeDef]: + by_name = {t.name.lower(): t for t in types} + if isinstance(frontmatter, dict): + declared: list[str] = [] + for key in ("type", "types"): + value = frontmatter.get(key) + if isinstance(value, str): + declared.append(value) + elif isinstance(value, list): + declared.extend(v for v in value if isinstance(v, str)) + if any(k in frontmatter for k in ("type", "types")): + return [by_name[n.lower()] for n in dict.fromkeys(declared) if n.lower() in by_name] + found = [t for t in types if t.path_globs and any(glob_to_regex(g).match(path) for g in t.path_globs)] + return sorted(found, key=lambda t: t.name.lower()) + + +def strategy(types: list[TypeDef], key: str) -> str: + declared = {t.merge[key] for t in types if key in t.merge} + if len(declared) > 1: + raise ValueError(f"type_conflict: merge strategy for {key}") + if declared: + return declared.pop() + if any(key in t.lifecycle_time_fields for t in types): + return "max" + if key == "tags" or any(key in t.unique_items_fields for t in types): + return "union" + return "conflict" + + +# ---------------------------------------------------------------- merge + +_DATE_TIME = re.compile( + r"^(\d{4})-(\d{2})-(\d{2})[Tt](\d{2}):(\d{2}):(\d{2})(\.\d+)?([Zz]|[+-]\d{2}:\d{2})$" +) +_FULL_DATE = re.compile(r"^\d{4}-\d{2}-\d{2}$") + + +def _instant(value: str) -> tuple[int, str] | None: + from datetime import datetime, timezone + + match = _DATE_TIME.match(value) + if not match: + return None + text = value.upper().replace("Z", "+00:00") + fraction = match.group(7) or "" + base = text.replace(fraction, "") if fraction else text + moment = datetime.fromisoformat(base).astimezone(timezone.utc) + digits = (fraction[1:] + "000000000")[:9] if fraction else "000000000" + return (int(moment.timestamp()), digits) + + +def compare(a: Any, b: Any) -> int | None: + if isinstance(a, bool) or isinstance(b, bool): + return None + if isinstance(a, (int, float)) and isinstance(b, (int, float)): + return (a > b) - (a < b) + if isinstance(a, str) and isinstance(b, str): + ia, ib = _instant(a), _instant(b) + if ia is not None and ib is not None: + return (ia > ib) - (ia < ib) + return (a > b) - (a < b) + return None + + +def as_list(value: Any, key: str) -> list[Any] | None: + if value is MISSING or value is None: + return [] + if isinstance(value, list): + return list(value) + if key == "tags" and isinstance(value, str): + return [value] + return None + + +def contains(items: list[Any], item: Any) -> bool: + return any(values_equal(item, other) for other in items) + + +def union_merge(base: list[Any], first: list[Any], second: list[Any]) -> list[Any]: + out = [x for x in first if not (contains(base, x) and not contains(second, x))] + for item in second: + if not contains(base, item) and not contains(out, item): + out.append(item) + return out + + +@dataclass +class MergeResult: + document: str + path: str + conflicts: list[dict[str, Any]] + + +def merge_body(base: str, first: str, second: str) -> str | None: + if first == second or second == base: + return first + if first == base: + return second + if first.startswith(base) and second.startswith(base): + first_tail = first[len(base) :] + second_tail = second[len(base) :] + separator = "\n" if first_tail and not first_tail.endswith("\n") else "" + return base + first_tail + separator + second_tail + return diff3(_split_lines(base), _split_lines(first), _split_lines(second)) + + +def _lcs_pairs(a: list[str], b: list[str]) -> list[tuple[int, int]]: + n, m = len(a), len(b) + table = [[0] * (m + 1) for _ in range(n + 1)] + for i in range(n - 1, -1, -1): + for j in range(m - 1, -1, -1): + table[i][j] = table[i + 1][j + 1] + 1 if a[i] == b[j] else max(table[i + 1][j], table[i][j + 1]) + pairs = [] + i = j = 0 + while i < n and j < m: + if a[i] == b[j]: + pairs.append((i, j)) + i += 1 + j += 1 + elif table[i + 1][j] >= table[i][j + 1]: + i += 1 + else: + j += 1 + return pairs + + +def diff3(base: list[str], first: list[str], second: list[str]) -> str | None: + match_first = dict(_lcs_pairs(base, first)) + match_second = dict(_lcs_pairs(base, second)) + out: list[str] = [] + b = f = s = 0 + while True: + # Find the next base line aligned in both sides at or after the cursors. + stable = None + for i in range(b, len(base)): + if i in match_first and i in match_second and match_first[i] >= f and match_second[i] >= s: + stable = i + break + if stable is None: + chunk = (base[b:], first[f:], second[s:]) + else: + chunk = (base[b:stable], first[f : match_first[stable]], second[s : match_second[stable]]) + base_chunk, first_chunk, second_chunk = chunk + if first_chunk == second_chunk or second_chunk == base_chunk: + out.extend(first_chunk) + elif first_chunk == base_chunk: + out.extend(second_chunk) + else: + return None + if stable is None: + return "".join(out) + out.append(base[stable]) + b, f, s = stable + 1, match_first[stable] + 1, match_second[stable] + 1 + + +def merge_records( + types: list[TypeDef], + base_text: str, + first_text: str, + second_text: str, + base_path: str, + first_path: str, + second_path: str, +) -> MergeResult: + base, first, second = (parse_document(t) for t in (base_text, first_text, second_text)) + conflicts: list[dict[str, Any]] = [] + eol = _line_ending(first) + + if first_text == second_text or second_text == base_text: + content_shortcut: str | None = first_text + elif first_text == base_text: + content_shortcut = second_text + else: + content_shortcut = None + + # Path merges with the conflict strategy. + if first_path == second_path or second_path == base_path: + path = first_path + elif first_path == base_path: + path = second_path + else: + path = first_path + conflicts.append({"kind": "path", "base": base_path, "first": first_path, "second": second_path}) + + if content_shortcut is not None: + return MergeResult(document=content_shortcut, path=path, conflicts=conflicts) + + mappings = all(isinstance(d.frontmatter, dict) for d in (base, first, second)) + result = parse_document(first_text) + if mappings: + record_types = matched_types(types, first_path, first.frontmatter) + keys: list[str] = [] + for doc in (first, second, base): + for key in doc.frontmatter: + if key not in keys: + keys.append(key) + for key in keys: + b = base.frontmatter.get(key, MISSING) + f = first.frontmatter.get(key, MISSING) + s = second.frontmatter.get(key, MISSING) + if states_equal(f, s) or states_equal(s, b): + continue + if states_equal(f, b): + _take(result, second, key, eol) + continue + kind = strategy(record_types, key) + if kind in ("max", "min"): + if f is MISSING or s is MISSING: + if f is MISSING: + _take(result, second, key, eol) + continue + order = compare(f, s) + if order is not None: + if (kind == "max" and order < 0) or (kind == "min" and order > 0): + _take(result, second, key, eol) + continue + elif kind == "union": + lists = [as_list(v, key) for v in (b, f, s)] + if all(item is not None for item in lists): + merged = union_merge(*lists) + if not merged and (f is MISSING or s is MISSING): + _remove(result, key) + elif not (isinstance(f, list) and values_equal(merged, f)): + _set(result, first, second, key, merged, eol) + continue + conflicts.append( + { + "kind": "field", + "field": key, + "base": None if b is MISSING else b, + "first": None if f is MISSING else f, + "second": None if s is MISSING else s, + } + ) + else: + fm_text = [render_frontmatter(d) for d in (base, first, second)] + if fm_text[1] == fm_text[2] or fm_text[2] == fm_text[0]: + pass + elif fm_text[1] == fm_text[0]: + result = parse_document(render_frontmatter(second) + first.body) + else: + conflicts.append({"kind": "frontmatter"}) + + body = merge_body(base.body, first.body, second.body) + if body is None: + conflicts.append({"kind": "body"}) + body = first.body + result.body = body + return MergeResult(document=render(result), path=path, conflicts=conflicts) + + +def render_frontmatter(doc: Document) -> str: + if not doc.has_frontmatter: + return "" + copy = Document(**{**doc.__dict__, "body": ""}) + return render(copy) + + +def _take(result: Document, side: Document, key: str, eol: str) -> None: + source = side.entry(key) + if source is None: + _remove(result, key) + return + target = result.entry(key) + if target is not None: + target.lines = list(source.lines) + else: + if not result.has_frontmatter: + result.has_frontmatter = True + result.open_delim = f"---{eol}" + result.close_delim = f"---{eol}" + _append_entry(result, Entry(key=key, lines=list(source.lines))) + + +def _append_entry(result: Document, entry: Entry) -> None: + """Insert a new entry after the last existing entry (Chapter 12A rule 4).""" + last = max((i for i, item in enumerate(result.items) if isinstance(item, Entry)), default=-1) + result.items.insert(last + 1, entry) + + +def _remove(result: Document, key: str) -> None: + result.items = [item for item in result.items if not (isinstance(item, Entry) and item.key == key)] + + +def _set(result: Document, first: Document, second: Document, key: str, value: Any, eol: str) -> None: + style = first.entry(key) or second.entry(key) + lines = emit_entry(key, value, style, eol) + target = result.entry(key) + if target is not None: + target.lines = lines + else: + _append_entry(result, Entry(key=key, lines=lines)) + + +# ------------------------------------------------------------ moves + + +def similarity(a: str, b: str) -> float: + set_a = {line.strip() for line in a.splitlines() if line.strip()} + set_b = {line.strip() for line in b.splitlines() if line.strip()} + if not set_a and not set_b: + return 1.0 + return len(set_a & set_b) / len(set_a | set_b) + + +def _identity_hint(content: str, id_field: str | None) -> str | None: + if not id_field: + return None + doc = parse_document(content) + value = doc.frontmatter.get(id_field) if isinstance(doc.frontmatter, dict) else None + return value if isinstance(value, str) and value else None + + +def detect_moves(disappeared: list[dict[str, Any]], appeared: list[dict[str, Any]], id_field: str | None) -> dict[str, Any]: + candidates = [] + for d in disappeared: + for a in appeared: + hint_d = _identity_hint(d["content"], id_field) + hint_a = _identity_hint(a["content"], id_field) + sim = similarity(d["content"], a["content"]) + if hint_d is not None and hint_a is not None: + if hint_d == hint_a: + candidates.append((0, -sim, d["path"], a["path"])) + continue + if d["content"] == a["content"]: + rank = 1 + elif d.get("file_id") is not None and d.get("file_id") == a.get("file_id") and sim >= 0.5: + rank = 2 + elif d["path"].rsplit("/", 1)[-1] == a["path"].rsplit("/", 1)[-1] and sim >= 0.8: + rank = 3 + else: + continue + candidates.append((rank, -sim, d["path"], a["path"])) + candidates.sort() + used_d: set[str] = set() + used_a: set[str] = set() + moves = [] + for _, _, from_path, to_path in candidates: + if from_path in used_d or to_path in used_a: + continue + used_d.add(from_path) + used_a.add(to_path) + moves.append({"from": from_path, "to": to_path}) + return { + "moves": sorted(moves, key=lambda m: m["from"]), + "deleted": sorted(d["path"] for d in disappeared if d["path"] not in used_d), + "created": sorted(a["path"] for a in appeared if a["path"] not in used_a), + } + + +# ---------------------------------------------------------------- regex + + +class InvalidPattern(Exception): + pass + + +_FORBIDDEN = [ + (re.compile(r"\\[pP]"), "Unicode class"), + (re.compile(r"\\[1-9]"), "backreference"), + (re.compile(r"\(\?(=|!|<=| bool: + """Unanchored search with ASCII-only classes and case folding.""" + for forbidden, what in _FORBIDDEN: + if forbidden.search(pattern): + raise InvalidPattern(what) + translated = re.sub(r"\\x\{([0-9A-Fa-f]+)\}", lambda m: chr(int(m.group(1), 16)), pattern) + translated = translated.replace("\\z", "\\Z") + try: + compiled = re.compile(translated, re.ASCII) + except re.error as exc: + raise InvalidPattern(str(exc)) from exc + return compiled.search(text) is not None + + +# ----------------------------------------------------------- body edits + + +class BodyEditError(Exception): + def __init__(self, code: str, reason: str | None = None): + super().__init__(code if reason is None else f"{code}: {reason}") + self.code = code + self.reason = reason + + +def body_digest(body: str) -> str: + import hashlib + + return "sha256:" + hashlib.sha256(body.encode("utf-8")).hexdigest() + + +def apply_edits(base: str, edits: list[dict[str, Any]]) -> str: + """Apply edits whose offsets count Unicode scalar values of `base`.""" + previous_end = 0 + previous_insert_at: int | None = None + out = [] + for edit in edits: + start, end, text = edit["start"], edit["end"], edit.get("text", "") + if not (isinstance(start, int) and isinstance(end, int)) or start < 0 or end < start or end > len(base): + raise BodyEditError("invalid_request", "offset_out_of_range") + if start < previous_end or (start == end and previous_insert_at == start): + raise BodyEditError("invalid_request", "edits_overlap_or_unordered") + out.append(base[previous_end:start]) + out.append(text) + previous_end = end + previous_insert_at = start if start == end else None + out.append(base[previous_end:]) + return "".join(out) + + +def rebase_body_edits(base: str | None, current: str, base_digest: str, edits: list[dict[str, Any]]) -> str: + if body_digest(current) == base_digest: + return apply_edits(current, edits) + if base is None or body_digest(base) != base_digest: + raise BodyEditError("concurrent_modification", "body_base_unavailable") + edited = apply_edits(base, edits) + merged = merge_body(base, current, edited) + if merged is None: + raise BodyEditError("concurrent_modification", "body_conflict") + return merged + + +# ------------------------------------------------------------- fixtures + +FIXTURE_OPERATIONS = { + "merge_records", + "merge_strategies", + "path_equivalence", + "allocate_path", + "derive_path", + "detect_moves", + "regex_match", + "apply_body_edits", +} + + +def setup_types(setup: dict[str, Any]) -> list[TypeDef]: + return [load_type(text) for text in (setup.get("types") or {}).values()] + + +def run_fixture(test: dict[str, Any], setup: dict[str, Any]) -> None: + operation = test["operation"] + data = test.get("input") or {} + expect = test.get("expect") or {} + + if operation == "merge_records": + paths = data.get("paths") or {"base": data["path"], "first": data["path"], "second": data["path"]} + result = merge_records( + setup_types(setup), data["base"], data["first"], data["second"], paths["base"], paths["first"], paths["second"] + ) + if result.document != expect["document"]: + raise AssertionError(f"merged document differs:\n{result.document!r}\nexpected:\n{expect['document']!r}") + conflicts = [{k: v for k, v in c.items() if k in ("kind", "field")} for c in result.conflicts] + if conflicts != expect.get("conflicts", []): + raise AssertionError(f"conflicts {conflicts} != expected {expect.get('conflicts')}") + if "path" in expect and expect["path"] != result.path: + raise AssertionError(f"path {result.path!r} != expected {expect['path']!r}") + return + + if operation == "merge_strategies": + document = parse_document(data["document"]) + types = matched_types(setup_types(setup), data["path"], document.frontmatter) + if "error" in expect: + field_name = expect["error"]["field"] + try: + strategy(types, field_name) + except ValueError as exc: + if expect["error"]["code"] not in str(exc): + raise AssertionError(f"unexpected error {exc}") from exc + return + raise AssertionError("expected an error") + for field_name, wanted in expect["strategies"].items(): + actual = strategy(types, field_name) + if actual != wanted: + raise AssertionError(f"{field_name}: strategy {actual!r} != expected {wanted!r}") + return + + if operation == "path_equivalence": + groups = equivalence_groups(data["paths"]) + if groups != expect["groups"]: + raise AssertionError(f"groups {groups} != expected {expect['groups']}") + return + + if operation == "allocate_path": + path = allocate_path(data["requested"], data.get("existing") or []) + if path != expect["path"]: + raise AssertionError(f"path {path!r} != expected {expect['path']!r}") + return + + if operation == "derive_path": + try: + path = derive_path(data["pattern"], data.get("frontmatter") or {}) + except PathError as exc: + if "error" not in expect or expect["error"].get("code") != exc.code or expect["error"].get("field") != exc.field: + raise AssertionError(f"unexpected {exc.code} for {exc.field}") from exc + return + if expect.get("path") != path: + raise AssertionError(f"path {path!r} != expected {expect.get('path')!r}") + return + + if operation == "detect_moves": + result = detect_moves(data.get("disappeared") or [], data.get("appeared") or [], data.get("id_field")) + for key in ("moves", "deleted", "created"): + if result[key] != expect.get(key, []): + raise AssertionError(f"{key} {result[key]} != expected {expect.get(key)}") + return + + if operation == "regex_match": + try: + matched = regex_match(data["pattern"], data["text"]) + except InvalidPattern as exc: + if expect.get("error", {}).get("code") != "invalid_pattern": + raise AssertionError(f"unexpected invalid pattern: {exc}") from exc + return + if "error" in expect or matched != expect["matches"]: + raise AssertionError(f"matches {matched} != expected {expect}") + return + + if operation == "apply_body_edits": + base = data.get("base") + digest = body_digest(base) if base is not None else data["body_base"] + if not data.get("base_available", True): + base = None + try: + body = rebase_body_edits(base, data["current"], digest, data["edits"]) + except BodyEditError as exc: + want = expect.get("error") or {} + if want.get("code") != exc.code or (want.get("details") or {}).get("reason") != exc.reason: + raise AssertionError(f"unexpected {exc}") from exc + return + if "error" in expect or body != expect["body"]: + raise AssertionError(f"body {body!r} != expected {expect}") + return + + raise AssertionError(f"no executable model for {operation}") diff --git a/tests/v0.3/README.md b/tests/v0.3/README.md index 37d9150..460cc0d 100644 --- a/tests/v0.3/README.md +++ b/tests/v0.3/README.md @@ -4,7 +4,7 @@ This directory is the parallel v0.3 conformance suite. It does not replace the existing `tests/level-*` v0.2.x suite. v0.3 conformance claims use the atomic profiles defined by the specification. -The tests below are grouped into eleven fixture sets for the rollout plan: +The tests below are grouped into thirteen fixture sets: 1. `schema_artifacts` 2. `migration` @@ -17,6 +17,8 @@ The tests below are grouped into eleven fixture sets for the rollout plan: 9. `event_action_interop` 10. `runtime_contracts` 11. `workflow_execution` +12. `merge` (since rc.5) +13. `watch` (since rc.5) The suite covers JSON Schema artifacts, type wrappers, first-class data contracts and projections, collection semantics, CEL host bindings, saved @@ -103,6 +105,11 @@ Future v0.3 adapters should support these operations: - `migrate_type` - `assess_type_pack` - `apply_type_pack` +- `load_types` +- `resolve_link` +- `merge_records`, `merge_strategies`, `path_equivalence`, `allocate_path`, + `derive_path`, `detect_moves`, `regex_match`, and `apply_body_edits` (since + rc.5; see below) ### Type-pack history @@ -139,3 +146,133 @@ Artifact checks cover schemas, examples, and migration output. Core operations, lifecycle behavior, CEL evaluation, runtime dispatch, and workflow execution use adapters or local prototype implementations. Stable-release adapter gates are tracked in [release/v0.3.0.md](../../release/v0.3.0.md). + +## Release-candidate markers + +Tests added after 0.3.0-rc.4 carry `since: 0.3.0-rc.5`. Existing tests whose +expected outcome changed carry `changed: 0.3.0-rc.5`. An engine that conforms +to rc.4 passes every test without either marker; the release notes in +`docs/releases/0.3.0-rc.5.md` list the changed tests. Requirements added in +rc.5 are marked with a `# since 0.3.0-rc.5` comment in `manifest.yaml`. + +## Concurrent-edit, regex, and body-edit operations (since rc.5) + +The merge, path, and move-detection fixtures use the operations below. They +are pure functions of their input and the group's `setup.types`, so +`scripts/check_v03_tests.py` executes them against the executable model in +`scripts/concurrent_edits_model.py`. The rc.5 Core Write tests also use +`update` with `add` and `remove`, and `load_types` to return type-loading +diagnostics. Tests that create two files differing only in case, such as +`paths.discovered_collision`, need a case-sensitive file system; an adapter on +a case-insensitive one reports them as skipped. + +### `merge_records` + +The extension for merge cases: base + two edits + declarations -> expected +result. Declarations come from the group's `setup.types`, as in every other +suite. + +```yaml +operation: merge_records +input: + path: tasks/T.md # the path of all three versions, or: + paths: { base: items/a.md, first: items/a.md, second: items/b.md } + base: | # exact source of the common base version + --- + ... + first: | # the earlier-ordered edited version + second: | # the later-ordered edited version +expect: + document: | # exact bytes of the merged version + conflicts: # every conflict, in frontmatter key order, then + - kind: field # frontmatter, body, path + field: status + path: items/b.md # optional: the merged path +``` + +`first` is the version ordered earlier, for example the one an engine +confirmed first. At a conflict the merged document holds the first version's +value. `expect.conflicts` lists conflicts by `kind` (`field`, `frontmatter`, +`body`, or `path`) and `field`; adapters may return the base, first, and +second values as well, which the expectation does not compare. Because +`expect.document` is exact bytes, these tests also check writer format +fidelity. + +The first group converts the mdbase-next prototype's merge fixtures +(`reference/prototype/crates/mdb-core/tests/merge_fixtures/*.json`): the +prototype's `theirs` (sequenced first) became `first` and its `ours` (the +incoming external edit) became `second`. The prototype's `x-merge: max` schema +annotation became `collection.merge: { completedDate: max }`. + +### `merge_strategies` + +```yaml +operation: merge_strategies +input: + path: items/a.md + document: | # source used for type matching +expect: + strategies: { firstSeen: min, tags: union, title: conflict } + # or + error: { code: type_conflict, field: rank } +``` + +### `path_equivalence` + +`input.paths` is a list of paths; `expect.groups` lists every group of two or +more equivalent paths, each sorted, and the groups sorted. + +### `allocate_path` + +`input.requested` is a path and `input.existing` the paths already in use; +`expect.path` is the path the collision rule chooses. + +### `derive_path` + +`input.pattern` and `input.frontmatter`; `expect.path`, or `expect.error` with +`code` and `field`. + +### `detect_moves` + +```yaml +operation: detect_moves +input: + id_field: id # optional settings.id_field + disappeared: # last observed state of records that disappeared + - { path: notes/a.md, content: "...", file_id: f1 } + appeared: # current state of records that appeared + - { path: archive/a.md, content: "...", file_id: f1 } +expect: + moves: [{ from: notes/a.md, to: archive/a.md }] # sorted by from + deleted: [] # sorted + created: [] # sorted +``` + +`file_id` stands for a platform file identity. Omit it to model a platform +that cannot supply one. All entries are observed within one observation +window. + +### `regex_match` + +`input.pattern` and `input.text`; `expect.matches` is whether the mdbase regex +profile (Chapter 10) finds an unanchored match, or `expect.error.code` is +`invalid_pattern`. Adapters run the pattern through the same engine they use +for CEL `matches()` and JSON Schema `pattern`. + +### `apply_body_edits` + +```yaml +operation: apply_body_edits +input: + base: | # the base body; body_base is its digest + current: | # the record's current body + edits: # offsets count Unicode scalar values of base + - { start: 9, end: 15, text: costs } + base_available: false # optional: the engine has no copy of the base +expect: + body: | # the body written, or: + error: { code: concurrent_modification, details: { reason: body_conflict } } +``` + +It models the body part of an `update` with `body_edits` (Chapter 12). + diff --git a/tests/v0.3/cel/match-determinism.yaml b/tests/v0.3/cel/match-determinism.yaml new file mode 100644 index 0000000..940a5cd --- /dev/null +++ b/tests/v0.3/cel/match-determinism.yaml @@ -0,0 +1,121 @@ +# Deterministic matching (Chapters 07 and 10): a match.expr that depends on the +# clock, file timestamps, or other records loads with a nondeterministic_match +# warning. Queries may use those bindings freely. +name: "deterministic match expressions" +spec_version: "0.3.0" +fixture_set: cel +category: cel +spec_ref: "v0.3/07, v0.3/10" + +groups: + - name: "nondeterministic match expressions are reported" + setup: + config: | + spec_version: "0.3.0" + types: + overdue.md: | + --- + kind: mdbase.type + name: overdue + version: 1 + match: + path_glob: "tasks/**/*.md" + expr: + $expr: 'has(raw.due) && raw.due < today()' + schema: + dialect: json-schema-2020-12 + value: + type: object + --- + recent.md: | + --- + kind: mdbase.type + name: recent + version: 1 + match: + path_glob: "inbox/**/*.md" + expr: + $expr: 'file.mtime > now() - duration("24h")' + schema: + dialect: json-schema-2020-12 + value: + type: object + --- + child.md: | + --- + kind: mdbase.type + name: child + version: 1 + match: + expr: + $expr: 'has(raw.parent) && link(raw.parent).asFile() != null' + schema: + dialect: json-schema-2020-12 + value: + type: object + --- + open_task.md: | + --- + kind: mdbase.type + name: open_task + version: 1 + match: + path_glob: "tasks/**/*.md" + expr: + $expr: 'has(raw.status) && raw.status == "open"' + schema: + dialect: json-schema-2020-12 + value: + type: object + --- + files: + tasks/a.md: | + --- + status: open + due: "2020-01-01" + --- + tests: + - name: "nondeterministic match expressions load with a warning" + id: cel_match.nondeterministic_match_warning + since: 0.3.0-rc.5 + covers: [cel_match.deterministic_match] + operation: load_types + input: {} + expect: + valid: true + diagnostics_contain: + - code: nondeterministic_match + severity: warning + type: overdue + details: { binding: today } + - code: nondeterministic_match + severity: warning + type: recent + details: { binding: file.mtime } + - code: nondeterministic_match + severity: warning + type: child + details: { binding: asFile } + + - name: "both deterministic and reported match expressions still match" + id: cel_match.match_deterministic_ok + since: 0.3.0-rc.5 + covers: [cel_match.deterministic_match] + operation: get_types + input: + path: "tasks/a.md" + expect: + valid: true + types: [open_task, overdue] + + - name: "queries may still use today()" + id: cel_match.query_today_allowed + since: 0.3.0-rc.5 + covers: [cel_match.deterministic_match] + operation: query + input: + where: 'has(raw.due) && raw.due < today()' + expect: + valid: true + results: + - path: "tasks/a.md" diff --git a/tests/v0.3/cel/regex-profile.yaml b/tests/v0.3/cel/regex-profile.yaml new file mode 100644 index 0000000..025efa2 --- /dev/null +++ b/tests/v0.3/cel/regex-profile.yaml @@ -0,0 +1,314 @@ +# The mdbase regex profile (Chapter 10), since 0.3.0-rc.5: RE2 syntax with +# ASCII-only \d, \w, \s, \b and case folding, over Unicode scalar values. The +# regex_match cases are pure and run against scripts/concurrent_edits_model.py; +# each one where a Unicode-aware engine would answer differently is a point +# where an rc.4 engine diverges. +name: regex profile +spec_version: "0.3.0" +fixture_set: cel +category: cel +spec_ref: "v0.3/06, v0.3/07, v0.3/10" +groups: + - name: ASCII-only classes and case folding over Unicode scalar values + tests: + - name: "\\w does not match a non-ASCII letter" + id: regex.word_ascii_only + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\w$" + text: "é" + expect: + matches: false + - name: "\\w matches an ASCII letter" + id: regex.word_ascii_letter + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\w$" + text: a + expect: + matches: true + - name: "\\W matches a non-ASCII letter" + id: regex.not_word_non_ascii + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\W$" + text: "é" + expect: + matches: true + - name: "\\d does not match an Arabic-Indic digit" + id: regex.digit_ascii_only + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\d$" + text: "١" + expect: + matches: false + - name: "\\d matches an ASCII digit" + id: regex.digit_ascii + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\d$" + text: "7" + expect: + matches: true + - name: "\\s does not match a no-break space" + id: regex.space_ascii_only + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\s$" + text: " " + expect: + matches: false + - name: "\\s matches a tab" + id: regex.space_tab + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^\\s$" + text: "\t" + expect: + matches: true + - name: "\\b sees a boundary between f and é" + id: regex.boundary_ascii + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "caf\\b" + text: "café" + expect: + matches: true + - name: "(?i) does not fold non-ASCII letters" + id: regex.ignore_case_ascii_only + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "(?i)^é$" + text: "É" + expect: + matches: false + - name: "(?i) folds ASCII letters" + id: regex.ignore_case_ascii + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "(?i)^abc$" + text: ABC + expect: + matches: true + - name: ". matches one Unicode scalar value" + id: regex.dot_scalar_value + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^.$" + text: "é" + expect: + matches: true + - name: ". matches a scalar value outside the BMP" + id: regex.dot_astral + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^.$" + text: "😀" + expect: + matches: true + - name: a bracket range can list non-ASCII letters + id: regex.explicit_range + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^[a-zà-ÿ]+$" + text: "café" + expect: + matches: true + - name: "\\x{...} names a scalar value" + id: regex.hex_escape + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "^caf\\x{E9}$" + text: "café" + expect: + matches: true + - name: a match is unanchored + id: regex.unanchored + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: b + text: abc + expect: + matches: true + - name: "\\p{L} is invalid" + id: regex.unicode_class_invalid + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "\\p{L}" + text: a + expect: + error: + code: invalid_pattern + - name: a backreference is invalid + id: regex.backreference_invalid + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "(a)\\1" + text: aa + expect: + error: + code: invalid_pattern + - name: look-ahead is invalid + id: regex.lookahead_invalid + since: "0.3.0-rc.5" + covers: [cel.regex_profile, core_read.regex_profile] + operation: regex_match + input: + pattern: "a(?=b)" + text: ab + expect: + error: + code: invalid_pattern + - name: "regex profile in CEL, JSON Schema, and match.where" + setup: + config: | + spec_version: "0.3.0" + types: + word.md: | + --- + kind: mdbase.type + name: word + version: 1 + match: + path_glob: "words/**/*.md" + where: + label: + matches: "^\\w+$" + schema: + dialect: json-schema-2020-12 + value: + type: object + properties: + code: { type: string, pattern: "^\\w+$" } + --- + files: + words/plain.md: | + --- + label: cafe + code: abc + --- + words/accented.md: | + --- + label: café + code: café + --- + tests: + - name: "CEL matches() uses ASCII-only \\w" + id: regex.cel_matches_ascii + since: "0.3.0-rc.5" + covers: [cel.regex_profile] + operation: evaluate_cel + input: + path: words/plain.md + expression: "\"é\".matches(\"^\\\\w$\") || !\"a\".matches(\"^\\\\w$\")" + expect: + valid: true + value: false + - name: an invalid literal pattern is a compile error + id: regex.cel_literal_invalid + since: "0.3.0-rc.5" + covers: [cel.regex_profile] + operation: evaluate_cel + input: + path: words/plain.md + expression: "\"a\".matches(\"\\\\p{L}\")" + expect: + valid: false + diagnostics: + - code: expression_compile_error + - name: "JSON Schema pattern uses ASCII-only \\w" + id: regex.schema_pattern_ascii + since: "0.3.0-rc.5" + covers: [core_read.regex_profile] + operation: validate + input: + path: words/accented.md + expect: + valid: false + issues: + - code: schema_pattern + field: code + - name: "match.where matches uses ASCII-only \\w" + id: regex.match_where_ascii + since: "0.3.0-rc.5" + covers: [core_read.regex_profile] + operation: get_types + input: + path: words/accented.md + expect: + valid: true + types: [] + - name: match.where matches an ASCII label + id: regex.match_where_plain + since: "0.3.0-rc.5" + covers: [core_read.regex_profile] + operation: get_types + input: + path: words/plain.md + expect: + valid: true + types: [word] + - name: unsupported pattern syntax + setup: + config: | + spec_version: "0.3.0" + types: + letters.md: | + --- + kind: mdbase.type + name: letters + version: 1 + schema: + dialect: json-schema-2020-12 + value: + type: object + properties: + name: { type: string, pattern: "^\\p{L}+$" } + --- + tests: + - name: a Unicode class in a schema pattern makes the type invalid + id: regex.schema_unicode_class_invalid + since: "0.3.0-rc.5" + covers: [core_read.regex_profile] + operation: load_types + input: {} + expect: + valid: false + diagnostics_contain: + - code: invalid_pattern + type: letters diff --git a/tests/v0.3/core/body-edits-update.yaml b/tests/v0.3/core/body-edits-update.yaml new file mode 100644 index 0000000..cf08fb3 --- /dev/null +++ b/tests/v0.3/core/body-edits-update.yaml @@ -0,0 +1,156 @@ +# Update with body_edits (Chapter 12), since 0.3.0-rc.5. +name: body edits through update +spec_version: "0.3.0" +fixture_set: core_collection +category: operations +spec_ref: v0.3/12 +groups: + - name: body_edits requests + setup: + config: | + spec_version: "0.3.0" + settings: + record_extensions: [md, base] + files: + notes/meeting.md: | + --- + title: Meeting + --- + Agenda + - budget + - hiring + views/open.base: | + views: + - type: table + name: Open + tests: + - name: body_edits apply directly through update + id: core_write.body_edits_direct + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: update + input: + path: notes/meeting.md + patch: + title: Weekly meeting + body_base: "sha256:1f72be3e4919a0546aef495f8cd10a6316b585386fec3bbf2096007d42ee57d7" + body_edits: + - start: 9 + end: 15 + text: costs + include_document: true + expect: + valid: true + document: | + --- + title: Weekly meeting + --- + Agenda + - costs + - hiring + - name: a stale base without base text fails without writing + id: core_write.body_edits_stale_base + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: update + input: + path: notes/meeting.md + body_base: "sha256:5d1af5fb9d7c67c2b30ab3023be2dec7d0b4361387924cc9f86d1978dc5153d4" + body_edits: + - start: 0 + end: 6 + text: Plan + expect: + valid: false + error: + code: concurrent_modification + details: + reason: body_base_unavailable + - name: body_base_text lets a stale base rebase + id: core_write.body_edits_base_text + since: "0.3.0-rc.5" + covers: [core_write.body_edits, merge.body_edit_rebase] + operation: update + input: + path: notes/meeting.md + body_base: "sha256:5d1af5fb9d7c67c2b30ab3023be2dec7d0b4361387924cc9f86d1978dc5153d4" + body_base_text: | + Agenda + - budget + body_edits: + - start: 0 + end: 6 + text: Plan + expect: + valid: true + body: | + Plan + - budget + - hiring + - name: body_base_text must match body_base + id: core_write.body_edits_base_text_mismatch + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: update + input: + path: notes/meeting.md + body_base: "sha256:1f72be3e4919a0546aef495f8cd10a6316b585386fec3bbf2096007d42ee57d7" + body_base_text: | + something else + body_edits: + - start: 0 + end: 6 + text: Plan + expect: + valid: false + error: + code: invalid_request + - name: body_edits cannot be combined with body + id: core_write.body_edits_with_body + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: update + input: + path: notes/meeting.md + body: | + x + body_base: "sha256:1f72be3e4919a0546aef495f8cd10a6316b585386fec3bbf2096007d42ee57d7" + body_edits: + - start: 0 + end: 0 + text: y + expect: + valid: false + error: + code: invalid_request + - name: body_edits require body_base + id: core_write.body_edits_without_base + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: update + input: + path: notes/meeting.md + body_edits: + - start: 0 + end: 0 + text: y + expect: + valid: false + error: + code: invalid_request + - name: a YAML document record has no body to edit + id: core_write.body_edits_yaml_document + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: update + input: + path: views/open.base + body_base: "sha256:e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" + body_edits: + - start: 0 + end: 0 + text: y + expect: + valid: false + error: + code: invalid_request diff --git a/tests/v0.3/core/core-collection.yaml b/tests/v0.3/core/core-collection.yaml index 0e3e775..0cff155 100644 --- a/tests/v0.3/core/core-collection.yaml +++ b/tests/v0.3/core/core-collection.yaml @@ -598,7 +598,7 @@ groups: resolved_links: parent: "tasks/valid.md" - - name: "collection links enforce validate_exists" + - name: "collection links report validate_exists" operation: validate input: path: "tasks/broken-parent.md" diff --git a/tests/v0.3/core/link-ambiguity.yaml b/tests/v0.3/core/link-ambiguity.yaml new file mode 100644 index 0000000..cf35802 --- /dev/null +++ b/tests/v0.3/core/link-ambiguity.yaml @@ -0,0 +1,84 @@ +# Link clarifications in 0.3.0-rc.5 (Chapters 04 and 08): settings.id_field has +# no default, and an ambiguous ID resolves to null with ambiguous_link and no +# filename fallback. +name: "link resolution clarifications" +spec_version: "0.3.0" +fixture_set: core_collection +category: links +spec_ref: "v0.3/04, v0.3/08" + +groups: + - name: "no id_field configured" + setup: + config: | + spec_version: "0.3.0" + files: + people/alice.md: | + --- + title: Alice by filename + --- + people/someone.md: | + --- + id: alice + title: Record whose id is alice + --- + tasks/t.md: | + --- + owner: "[[alice]]" + --- + tests: + - name: "an id field is not used for resolution unless configured" + id: links.no_default_id_field + since: 0.3.0-rc.5 + covers: [links.id_field_no_default] + operation: resolve_link + input: + path: "tasks/t.md" + field: owner + expect: + valid: true + resolved: "people/alice.md" + + - name: "duplicate configured ids" + setup: + config: | + spec_version: "0.3.0" + settings: + id_field: id + files: + tasks/a.md: | + --- + id: shared + --- + tasks/b.md: | + --- + id: shared + --- + tasks/shared.md: | + --- + title: Filename match + --- + tasks/child.md: | + --- + parent: "[[shared]]" + --- + tests: + - name: "the ambiguity is a cross-record warning on the referring record" + id: links.ambiguous_id_validate + since: 0.3.0-rc.5 + covers: [links.ambiguous_id_link] + operation: validate + input: + path: "tasks/child.md" + setup: + config: | + spec_version: "0.3.0" + settings: + id_field: id + validation: warn + expect: + valid: true + issues: + - code: ambiguous_link + severity: warning + field: parent diff --git a/tests/v0.3/core/links-and-discovery.yaml b/tests/v0.3/core/links-and-discovery.yaml index d0be1b7..2aa07bf 100644 --- a/tests/v0.3/core/links-and-discovery.yaml +++ b/tests/v0.3/core/links-and-discovery.yaml @@ -213,7 +213,8 @@ groups: - name: "ambiguous id resolution fails without filename fallback" id: links.duplicate_id_ambiguous - covers: [links.configured_id_resolution] + changed: 0.3.0-rc.5 + covers: [links.configured_id_resolution, links.ambiguous_id_link] operation: resolve_link input: path: "tasks/child.md" @@ -221,3 +222,7 @@ groups: expect: valid: true resolved: null + diagnostics: + - code: ambiguous_link + field: parent + details: { candidates: ["tasks/a.md", "tasks/b.md"] } diff --git a/tests/v0.3/core/list-ops-and-fidelity.yaml b/tests/v0.3/core/list-ops-and-fidelity.yaml new file mode 100644 index 0000000..f30cd9c --- /dev/null +++ b/tests/v0.3/core/list-ops-and-fidelity.yaml @@ -0,0 +1,374 @@ +# Core Write additions in 0.3.0-rc.5: list operations (Chapter 12), writer format +# fidelity (Chapter 12A), and path collisions on create and rename (Chapters +# 02 and 12). expect.document is the exact post-write source. +name: "core write additions" +spec_version: "0.3.0" +fixture_set: core_collection +category: operations +spec_ref: "v0.3/02, v0.3/03, v0.3/12, v0.3/12A" + +groups: + - name: "list operations and format fidelity" + setup: + config: | + spec_version: "0.3.0" + types: + note.md: | + --- + kind: mdbase.type + name: note + version: 1 + match: + path_glob: "notes/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + --- + files: + notes/a.md: | + --- + title: 'A' # quoted on purpose + status: open + tags: [x, y] # flow style + reviewers: + - ann + --- + Body of A. + notes/single-tag.md: | + --- + tags: solo + --- + notes/scalar-list.md: | + --- + labels: not-a-list + --- + tests: + - name: "add appends missing items and keeps the flow style and comments" + id: core_write.add_flow_style + since: 0.3.0-rc.5 + covers: [core_write.list_operations, core_write.writer_format_fidelity] + operation: update + input: + path: "notes/a.md" + add: { tags: [y, z] } + include_document: true + expect: + valid: true + document: | + --- + title: 'A' # quoted on purpose + status: open + tags: [x, y, z] # flow style + reviewers: + - ann + --- + Body of A. + + - name: "add to a block list keeps the block style" + id: core_write.add_block_style + since: 0.3.0-rc.5 + covers: [core_write.list_operations, core_write.writer_format_fidelity] + operation: update + input: + path: "notes/a.md" + add: { reviewers: [bo] } + include_document: true + expect: + valid: true + document: | + --- + title: 'A' # quoted on purpose + status: open + tags: [x, y] # flow style + reviewers: + - ann + - bo + --- + Body of A. + + - name: "removing every item leaves an empty list" + id: core_write.remove_to_empty + since: 0.3.0-rc.5 + covers: [core_write.list_operations] + operation: update + input: + path: "notes/a.md" + remove: { tags: [x, y, absent] } + expect: + valid: true + frontmatter: { tags: [] } + + - name: "add creates a missing field" + id: core_write.add_creates_field + since: 0.3.0-rc.5 + covers: [core_write.list_operations] + operation: update + input: + path: "notes/a.md" + add: { labels: [new] } + expect: + valid: true + frontmatter: { labels: [new] } + + - name: "a tags string is read as a one-item list" + id: core_write.tags_string + since: 0.3.0-rc.5 + covers: [core_write.list_operations] + operation: update + input: + path: "notes/single-tag.md" + add: { tags: [more] } + expect: + valid: true + frontmatter: { tags: [solo, more] } + + - name: "a non-list value cannot take list operations" + id: core_write.non_list_rejected + since: 0.3.0-rc.5 + covers: [core_write.list_operations] + operation: update + input: + path: "notes/scalar-list.md" + add: { labels: [x] } + expect: + valid: false + error: + code: invalid_request + + - name: "naming one field in patch and add is invalid" + id: core_write.patch_and_add_conflict + since: 0.3.0-rc.5 + covers: [core_write.list_operations] + operation: update + input: + path: "notes/a.md" + patch: { tags: [q] } + add: { tags: [z] } + expect: + valid: false + error: + code: invalid_request + + - name: "adding and removing the same item is invalid" + id: core_write.add_remove_same_item + since: 0.3.0-rc.5 + covers: [core_write.list_operations] + operation: update + input: + path: "notes/a.md" + add: { tags: [z] } + remove: { tags: [z] } + expect: + valid: false + error: + code: invalid_request + + - name: "a patch re-emits only the changed entry" + id: core_write.patch_minimal_bytes + since: 0.3.0-rc.5 + covers: [core_write.writer_format_fidelity] + operation: update + input: + path: "notes/a.md" + patch: { status: done } + include_document: true + expect: + valid: true + document: | + --- + title: 'A' # quoted on purpose + status: done + tags: [x, y] # flow style + reviewers: + - ann + --- + Body of A. + + - name: "a new key is appended after the last entry" + id: core_write.new_key_appended + since: 0.3.0-rc.5 + covers: [core_write.writer_format_fidelity] + operation: update + input: + path: "notes/a.md" + patch: { owner: ann } + unset: [status] + include_document: true + expect: + valid: true + document: | + --- + title: 'A' # quoted on purpose + tags: [x, y] # flow style + reviewers: + - ann + owner: ann + --- + Body of A. + + - name: "format fidelity for YAML document records" + setup: + config: | + spec_version: "0.3.0" + settings: + record_extensions: [md, base] + files: + views/tasks.base: | + # Owned by Obsidian + filters: + and: + - 'status == "open"' # keep this filter first + views: + - type: table + name: Open tasks + limit: 10 + tests: + - name: "a structured update keeps comments and untouched entries" + id: core_write.yaml_document_fidelity + since: 0.3.0-rc.5 + covers: [core_write.writer_format_fidelity] + operation: update + input: + path: "views/tasks.base" + patch: { limit: 25 } + include_document: true + expect: + valid: true + document: | + # Owned by Obsidian + filters: + and: + - 'status == "open"' # keep this filter first + views: + - type: table + name: Open tasks + limit: 25 + + - name: "path collisions" + setup: + config: | + spec_version: "0.3.0" + types: + task.md: | + --- + kind: mdbase.type + name: task + version: 1 + match: + path_glob: "tasks/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + required: [title] + properties: + title: { type: string } + collection: + path: + pattern: "tasks/{title}.md" + --- + files: + tasks/Call Bob.md: | + --- + title: Call Bob + --- + tasks/todo.md: | + --- + title: todo + --- + tests: + - name: "an explicit path equivalent to an existing record fails" + id: paths.explicit_create_conflict + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: create + input: + path: "tasks/CALL BOB.md" + type: task + frontmatter: { title: Another } + expect: + valid: false + error: + code: path_conflict + + - name: "a derived path that collides gets the first free suffix" + id: paths.derived_create_suffix + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: create + input: + type: task + frontmatter: { title: call bob } + expect: + valid: true + path: "tasks/call bob (2).md" + frontmatter: { title: call bob } + + - name: "a derived path value containing a slash is rejected" + id: paths.derived_slash_rejected + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: create + input: + type: task + frontmatter: { title: "Q3/Q4 plan" } + expect: + valid: false + error: + code: path_value_invalid + field: title + + - name: "a case-only rename of the same record is allowed" + id: paths.case_only_rename + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: rename + input: + from: "tasks/todo.md" + to: "tasks/Todo.md" + expect: + valid: true + path: "tasks/Todo.md" + + - name: "a rename onto another record's equivalent path fails" + id: paths.rename_conflict + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: rename + input: + from: "tasks/todo.md" + to: "tasks/call bob.md" + expect: + valid: false + error: + code: path_conflict + + - name: "discovered path collisions" + setup: + config: | + spec_version: "0.3.0" + files: + notes/Plan.md: | + --- + title: upper + --- + notes/plan.md: | + --- + title: lower + --- + tests: + - name: "both files are records and both report path_collision" + id: paths.discovered_collision + since: 0.3.0-rc.5 + covers: [core_read.path_equivalence] + operation: read + input: + path: "notes/plan.md" + expect: + valid: true + frontmatter: { title: lower } + diagnostics: + - code: path_collision + severity: warning + details: { paths: ["notes/Plan.md", "notes/plan.md"] } diff --git a/tests/v0.3/core/paths.yaml b/tests/v0.3/core/paths.yaml new file mode 100644 index 0000000..2a2eff1 --- /dev/null +++ b/tests/v0.3/core/paths.yaml @@ -0,0 +1,215 @@ +# Executable cases for path keys, the collision suffix rule, and path-pattern +# values (Chapters 02 and 07). Adapter-level create and rename cases are in +# core-write.yaml. +name: "path equivalence, collisions, and path-pattern values" +spec_version: "0.3.0" +fixture_set: core_collection +category: paths +spec_ref: "v0.3/02, v0.3/07" +groups: + - name: path keys and derived paths + tests: + - name: paths that differ in case or Unicode normalization share a path key + id: paths.case_and_nfc + since: 0.3.0-rc.5 + covers: [core_read.path_equivalence] + operation: path_equivalence + input: + paths: ["Notes/Café.md", "notes/café.md", "NOTES/CAFÉ.md", notes/other.md] + expect: + groups: + - ["NOTES/CAFÉ.md", "Notes/Café.md", "notes/café.md"] + - name: "full case folding makes Straße and STRASSE equivalent" + id: paths.full_case_folding + since: 0.3.0-rc.5 + covers: [core_read.path_equivalence] + operation: path_equivalence + input: + paths: ["Straße.md", STRASSE.md, strasse-2.md] + expect: + groups: + - [STRASSE.md, "Straße.md"] + - name: different names are not equivalent + id: paths.distinct_paths + since: 0.3.0-rc.5 + covers: [core_read.path_equivalence] + operation: path_equivalence + input: + paths: [tasks/a.md, tasks/b.md, tasks/a.base] + expect: + groups: [] + - name: a free path is used as requested + id: paths.free_path_unchanged + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: allocate_path + input: + requested: tasks/Call Bob.md + existing: [tasks/Other.md] + expect: + path: tasks/Call Bob.md + - name: a taken path gets the first free suffix + id: paths.first_suffix + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: allocate_path + input: + requested: tasks/Call Bob.md + existing: [tasks/Call Bob.md] + expect: + path: "tasks/Call Bob (2).md" + - name: suffix candidates are checked by path key + id: paths.suffix_uses_path_keys + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: allocate_path + input: + requested: tasks/call bob.md + existing: [tasks/Call Bob.md, "tasks/CALL BOB (2).md"] + expect: + path: "tasks/call bob (3).md" + - name: an existing suffix is not parsed + id: paths.existing_suffix_not_parsed + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: allocate_path + input: + requested: "tasks/X (2).md" + existing: ["tasks/X (2).md"] + expect: + path: "tasks/X (2) (2).md" + - name: the suffix goes before the final extension only + id: paths.suffix_before_final_extension + since: 0.3.0-rc.5 + covers: [core_write.path_collisions] + operation: allocate_path + input: + requested: views/tasks.view.base + existing: [views/tasks.view.base] + expect: + path: "views/tasks.view (2).base" + - name: a string value fills its placeholder + id: paths.derive_plain + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: Call Bob + expect: + path: tasks/Call Bob.md + - name: a number uses its JSON representation + id: paths.derive_number + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: 42 + expect: + path: tasks/42.md + - name: a value containing a slash cannot create a folder + id: paths.derive_slash + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: Q3/Q4 plan + expect: + error: + code: path_value_invalid + field: title + - name: a value containing a backslash is invalid + id: paths.derive_backslash + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: "a\\b" + expect: + error: + code: path_value_invalid + field: title + - name: a value beginning with a dot is invalid + id: paths.derive_dot + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: ".." + expect: + error: + code: path_value_invalid + field: title + - name: a value that would hide the file is invalid + id: paths.derive_hidden + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: ".draft" + expect: + error: + code: path_value_invalid + field: title + - name: an empty value is invalid + id: paths.derive_empty + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: "" + expect: + error: + code: path_value_invalid + field: title + - name: a list value is invalid + id: paths.derive_list + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: [a] + expect: + error: + code: path_value_invalid + field: title + - name: a missing value is path_value_missing + id: paths.derive_missing + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: {} + expect: + error: + code: path_value_missing + field: title + - name: a null value is path_value_missing + id: paths.derive_null + since: 0.3.0-rc.5 + covers: [core_write.path_pattern_values] + operation: derive_path + input: + pattern: "tasks/{title}.md" + frontmatter: + title: null + expect: + error: + code: path_value_missing + field: title diff --git a/tests/v0.3/core/validation-tiers.yaml b/tests/v0.3/core/validation-tiers.yaml new file mode 100644 index 0000000..d3ec64c --- /dev/null +++ b/tests/v0.3/core/validation-tiers.yaml @@ -0,0 +1,449 @@ +# Validation tiers (Chapter 04), uniqueness enforcement and scope (Chapter 07), +# and non-blocking link checks (Chapter 08). These tests need a collection +# engine and run through adapters. +name: "validation tiers and cross-record checks" +spec_version: "0.3.0" +fixture_set: core_collection +category: validation +spec_ref: "v0.3/04, v0.3/07, v0.3/08" + +groups: + - name: "uniqueness reports by default and enforces on opt-in" + setup: + config: | + spec_version: "0.3.0" + settings: + validation: error + types: + task.md: | + --- + kind: mdbase.type + name: task + version: 1 + match: + path_glob: "tasks/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + required: [title] + properties: + title: { type: string } + ref: {} + code: { type: string } + note: { type: string } + collection: + unique: + - field: ref + - field: code + enforce: write + --- + note.md: | + --- + kind: mdbase.type + name: note + version: 1 + match: + path_glob: "notes/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + --- + files: + tasks/a.md: | + --- + title: A + ref: R-1 + code: C-1 + --- + tasks/numeric.md: | + --- + title: Numeric + ref: 7 + --- + tasks/dup-1.md: | + --- + title: Dup one + code: C-9 + --- + tasks/dup-2.md: | + --- + title: Dup two + code: C-9 + --- + notes/n.md: | + --- + code: C-2 + --- + tests: + - name: "a report-mode duplicate is written and reported as a warning" + id: collection.unique_report_default_writes + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_enforcement_modes, core_write.validation_tiers] + operation: create + input: + path: "tasks/b.md" + type: task + frontmatter: { title: B, ref: R-1 } + expect: + valid: true + diagnostics: + - code: duplicate_value + severity: warning + field: ref + details: { paths: ["tasks/a.md"] } + + - name: "validate reports an existing duplicate with the level's severity" + id: collection.unique_report_validate + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_enforcement_modes] + operation: validate + input: + path: "tasks/dup-1.md" + expect: + valid: false + issues: + - code: duplicate_value + severity: error + field: code + details: { paths: ["tasks/dup-2.md"] } + + - name: "an enforce: write duplicate is rejected before writing" + id: collection.unique_enforce_write_rejects + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_enforcement_modes, core_write.validation_tiers] + operation: create + input: + path: "tasks/c.md" + type: task + frontmatter: { title: C, code: C-1 } + expect: + valid: false + error: + code: duplicate_value + + - name: "enforce: write applies at validation level off" + id: collection.unique_enforce_write_level_off + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_enforcement_modes] + operation: create + setup: + config: | + spec_version: "0.3.0" + settings: + validation: "off" + input: + path: "tasks/c.md" + type: task + frontmatter: { title: C, code: C-1 } + expect: + valid: false + error: + code: duplicate_value + + - name: "a write that leaves an existing duplicate unchanged succeeds" + id: collection.unique_enforce_write_unrelated_edit + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_enforcement_modes] + operation: update + input: + path: "tasks/dup-1.md" + patch: { note: still duplicated } + expect: + valid: true + frontmatter: { code: C-9, note: still duplicated } + + - name: "type scope ignores records of other types" + id: collection.unique_type_scope + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_scope] + operation: create + input: + path: "tasks/d.md" + type: task + frontmatter: { title: D, code: C-2 } + expect: + valid: true + diagnostics: [] + + - name: "values are compared without coercion" + id: collection.unique_no_coercion + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_scope] + operation: create + input: + path: "tasks/e.md" + type: task + frontmatter: { title: E, ref: "7" } + expect: + valid: true + diagnostics: [] + + - name: "collection and path_glob scopes" + setup: + config: | + spec_version: "0.3.0" + types: + task.md: | + --- + kind: mdbase.type + name: task + version: 1 + match: + path_glob: "tasks/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + collection: + read_defaults: + slug: default-slug + unique: + - field: code + scope: collection + - field: slug + scope: path_glob + path_glob: "tasks/active/**" + --- + files: + tasks/active/a.md: | + --- + code: K-1 + slug: alpha + --- + tasks/active/b.md: | + --- + slug: alpha + --- + tasks/archive/c.md: | + --- + slug: beta + --- + tasks/active/d.md: | + --- + slug: beta + --- + tasks/active/no-slug-1.md: | + --- + code: K-3 + --- + tasks/active/no-slug-2.md: | + --- + code: K-4 + --- + notes/n.md: | + --- + code: K-1 + --- + tests: + - name: "collection scope compares governed records with records of any type" + id: collection.unique_collection_scope + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_scope] + operation: validate + input: + path: "tasks/active/a.md" + expect: + valid: false + issues: + - code: duplicate_value + field: code + details: { paths: ["notes/n.md"] } + + - name: "a record outside the declaring type is not governed" + id: collection.unique_ungoverned_record + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_scope] + operation: validate + input: + path: "notes/n.md" + expect: + valid: true + diagnostics: [] + + - name: "path_glob scope governs and compares only matching paths" + id: collection.unique_path_glob_scope + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_scope] + operation: validate + input: + path: "tasks/active/d.md" + expect: + valid: true + diagnostics: [] + + - name: "path_glob scope reports duplicates inside the glob" + id: collection.unique_path_glob_duplicate + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_scope] + operation: validate + input: + path: "tasks/active/b.md" + expect: + valid: false + issues: + - code: duplicate_value + field: slug + details: { paths: ["tasks/active/a.md"] } + + - name: "read defaults never take part in uniqueness" + id: collection.unique_raw_values_only + since: 0.3.0-rc.5 + covers: [collection_semantics.unique_scope] + operation: validate + input: + path: "tasks/active/no-slug-1.md" + expect: + valid: true + diagnostics: [] + + - name: "cross-record link checks never block writes" + setup: + config: | + spec_version: "0.3.0" + settings: + validation: error + types: + task.md: | + --- + kind: mdbase.type + name: task + version: 1 + match: + path_glob: "tasks/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + properties: + title: { type: string } + parent: { type: string } + owner: { type: string } + collection: + links: + parent: + target_type: task + validate_exists: true + owner: + target_type: person + --- + person.md: | + --- + kind: mdbase.type + name: person + version: 1 + match: + path_glob: "people/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + --- + files: + tasks/a.md: | + --- + title: A + --- + tasks/b.md: | + --- + title: B + --- + tests: + - name: "a broken link is written at level error and reported as a warning" + id: links.validate_exists_non_blocking + since: 0.3.0-rc.5 + covers: [core_write.validation_tiers, collection_semantics.cross_record_reporting] + operation: update + input: + path: "tasks/a.md" + patch: { parent: "[[missing]]" } + expect: + valid: true + frontmatter: { parent: "[[missing]]" } + diagnostics: + - code: link_not_found + severity: warning + field: parent + + - name: "a link to a record of the wrong type is reported, not rejected" + id: links.target_type_non_blocking + since: 0.3.0-rc.5 + covers: [collection_semantics.cross_record_reporting] + operation: update + input: + path: "tasks/a.md" + patch: { owner: "[[b]]" } + expect: + valid: true + diagnostics: + - code: link_target_type_mismatch + severity: warning + field: owner + + - name: "a single-record issue still rejects at level error" + id: core_write.single_record_blocks + since: 0.3.0-rc.5 + covers: [core_write.validation_tiers] + operation: update + input: + path: "tasks/a.md" + patch: { title: 5 } + expect: + valid: false + diagnostics: + - code: schema_type + severity: error + + - name: "invalid records are read and reported" + setup: + config: | + spec_version: "0.3.0" + settings: + validation: error + types: + task.md: | + --- + kind: mdbase.type + name: task + version: 1 + match: + path_glob: "tasks/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + properties: + priority: { type: integer, maximum: 5 } + --- + files: + tasks/edited-elsewhere.md: | + --- + priority: 9 + --- + Written by another tool. + tests: + - name: "a query returns an invalid record" + id: core_write.invalid_record_queried + since: 0.3.0-rc.5 + covers: [core_read.invalid_records_reported] + operation: query + input: + types: [task] + expect: + valid: true + results: + - path: "tasks/edited-elsewhere.md" + + - name: "an update that repairs an externally invalidated record succeeds" + id: core_write.update_invalid_record + since: 0.3.0-rc.5 + covers: [core_read.invalid_records_reported, core_write.validation_tiers] + operation: update + input: + path: "tasks/edited-elsewhere.md" + patch: { priority: 3 } + expect: + valid: true + frontmatter: { priority: 3 } + diagnostics: [] diff --git a/tests/v0.3/manifest.yaml b/tests/v0.3/manifest.yaml index af3cfdf..d35d748 100644 --- a/tests/v0.3/manifest.yaml +++ b/tests/v0.3/manifest.yaml @@ -31,6 +31,9 @@ claim_profiles: - portable_globs - raw_schema_validation - machine_readable_diagnostics + - invalid_records_reported # since 0.3.0-rc.5 + - path_equivalence # since 0.3.0-rc.5 + - regex_profile # since 0.3.0-rc.5 - id: collection_semantics status: coverage_complete requires: [core_read] @@ -40,6 +43,10 @@ claim_profiles: - simple_path_policy - display_metadata - independent_multi_type_validation + - unique_enforcement_modes # since 0.3.0-rc.5 + - unique_scope # since 0.3.0-rc.5 + - cross_record_reporting # since 0.3.0-rc.5 + - merge_declarations # since 0.3.0-rc.5 - id: data_contracts status: coverage_complete requires: [core_read] @@ -64,6 +71,7 @@ claim_profiles: - text_helpers - operational_limits - expression_diagnostics + - regex_profile # since 0.3.0-rc.5 - id: cel_match status: draft requires: [core_read, cel] @@ -73,6 +81,7 @@ claim_profiles: - combined_match_semantics - boolean_match_result - match_evaluation_diagnostics + - deterministic_match # since 0.3.0-rc.5 - id: cel_query status: draft requires: [collection_semantics, cel] @@ -92,6 +101,8 @@ claim_profiles: - configured_id_resolution - frontmatter_wikilinks - backlinks + - id_field_no_default # since 0.3.0-rc.5 + - ambiguous_id_link # since 0.3.0-rc.5 - id: core_write status: draft requires: [collection_semantics] @@ -99,6 +110,28 @@ claim_profiles: - null_and_unset - batch_preflight - validation_levels + - validation_tiers # since 0.3.0-rc.5 + - list_operations # since 0.3.0-rc.5 + - writer_format_fidelity # since 0.3.0-rc.5 + - path_collisions # since 0.3.0-rc.5 + - path_pattern_values # since 0.3.0-rc.5 + - body_edits # since 0.3.0-rc.5 + - id: merge # since 0.3.0-rc.5 + status: coverage_complete + requires: [core_write] + requirements: + - field_three_way + - default_strategies + - declared_strategies + - max_min_semantics + - union_semantics + - body_three_way + - append_append_union + - conflict_reporting + - invalid_merge_reported + - merge_format_fidelity + - path_merge + - body_edit_rebase # since 0.3.0-rc.5 - id: type_packs status: draft requires: [data_contracts, core_write] @@ -165,6 +198,7 @@ claim_profiles: - per_subject_ordering - notification_coalescing - listener_isolation + - move_detection # since 0.3.0-rc.5 fixture_sets: - id: schema_artifacts @@ -187,6 +221,11 @@ fixture_sets: - core/core-write.yaml - core/links-and-discovery.yaml - core/yaml-document-records.yaml + - core/validation-tiers.yaml + - core/list-ops-and-fidelity.yaml + - core/paths.yaml + - core/link-ambiguity.yaml + - core/body-edits-update.yaml - id: data_contracts description: Data contracts, type implementations, projections, stable digests, and conflicts. coverage_targets: [data_contracts] @@ -207,6 +246,8 @@ fixture_sets: coverage_targets: [cel, cel_match, cel_query, links] files: - cel/cel-profile.yaml + - cel/match-determinism.yaml + - cel/regex-profile.yaml - id: views description: Ordinary view records resolve named queries, invocation context, and advisory presentation. coverage_targets: [cel_query] @@ -227,3 +268,14 @@ fixture_sets: coverage_targets: [runtime/0.2] files: - workflow-execution/workflow-execution.yaml + - id: merge + description: Three-way record merge and merge declarations, including the mdbase-next prototype merge fixtures (since 0.3.0-rc.5). + coverage_targets: [merge, collection_semantics] + files: + - merge/merge.yaml + - merge/body-edits.yaml + - id: watch + description: Move detection for externally moved files (since 0.3.0-rc.5). + coverage_targets: [watch] + files: + - watch/move-detection.yaml diff --git a/tests/v0.3/merge/body-edits.yaml b/tests/v0.3/merge/body-edits.yaml new file mode 100644 index 0000000..7824767 --- /dev/null +++ b/tests/v0.3/merge/body-edits.yaml @@ -0,0 +1,304 @@ +# Body edits (Chapters 12 and 12A), since 0.3.0-rc.5. Each apply_body_edits +# case gives the base body, the record's current body, and the edits, whose +# offsets count Unicode scalar values of the base. base_available: false +# models an engine that has neither body_base_text nor a retained copy of the +# base. Expect either the written body or the error and details.reason. +name: body edits against a known base +spec_version: "0.3.0" +fixture_set: merge +category: merge +spec_ref: "v0.3/12, v0.3/12A" +groups: + - name: apply and rebase body edits + tests: + - name: edits apply directly when the body is the base + id: body_edits.direct_replace + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + edits: + - start: 9 + end: 15 + text: costs + expect: + body: | + Agenda + - costs + - hiring + - name: every offset refers to the base body + id: body_edits.base_coordinates + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + edits: + - start: 0 + end: 6 + text: Plan + - start: 25 + end: 25 + text: | + - travel + expect: + body: | + Plan + - budget + - hiring + - travel + - name: "offsets count Unicode scalar values, not bytes or UTF-16 units" + id: body_edits.scalar_offsets + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase] + operation: apply_body_edits + input: + base: | + café 😀 + bar + current: | + café 😀 + bar + edits: + - start: 7 + end: 10 + text: baz + expect: + body: | + café 😀 + baz + - name: an empty text deletes its range + id: body_edits.deletion + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + edits: + - start: 7 + end: 16 + text: "" + expect: + body: | + Agenda + - hiring + - name: edits rebase onto a concurrent change to other lines + id: body_edits.rebase_disjoint_lines + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda (Monday) + - budget + - hiring + edits: + - start: 18 + end: 24 + text: interviews + expect: + body: | + Agenda (Monday) + - budget + - interviews + - name: "an append rebases onto a concurrent append, current text first" + id: body_edits.rebase_append_append + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + - rent + edits: + - start: 25 + end: 25 + text: | + - travel + expect: + body: | + Agenda + - budget + - hiring + - rent + - travel + - name: an edit to a line that changed concurrently is a conflict + id: body_edits.rebase_conflict_same_line + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget 2027 + - hiring + edits: + - start: 9 + end: 15 + text: costs + expect: + error: + code: concurrent_modification + details: + reason: body_conflict + - name: a changed body without the base text cannot be rebased + id: body_edits.base_unavailable + since: "0.3.0-rc.5" + covers: [merge.body_edit_rebase, core_write.body_edits] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda (Monday) + - budget + - hiring + edits: + - start: 9 + end: 15 + text: costs + base_available: false + expect: + error: + code: concurrent_modification + details: + reason: body_base_unavailable + - name: an offset past the end of the base is invalid + id: body_edits.out_of_range + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + edits: + - start: 20 + end: 99 + text: x + expect: + error: + code: invalid_request + details: + reason: offset_out_of_range + - name: overlapping edits are invalid + id: body_edits.overlapping + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + edits: + - start: 0 + end: 6 + text: A + - start: 3 + end: 8 + text: B + expect: + error: + code: invalid_request + details: + reason: edits_overlap_or_unordered + - name: edits out of order are invalid + id: body_edits.unordered + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + edits: + - start: 9 + end: 15 + text: A + - start: 0 + end: 6 + text: B + expect: + error: + code: invalid_request + details: + reason: edits_overlap_or_unordered + - name: two insertions at one offset are invalid + id: body_edits.same_offset_insertions + since: "0.3.0-rc.5" + covers: [core_write.body_edits] + operation: apply_body_edits + input: + base: | + Agenda + - budget + - hiring + current: | + Agenda + - budget + - hiring + edits: + - start: 6 + end: 6 + text: A + - start: 6 + end: 6 + text: B + expect: + error: + code: invalid_request + details: + reason: edits_overlap_or_unordered diff --git a/tests/v0.3/merge/merge.yaml b/tests/v0.3/merge/merge.yaml new file mode 100644 index 0000000..0e472d6 --- /dev/null +++ b/tests/v0.3/merge/merge.yaml @@ -0,0 +1,1717 @@ +# Three-way record merge fixtures (Chapter 12A). +# +# Each merge_records test supplies the base, first, and second versions of one +# record. "first" is the earlier-ordered version; at a conflict the merged +# document keeps its value. expect.document is the exact merged bytes and +# expect.conflicts lists every conflict by kind and field, in key order. +# +# The first group converts the 13 fixtures of the mdbase-next feasibility +# prototype (reference/prototype/crates/mdb-core/tests/merge_fixtures): the +# prototype's "theirs" (sequenced first) is "first" and its "ours" (the +# incoming external edit) is "second". The prototype spelled the max +# declaration as an x-merge schema annotation; here it is collection.merge. +name: "three-way record merge" +spec_version: "0.3.0" +fixture_set: merge +category: merge +spec_ref: "v0.3/07, v0.3/12A" +groups: + - name: "prototype merge fixtures (TaskNotes-like task type)" + setup: + config: | + spec_version: "0.3.0" + types: + task.md: | + --- + kind: mdbase.type + name: task + match: + path_glob: "tasks/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + required: [title] + properties: + title: { type: string, minLength: 1 } + status: { type: string, enum: [open, doing, done] } + priority: { type: integer, minimum: 1, maximum: 5 } + blocked_by: { type: array, uniqueItems: true, items: { type: string } } + completedDate: { type: string } + collection: + read_defaults: + status: open + unique: + - field: code + scope: type + enforce: write + links: + blocked_by[]: + target_type: task + path: + pattern: "tasks/{title}.md" + merge: + completedDate: max + lifecycle: + on_create: + - if: '!has(raw.id)' + set: + id: { ulid: true } + - set: + dateCreated: { now: true } + dateModified: { now: true } + on_update: + set: + dateModified: { now: true } + --- + # Task + tests: + - name: "Both append different lines at the end: union (theirs first), no conflict. (prototype fixture body-append-append)" + id: merge.body_append_append + since: 0.3.0-rc.5 + covers: [merge.append_append_union] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + - one + first: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + - one + - theirs + second: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + - one + - ours + expect: + document: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + - one + - theirs + - ours + conflicts: [] + - name: "Different body lines edited: both kept. (prototype fixture body-disjoint-hunks)" + id: merge.body_disjoint_hunks + since: 0.3.0-rc.5 + covers: [merge.body_three_way] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + l1 + l2 + l3 + l4 + l5 + l6 + first: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + l1 theirs + l2 + l3 + l4 + l5 + l6 + second: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + l1 + l2 + l3 + l4 + l5 + l6 ours + expect: + document: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + l1 theirs + l2 + l3 + l4 + l5 + l6 ours + conflicts: [] + - name: "Same line edited differently: conflict; theirs stays. (prototype fixture body-overlap)" + id: merge.body_overlap + since: 0.3.0-rc.5 + covers: [merge.body_three_way, merge.conflict_reporting] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + x + y + z + first: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + x + y theirs + z + second: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + x + y ours + z + expect: + document: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + x + y theirs + z + conflicts: + - kind: body + - name: "completedDate declared merge: max. (prototype fixture declared-max)" + id: merge.declared_max + since: 0.3.0-rc.5 + covers: [merge.declared_strategies, merge.max_min_semantics] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + completedDate: 2026-10-01 + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + completedDate: 2026-10-05 + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + completedDate: 2026-10-04 + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + completedDate: 2026-10-05 + --- + line1 + line2 + line3 + conflicts: [] + - name: "Different keys changed on each side: both kept. (prototype fixture disjoint-fields)" + id: merge.disjoint_fields + since: 0.3.0-rc.5 + covers: [merge.field_three_way] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: doing + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: open + priority: 4 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: doing + priority: 4 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: [] + - name: "A field change and a body change merge. (prototype fixture field-vs-body)" + id: merge.field_vs_body + since: 0.3.0-rc.5 + covers: [merge.field_three_way, merge.body_three_way] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 ours + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 ours + line3 + conflicts: [] + - name: "Untouched entries keep comments/quoting; the changed key is copied verbatim from the side that changed it. (prototype fixture format-preserved)" + id: merge.format_preserved + since: 0.3.0-rc.5 + covers: [merge.merge_format_fidelity] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 # p + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: 'doing' # moved + priority: 2 # p + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: open + priority: 3 # p + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: 'doing' # moved + priority: 3 # p + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: [] + - name: "dateModified changed on both sides: the later timestamp wins, no conflict. (prototype fixture lifecycle-now-max)" + id: merge.lifecycle_now_max + since: 0.3.0-rc.5 + covers: [merge.default_strategies, merge.max_min_semantics] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-03T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-02T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-03T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: [] + - name: "Ours sets priority 9 (violates maximum 5) via an external edit: accepted, validity reported, not a conflict. (prototype fixture merge-invalid-is-reported)" + id: merge.merge_invalid_is_reported + since: 0.3.0-rc.5 + covers: [merge.invalid_merge_reported] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: doing + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: open + priority: 9 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: doing + priority: 9 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: [] + - name: "Default merge: conflict. The earlier-sequenced value (theirs) stays. (prototype fixture same-field-different-values)" + id: merge.same_field_different_values + since: 0.3.0-rc.5 + covers: [merge.field_three_way, merge.conflict_reporting] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: doing + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: doing + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: + - kind: field + field: status + - name: "Both sides set the same value: no conflict. (prototype fixture same-field-same-value)" + id: merge.same_field_same_value + since: 0.3.0-rc.5 + covers: [merge.field_three_way] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: done + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: [] + - name: "Both add different tags: union. (prototype fixture tags-union-add-add)" + id: merge.tags_union_add_add + since: 0.3.0-rc.5 + covers: [merge.default_strategies, merge.union_semantics] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b, c] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b, d] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b, c, d] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: [] + - name: "One side removes b, the other adds c: remove and add both apply. (prototype fixture tags-union-remove-add)" + id: merge.tags_union_remove_add + since: 0.3.0-rc.5 + covers: [merge.union_semantics] + operation: merge_records + input: + path: tasks/T.md + base: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + first: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, b, c] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + second: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + expect: + document: | + --- + type: task + title: T # keep me + status: open + priority: 2 + tags: [a, c] + dateModified: 2026-10-01T00:00:00.000Z + --- + line1 + line2 + line3 + conflicts: [] + - name: "declared strategies, defaults, paths, and edge cases" + setup: + config: | + spec_version: "0.3.0" + types: + item.md: | + --- + kind: mdbase.type + name: item + version: 1 + match: + path_glob: "items/**/*.md" + schema: + dialect: json-schema-2020-12 + value: + type: object + properties: + title: { type: string } + score: { type: number } + reviewers: { type: array, uniqueItems: true, items: { type: string } } + labels: { type: array, items: { type: string } } + collection: + merge: + firstSeen: min + lastSeen: max + score: max + dateModified: conflict + lifecycle: + on_update: + set: + dateModified: { now: true } + --- + tests: + - name: a field declared min keeps the lesser of two changed values + id: merge.declared_min + since: 0.3.0-rc.5 + covers: [merge.declared_strategies, merge.max_min_semantics] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + firstSeen: 2026-10-03 + --- + Body. + first: | + --- + title: Item + score: 1 + firstSeen: 2026-10-02 + --- + Body. + second: | + --- + title: Item + score: 1 + firstSeen: 2026-10-01 + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + firstSeen: 2026-10-01 + --- + Body. + conflicts: [] + - name: a declaration replaces the max default of a lifecycle now field + id: merge.declaration_overrides_default + since: 0.3.0-rc.5 + covers: [merge.declared_strategies, merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + dateModified: 2026-10-01T00:00:00Z + --- + Body. + first: | + --- + title: Item + score: 1 + dateModified: 2026-10-02T00:00:00Z + --- + Body. + second: | + --- + title: Item + score: 1 + dateModified: 2026-10-03T00:00:00Z + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + dateModified: 2026-10-02T00:00:00Z + --- + Body. + conflicts: + - kind: field + field: dateModified + - name: "max compares date-times with offsets as instants, not as text" + id: merge.max_compares_instants + since: 0.3.0-rc.5 + covers: [merge.max_min_semantics] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + lastSeen: 2026-10-01T00:00:00Z + --- + Body. + first: | + --- + title: Item + score: 1 + lastSeen: 2026-10-02T09:00:00+10:00 + --- + Body. + second: | + --- + title: Item + score: 1 + lastSeen: 2026-10-01T23:30:00Z + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + lastSeen: 2026-10-01T23:30:00Z + --- + Body. + conflicts: [] + - name: max of a number and a string is a conflict + id: merge.max_incomparable_conflicts + since: 0.3.0-rc.5 + covers: [merge.max_min_semantics, merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 5 + --- + Body. + second: | + --- + title: Item + score: high + --- + Body. + expect: + document: | + --- + title: Item + score: 5 + --- + Body. + conflicts: + - kind: field + field: score + - name: under max a value survives a concurrent removal + id: merge.max_value_beats_removal + since: 0.3.0-rc.5 + covers: [merge.max_min_semantics] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + --- + Body. + second: | + --- + title: Item + score: 3 + --- + Body. + expect: + document: | + --- + title: Item + score: 3 + --- + Body. + conflicts: [] + - name: an array with uniqueItems merges as a set by default and keeps block style + id: merge.unique_items_default_union + since: 0.3.0-rc.5 + covers: [merge.default_strategies, merge.union_semantics, merge.merge_format_fidelity] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + reviewers: + - ann + --- + Body. + first: | + --- + title: Item + score: 1 + reviewers: + - ann + - bo + --- + Body. + second: | + --- + title: Item + score: 1 + reviewers: + - cy + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + reviewers: + - bo + - cy + --- + Body. + conflicts: [] + - name: an array without uniqueItems has the conflict default + id: merge.plain_array_conflicts + since: 0.3.0-rc.5 + covers: [merge.default_strategies, merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + labels: [x] + --- + Body. + first: | + --- + title: Item + score: 1 + labels: [x, y] + --- + Body. + second: | + --- + title: Item + score: 1 + labels: [x, z] + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + labels: [x, y] + --- + Body. + conflicts: + - kind: field + field: labels + - name: union treats a removed key as an empty list and re-emits the result + id: merge.union_with_removed_key + since: 0.3.0-rc.5 + covers: [merge.union_semantics] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + reviewers: [ann] + --- + Body. + first: | + --- + title: Item + score: 1 + --- + Body. + second: | + --- + title: Item + score: 1 + reviewers: [ann, bo] + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + reviewers: [bo] + --- + Body. + conflicts: [] + - name: an empty union result with a removed key leaves the key missing + id: merge.union_empty_removes_key + since: 0.3.0-rc.5 + covers: [merge.union_semantics] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + reviewers: [ann] + --- + Body. + first: | + --- + title: Item + score: 1 + --- + Body. + second: | + --- + title: Item + score: 1 + reviewers: [] + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + --- + Body. + conflicts: [] + - name: a tags string reads as a one-item list + id: merge.tags_string_is_one_item + since: 0.3.0-rc.5 + covers: [merge.union_semantics] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + tags: x + --- + Body. + first: | + --- + title: Item + score: 1 + tags: [x, y] + --- + Body. + second: | + --- + title: Item + score: 1 + tags: z + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + tags: [y, z] + --- + Body. + conflicts: [] + - name: a key removed on one side and unchanged on the other is removed + id: merge.one_side_removes + since: 0.3.0-rc.5 + covers: [merge.field_three_way] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + note: old + --- + Body. + first: | + --- + title: Item + score: 1 + note: old + --- + Body. + second: | + --- + title: Item + score: 1 + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + --- + Body. + conflicts: [] + - name: a key added on both sides with different values is a conflict + id: merge.added_both_sides_differently + since: 0.3.0-rc.5 + covers: [merge.field_three_way, merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 1 + owner: ann + --- + Body. + second: | + --- + title: Item + score: 1 + owner: bo + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + owner: ann + --- + Body. + conflicts: + - kind: field + field: owner + - name: setting null and removing a key are different changes + id: merge.null_is_not_missing + since: 0.3.0-rc.5 + covers: [merge.field_three_way, merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + owner: ann + --- + Body. + first: | + --- + title: Item + score: 1 + owner: null + --- + Body. + second: | + --- + title: Item + score: 1 + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + owner: null + --- + Body. + conflicts: + - kind: field + field: owner + - name: a key added by the second side goes after the last entry and before trailing comments + id: merge.new_key_after_last_entry + since: 0.3.0-rc.5 + covers: [merge.merge_format_fidelity] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item # name + score: 1 + # end of fields + --- + Body. + first: | + --- + title: Renamed # name + score: 1 + # end of fields + --- + Body. + second: | + --- + title: Item # name + score: 1 + owner: "bo" + # end of fields + --- + Body. + expect: + document: | + --- + title: Renamed # name + score: 1 + owner: "bo" + # end of fields + --- + Body. + conflicts: [] + - name: append-append inserts a line break after an append that lacks one + id: merge.append_without_final_newline + since: 0.3.0-rc.5 + covers: [merge.append_append_union] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + a + first: |- + --- + title: Item + score: 1 + --- + a + b + second: | + --- + title: Item + score: 1 + --- + a + c + expect: + document: | + --- + title: Item + score: 1 + --- + a + b + c + conflicts: [] + - name: two insertions at the same place inside the body are a conflict + id: merge.interior_insertions_conflict + since: 0.3.0-rc.5 + covers: [merge.body_three_way, merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + a + z + first: | + --- + title: Item + score: 1 + --- + a + b + z + second: | + --- + title: Item + score: 1 + --- + a + c + z + expect: + document: | + --- + title: Item + score: 1 + --- + a + b + z + conflicts: + - kind: body + - name: "when the first version is unchanged the merged version is the second, byte for byte" + id: merge.unchanged_first_takes_second_bytes + since: 0.3.0-rc.5 + covers: [merge.field_three_way, merge.merge_format_fidelity] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 1 + --- + Body. + second: | + --- + # reviewed + title: Item + score: 1 + --- + Body. + expect: + document: | + --- + # reviewed + title: Item + score: 1 + --- + Body. + conflicts: [] + - name: two identical edits merge to that edit + id: merge.identical_edits + since: 0.3.0-rc.5 + covers: [merge.field_three_way, merge.body_three_way] + operation: merge_records + input: + path: items/a.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 2 + --- + Body! + second: | + --- + title: Item + score: 2 + --- + Body! + expect: + document: | + --- + title: Item + score: 2 + --- + Body! + conflicts: [] + - name: a move on one side wins over an unchanged path + id: merge.path_one_side_moves + since: 0.3.0-rc.5 + covers: [merge.path_merge] + operation: merge_records + input: + paths: + base: items/a.md + first: items/a.md + second: items/b.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 1 + --- + Body. + second: | + --- + title: Item + score: 2 + --- + Body. + expect: + document: | + --- + title: Item + score: 2 + --- + Body. + conflicts: [] + path: items/b.md + - name: two different moves are a path conflict + id: merge.path_two_moves_conflict + since: 0.3.0-rc.5 + covers: [merge.path_merge, merge.conflict_reporting] + operation: merge_records + input: + paths: + base: items/a.md + first: items/b.md + second: items/c.md + base: | + --- + title: Item + score: 1 + --- + Body. + first: | + --- + title: Item + score: 1 + --- + Body. + second: | + --- + title: Item + score: 1 + --- + Body. + expect: + document: | + --- + title: Item + score: 1 + --- + Body. + conflicts: + - kind: path + path: items/b.md + - name: frontmatter that is not a mapping merges as one unit + id: merge.non_mapping_frontmatter + since: 0.3.0-rc.5 + covers: [merge.conflict_reporting] + operation: merge_records + input: + path: items/a.md + base: | + --- + - x + --- + Body. + first: | + --- + - y + --- + Body. + second: | + --- + - z + --- + Body. + expect: + document: | + --- + - y + --- + Body. + conflicts: + - kind: frontmatter + - name: declared and default strategies for one record + id: merge.strategies_item + since: 0.3.0-rc.5 + covers: [merge.default_strategies, merge.declared_strategies, collection_semantics.merge_declarations] + operation: merge_strategies + input: + path: items/a.md + document: | + --- + title: Item + score: 1 + --- + Body. + expect: + strategies: + firstSeen: min + lastSeen: max + score: max + dateModified: conflict + reviewers: union + tags: union + labels: conflict + title: conflict + - name: merge declarations across several matched types + setup: + config: | + spec_version: "0.3.0" + types: + alpha.md: | + --- + kind: mdbase.type + name: alpha + version: 1 + match: + path_glob: "shared/**/*.md" + schema: + dialect: json-schema-2020-12 + value: { type: object } + collection: + merge: + rank: max + note: union + --- + beta.md: | + --- + kind: mdbase.type + name: beta + version: 1 + match: + path_glob: "shared/**/*.md" + schema: + dialect: json-schema-2020-12 + value: { type: object } + collection: + merge: + rank: min + note: union + --- + tests: + - name: different declarations for one field from two matched types are a type_conflict + id: merge.declaration_type_conflict + since: 0.3.0-rc.5 + covers: [collection_semantics.merge_declarations] + operation: merge_strategies + input: + path: shared/x.md + document: | + --- + rank: 1 + --- + expect: + error: + code: type_conflict + field: rank + - name: identical declarations from two matched types coalesce + id: merge.declaration_coalesce + since: 0.3.0-rc.5 + covers: [collection_semantics.merge_declarations] + operation: merge_strategies + input: + path: shared/x.md + document: | + --- + note: [a] + --- + fields: [note] + expect: + strategies: + note: union diff --git a/tests/v0.3/watch/move-detection.yaml b/tests/v0.3/watch/move-detection.yaml new file mode 100644 index 0000000..fb137c4 --- /dev/null +++ b/tests/v0.3/watch/move-detection.yaml @@ -0,0 +1,433 @@ +# Move detection (Chapter 12A). Each test lists what disappeared and what +# appeared within one observation window. file_id stands for a platform file +# identity such as a device and inode; omit it when the platform cannot supply +# one. Cases follow the identity measurements in the mdbase-next prototype +# (SPEC-CHANGES SC9): rename, rename then edit, delete and create with the +# same bytes, delete and create of an edited file, and unrelated files. +name: "move detection" +spec_version: "0.3.0" +fixture_set: watch +category: watch +spec_ref: v0.3/12A +groups: + - name: pairing disappearances and appearances within one observation window + tests: + - name: a rename with unchanged bytes is a move + id: watch.moves_rename_same_content + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + file_id: f1 + appeared: + - path: archive/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + file_id: f1 + expect: + moves: + - from: notes/a.md + to: archive/a.md + deleted: [] + created: [] + - name: a rename followed by an edit keeps the file identity + id: watch.moves_rename_then_edit + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + file_id: f1 + appeared: + - path: notes/b.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - other 1 + - other 2 + - other 3 + file_id: f1 + expect: + moves: + - from: notes/a.md + to: notes/b.md + deleted: [] + created: [] + - name: "delete and create with identical bytes is a move (git checkout)" + id: watch.moves_delete_create_same_bytes + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + file_id: f1 + appeared: + - path: notes/b.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + file_id: f2 + expect: + moves: + - from: notes/a.md + to: notes/b.md + deleted: [] + created: [] + - name: delete and create of an edited file with the same name in another folder is a move + id: watch.moves_same_name_edited + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + file_id: f1 + appeared: + - path: archive/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item nine + file_id: f2 + expect: + moves: + - from: notes/a.md + to: archive/a.md + deleted: [] + created: [] + - name: the same name with similarity below 0.8 and no file identity is not a move + id: watch.moves_same_name_too_different + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + appeared: + - path: archive/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - other 1 + - other 2 + - other 3 + expect: + moves: [] + deleted: [notes/a.md] + created: [archive/a.md] + - name: an unrelated delete and create are never paired + id: watch.moves_unrelated_files + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + file_id: f1 + appeared: + - path: lists/shopping.md + content: | + --- + title: Shopping + --- + - milk + - bread + file_id: f2 + expect: + moves: [] + deleted: [notes/a.md] + created: [lists/shopping.md] + - name: equal id_field values pair even when the content changed completely + id: watch.moves_identity_hint_pairs + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + id_field: id + disappeared: + - path: tasks/Call Bob.md + content: | + --- + id: t1 + title: Call Bob + --- + Old notes. + appeared: + - path: tasks/Call Robert.md + content: | + --- + id: t1 + title: Call Robert + --- + Rewritten completely. + expect: + moves: + - from: tasks/Call Bob.md + to: tasks/Call Robert.md + deleted: [] + created: [] + - name: "different id_field values never pair, even with the same name and file identity" + id: watch.moves_identity_hint_vetoes + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + id_field: id + disappeared: + - path: tasks/Call Bob.md + content: | + --- + id: t1 + title: Call Bob + --- + Old notes. + file_id: f1 + appeared: + - path: done/Call Bob.md + content: | + --- + id: t2 + title: Call Bob + --- + Old notes. + file_id: f1 + expect: + moves: [] + deleted: [tasks/Call Bob.md] + created: [done/Call Bob.md] + - name: "when two appearances match equally, the smaller path pairs" + id: watch.moves_copies_pair_deterministically + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + appeared: + - path: z/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + - path: b/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + expect: + moves: + - from: notes/a.md + to: b/a.md + deleted: [] + created: [z/a.md] + - name: identical bytes outrank a same-name match + id: watch.moves_identical_bytes_beat_name + since: 0.3.0-rc.5 + covers: [watch.move_detection] + operation: detect_moves + input: + disappeared: + - path: notes/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + appeared: + - path: archive/a.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item nine + - path: archive/renamed.md + content: | + --- + title: Weekly review + --- + - item 1 + - item 2 + - item 3 + - item 4 + - item 5 + - item 6 + - item 7 + - item 8 + - item 9 + expect: + moves: + - from: notes/a.md + to: archive/renamed.md + deleted: [] + created: [archive/a.md] From 619c5acf7d97fabe5fa76f42e79c4de35be75c5d Mon Sep 17 00:00:00 2001 From: callumalpass Date: Sun, 4 Oct 2026 14:30:42 +1100 Subject: [PATCH 4/5] Add the 0.3.0-rc.5 changelog entry and release notes The rc.5 entry covers concurrent edits, reported validity, the regex profile, and body_edits, and folds in the changes made since rc.4 to YAML document records and Bases, seed type upgrades, and saved-view identification. The release notes list the behavior changes, the one existing test whose expectation changed and the new tests an rc.4 engine fails, the three choices decided for rc.5 (unsupported patterns invalidate the type, body-edit conflicts are per line, time-dependent match.expr becomes an error in 0.3.0 stable), and the remaining provisional choices. --- CHANGELOG.md | 101 ++++++++++- docs/releases/0.3.0-rc.5.md | 339 ++++++++++++++++++++++++++++++++++++ 2 files changed, 439 insertions(+), 1 deletion(-) create mode 100644 docs/releases/0.3.0-rc.5.md diff --git a/CHANGELOG.md b/CHANGELOG.md index 0bb08cc..5220515 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,7 +2,104 @@ All notable changes to this specification and conformance suite are documented here. -## Unreleased +## 0.3.0-rc.5 (draft, untagged) + +The fifth release candidate specifies the data semantics that let several +tools, devices, and people edit one collection at once: validity is reported +rather than guaranteed, concurrent edits merge field by field, and writers +change only the bytes they mean to change. Most changes relax write-time +checks; the few tightenings are reported as diagnostics. Collections keep +`spec_version: "0.3.0"`. Replication, logs, sequencers, conflict envelopes, +and delete-versus-update policy remain implementation concerns. Release +notes, including every provisional choice and every test whose expectation +changed: [docs/releases/0.3.0-rc.5.md](./docs/releases/0.3.0-rc.5.md). + +rc.5 also ships the changes made since rc.4 to YAML document records and +Bases, seed type upgrades, and saved-view identification. + +### Concurrent edits, reported validity, regex profile, and body edits + +#### Added + +- Chapter 12A, Concurrent Edits: record identity without IDs in files, move + detection, the three-way record merge, and the writer format fidelity rule. +- `collection.merge` declares `conflict`, `max`, `min`, or `union` per + top-level field. Defaults: `max` for fields that lifecycle assigns with + `now` or `today`; `union` for `tags` and for arrays declared + `uniqueItems: true`; `conflict` otherwise. +- Update accepts `add` and `remove` list operations, applied to the current + value instead of replacing the list. +- Update accepts `body_edits`: text edits whose offsets count Unicode scalar + values of a base body identified by `body_base`, a SHA-256 digest. They + apply directly to an unchanged body and otherwise rebase with the + three-way body merge, failing with `body_conflict` or + `body_base_unavailable`. They are mutually exclusive with `body` and + `document`. +- The mdbase regex profile (Chapter 10): RE2 syntax with ASCII-only `\d`, + `\w`, `\s`, `\b`, and case folding over Unicode scalar values, used by CEL + `matches()`, JSON Schema `pattern`, and `match.where` `matches`. Unicode + classes, backreferences, and look-around are `invalid_pattern`. +- `collection.unique[].enforce: write | report`, defaulting to `report`. +- Path equivalence: paths that are equal after NFC normalization and full + Unicode case folding name one record path. Explicit paths that collide fail + with `path_conflict`; derived paths and concurrently created records get a + deterministic ` (n)` suffix; discovered collisions report `path_collision`. +- The `merge` conformance profile (requires `core_write`), in the profile + list and in `schemas/v0.3/conformance-claim.schema.json`; `collection.merge` + and `unique[].enforce` in `schemas/v0.3/type-file.schema.json`. +- Diagnostic codes `path_collision`, `path_value_invalid`, `ambiguous_link`, + `link_target_type_mismatch`, `nondeterministic_match`, `invalid_pattern`, + and `duplicate_value` are listed as core codes. +- The v0.3 suite gains 147 tests marked `since: 0.3.0-rc.5`, including the 13 + merge fixtures of the mdbase-next prototype in a new `merge_records` format + and new `merge` and `watch` fixture sets. `scripts/check_v03_tests.py` runs + the 95 pure-function fixtures against an executable model + (`scripts/concurrent_edits_model.py`) in CI. + +#### Changed + +- Validity is reported, never guaranteed. Checks are split into request and + safety, single-record, and cross-record tiers. At level `error` only + single-record issues reject a write made through an engine; cross-record + issues (uniqueness in `report` mode, `validate_exists`, `target_type`, + ambiguous links, path collisions) are reported and never block, and a + successful write reports them as warnings. +- Uniqueness scopes are exact: a rule governs records of its declaring type, + `scope` selects the records they must differ from, and raw values compare + without coercion. The default scope is `type`. +- A `match.expr` that uses `now()`, `today()`, `file.mtime`, `file.ctime`, or + helpers that read other records loads with a `nondeterministic_match` + warning. It becomes an error that invalidates the type in 0.3.0 stable. +- CEL is the one expression language. Obsidian Bases filters and formulas are + an adapter dialect that never affects membership, validation, lifecycle, + merge, or portable queries. +- Writers re-emit only changed top-level frontmatter entries and keep flow or + block collection style, for Markdown records and YAML document records. +- A taken derived path receives a ` (n)` suffix instead of failing. +- `collection.path.pattern` values containing `/` or `\`, beginning with `.`, + empty, or not scalars fail with `path_value_invalid`. +- An ambiguous configured ID resolves to null with `ambiguous_link` and no + filename fallback. Filename tiebreakers are required and end in code-point + order. +- `if_revision` is opt-in; transports and SDKs must not add it on a caller's + behalf. +- Lifecycle states that it has no sequence provider; implementations may + offer one only under an `x-*` extension. v0.2 `generated: sequence` is + reported as unsupported by v0.2 migration. +- Test `links.duplicate_id_ambiguous` now also expects the `ambiguous_link` + diagnostic. + +#### Clarified + +- `settings.id_field` has no default; engines that resolve through `id` + without configuration are non-conforming. It also serves as the + move-detection identity hint. + +#### Migration + +- Collections need no edit. Add `enforce: write` to uniqueness rules that must + keep blocking duplicate writes. Chapter 13 ("Changes Since rc.4") lists + every behavior change and how engines report the tightenings. ### YAML document records and Bases as records @@ -15,6 +112,8 @@ All notable changes to this specification and conformance suite are documented h through `x-obsidian.bases.include` and the saved-view source operations for Bases become transitional. +### Seed type upgrades + - A seed type resource may declare `upgrade_from`: one baseline or a list of baselines, each a digest-pinned starter the publisher previously shipped (`{ digest, document, version? }`), so a collection can upgrade from any diff --git a/docs/releases/0.3.0-rc.5.md b/docs/releases/0.3.0-rc.5.md new file mode 100644 index 0000000..6c0a5e0 --- /dev/null +++ b/docs/releases/0.3.0-rc.5.md @@ -0,0 +1,339 @@ +# mdbase-spec v0.3.0-rc.5 — Release Notes (Draft) + +**Status:** draft on branch `spec/v0.3.0-rc.5-draft`; not tagged. + +The fifth release candidate of mdbase 0.3.0 specifies what happens when a +collection is edited by several tools at once: an application writing through +an engine, a person in a text editor, a synchronization tool, an agent on +another device. The rules come from the mdbase-next feasibility prototype and +its invariant-confluence analysis (`SPEC-CHANGES.md`, changes SC0 to SC11). + +The changes stay inside the 0.3.0 release-candidate series. Most of them relax +write-time checks, and the few tightenings are reported as diagnostics, so no +collection that worked under rc.4 stops loading. Collections keep +`spec_version: "0.3.0"` and need no edit. Engines that conform to rc.4 keep +claiming `0.3.0-rc.4` until they pass the rc.5 suite; 0.3.0 is declared stable +once an engine does. + +Four principles hold the release together: + +- **Plain Markdown.** No user file ever requires mdbase metadata. Record + identity, revisions, and merge state stay outside records. +- **Validity is reported, not guaranteed.** Any tool can write any bytes, so + conforming tools read and report invalid records instead of refusing them. + Only checks within a single record may reject a write made through an + engine. +- **Field-level merge.** Two concurrent edits of a record combine field by + field with declared strategies, so only real disagreements are conflicts. +- **Data semantics only.** The specification defines what a merge produces. + When tools merge, how edits travel, logs, sequencers, conflict envelopes, + and delete-versus-update policy are implementation concerns. + +## New + +| Change | Where | From | +| --- | --- | --- | +| Validation principle and three validation tiers | Chapter 04 | SC0, SC4 | +| `collection.merge`: `conflict`, `max`, `min`, `union`, with defaults | Chapter 07 | SC1 | +| `add` and `remove` list operations on update | Chapter 12 | SC2 | +| `collection.unique[].enforce: write \| report` and exact scopes | Chapter 07 | SC3 | +| Path keys (NFC + full case folding) and the ` (n)` collision rule | Chapter 02 | SC5 | +| `nondeterministic_match` for time-dependent or cross-record `match.expr` | Chapters 07, 10 | SC6 | +| No sequence provider | Chapter 09 | SC7 | +| CEL as the one expression language; Bases as an adapter dialect | Chapters 04, 10 | SC8 | +| Record identity without in-file IDs; move detection | Chapter 12A | SC9 | +| Append-append body union | Chapter 12A | SC10 | +| Writer format fidelity | Chapters 03, 12A | SC11 | +| `body_edits`: text edits against a `body_base` digest, rebased with the body merge | Chapters 12, 12A | ADR 0014 readiness | +| The mdbase regex profile: RE2 syntax, ASCII-only classes and case folding, everywhere | Chapters 06, 07, 10 | WASM size, replay determinism | +| `merge` conformance profile | Chapter 14 | — | +| `collection.merge`, `unique[].enforce`, and the `merge` profile in `schemas/v0.3/` | — | — | + +Drift found by the analysis of the current engine is fixed in the text and +pinned by tests: `settings.id_field` has no default, path-pattern values may +not contain `/`, and an ambiguous ID resolves to null with `ambiguous_link` +and no filename fallback. rc.4 already said all three; engines did not follow +it. + +## Behavior changes since rc.4 + +- **Uniqueness reports by default.** At level `error`, rc.4 rejected a write + that created a duplicate. In rc.5 a duplicate is reported, and the write + succeeds with a warning, unless the rule says `enforce: write`. +- **Link checks never block.** `validate_exists` and `target_type` are + reported, never used to reject a write. +- **Successful writes may carry warnings.** A write that leaves a + cross-record issue returns `valid: true` with `warning` diagnostics. +- **Concurrent edits merge.** Tools that reconcile edits merge field by + field: timestamps maintained by lifecycle take the later value, and `tags` + and `uniqueItems` arrays merge as sets. +- **A taken derived path is suffixed** with ` (2)` instead of failing. +- **Writes are byte-minimal.** Untouched frontmatter entries, comments, and + quoting are preserved exactly, including in YAML document records. +- **Body edits.** An update may send `body_edits` against a `body_base` + digest instead of a whole body. It works without any live-collaboration + machinery: an engine applies the edits directly when the body is unchanged, + and rebases them with the line-based body merge otherwise. +- **One regex flavor.** CEL `matches()`, JSON Schema `pattern`, and + `match.where` `matches` all use RE2 syntax with ASCII-only `\d`, `\w`, + `\s`, `\b`, and case folding, over Unicode scalar values. rc.4 engines that + used Unicode classes diverge on non-ASCII text, for example `\w` against + `é`. + +Tightenings, each reported as a diagnostic: + +- **Time-dependent or cross-record `match.expr`** loads with a + `nondeterministic_match` warning and is evaluated as before. It becomes an + error in 0.3.0 stable. +- **Equivalent paths collide.** `Todo.md` and `todo.md` are one path. Existing + equivalent files stay records and report `path_collision`; an explicit + create or rename onto an equivalent path fails with `path_conflict`. +- **Derived paths stay in one folder.** Existing records are untouched; a + create that would derive a value with `/` or `\`, a leading `.`, or an empty + value fails with `path_value_invalid`. +- **Ambiguous IDs** resolve to null and report `ambiguous_link`. +- **Unicode regex classes** (`\p{L}`), backreferences, and look-around are + invalid patterns: a JSON Schema or `match.where` pattern using them makes + its type invalid with `invalid_pattern`, and a CEL literal is an + `expression_compile_error`. + +Chapter 13 ("Changes Since rc.4") has the full tables, the list of what needs +no change, and the engine checklist. + +## Conformance + +The rc.5 additions live in the v0.3 suite (`tests/v0.3/`) and carry +`since: 0.3.0-rc.5`. There are 147 new tests; 95 are pure functions that CI +executes against an executable model (`scripts/concurrent_edits_model.py`) +through `scripts/check_v03_tests.py`: + +- `merge/merge.yaml` (new `merge` fixture set): the 13 prototype merge + fixtures in a new `merge_records` format (base, first, and second versions + plus type declarations, with the exact merged bytes and conflicts + expected), and 24 more merge and strategy cases +- `core/paths.yaml`: path keys, collision suffixes, and path-pattern values +- `watch/move-detection.yaml` (new `watch` fixture set): the SC9 pairing + cases, including the false-pairing guard and the identity hint +- `cel/regex-profile.yaml`: 18 pure `regex_match` cases where ASCII and + Unicode classes differ, plus CEL, JSON Schema, and `match.where` cases +- `merge/body-edits.yaml`: 12 pure `apply_body_edits` cases (direct, rebased, + conflict, base unavailable, invalid ranges); `core/body-edits-update.yaml` + covers `update` with `body_edits` +- adapter-target suites `core/validation-tiers.yaml`, + `core/list-ops-and-fidelity.yaml`, `core/link-ambiguity.yaml`, and + `cel/match-determinism.yaml` + +`tests/v0.3/README.md` documents the `since`/`changed` markers and the new +operations, including the `merge_records` format. + +### Tests whose expected outcome changed from rc.4 + +An engine that conforms to rc.4 diverges on exactly these existing tests: + +| Test | rc.4 expectation | rc.5 expectation | +| --- | --- | --- | +| `links.duplicate_id_ambiguous` (`core/links-and-discovery.yaml`) | resolves to null | resolves to null **and** reports an `ambiguous_link` diagnostic with `details.candidates` | + +One existing test was renamed without changing its expectation: "collection +links enforce validate_exists" (`core/core-collection.yaml`, a `validate` +test) is now "collection links report validate_exists". `validate` still +reports `link_not_found` with the level's severity. + +No other rc.4 test contradicted the new semantics: rc.4 had no test that a +duplicate or a broken link rejects a write, no derived-path collision test, +and no format test for structured writes. + +### New tests that contradict rc.4 engine behavior + +These new tests encode the relaxations and tightenings above, so an rc.4 +engine fails them even though rc.4 had no test for the case: + +| Test | rc.4 engine behavior | +| --- | --- | +| `collection.unique_report_default_writes` | rejects the duplicate at level `error` | +| `links.validate_exists_non_blocking`, `links.target_type_non_blocking` | reject the write at level `error` | +| `paths.derived_create_suffix` | fails with `path_conflict` | +| `paths.explicit_create_conflict`, `paths.rename_conflict`, `paths.discovered_collision` | treat case-different paths as distinct | +| `paths.derived_slash_rejected`, `paths.derive_slash` | create a subfolder | +| `links.no_default_id_field` | some engines default `id_field` to `id` | +| `links.ambiguous_id_validate` | no `ambiguous_link` diagnostic | +| `cel_match.nondeterministic_match_warning` | no diagnostic | +| `core_write.patch_minimal_bytes`, `core_write.new_key_appended`, `core_write.yaml_document_fidelity`, `core_write.add_flow_style`, `core_write.add_block_style` | may re-emit the frontmatter | +| every `core_write` list-operation test | `add` and `remove` are unknown request members | +| every `core_write.body_edits_*` test | `body_edits` is an unknown request member | +| `regex.word_ascii_only`, `regex.not_word_non_ascii`, `regex.digit_ascii_only`, `regex.space_ascii_only`, `regex.boundary_ascii`, `regex.cel_matches_ascii`, `regex.schema_pattern_ascii`, `regex.match_where_ascii` | a Unicode-aware engine treats `é` as a word character, `١` as a digit, a no-break space as white space, and finds no boundary inside `café` | +| `regex.ignore_case_ascii_only` | Unicode case folding matches `É` with `(?i)é` | +| `regex.unicode_class_invalid`, `regex.cel_literal_invalid`, `regex.schema_unicode_class_invalid` | accepts `\p{L}` | +| `regex.dot_astral` | a UTF-16 engine without the Unicode flag sees two code units | + +## Decided for rc.5 + +These choices were open during drafting and are now settled. + +- **Unsupported patterns invalidate the type.** A JSON Schema `pattern` or + `match.where` pattern that uses a Unicode class such as `\p{L}`, a + backreference, or look-around makes its type invalid with + `invalid_pattern`; it is not skipped with a warning, because silently + dropping a constraint is worse than failing to load. This is the one rc.5 + tightening that can stop a type that loaded under rc.4 from loading. +- **Body-edit conflicts are detected per line.** Rebasing `body_edits` uses + the line-based body merge, so edits to different words of one line + conflict. A finer granularity may be added later as a compatible + refinement, since it only turns some conflicts into merges. +- **Time-dependent `match.expr` becomes an error in 0.3.0 stable.** In rc.5 a + `match.expr` that uses a time-dependent or cross-record binding loads with + a `nondeterministic_match` warning and is evaluated as before. In 0.3.0 + stable it is an error that invalidates the type. The release candidates + are the migration window. + +## Provisional choices + +The design input left the following details open. Each has a provisional +choice in the release candidate, marked **Provisional (rc.5)** in the chapter +where it is normative or listed here, and each may change before 0.3.0 is +declared stable. + +### Merge + +1. **Where merge strategies are declared.** SC1 says "a field may declare + `merge:`" and the prototype used an `x-merge` keyword inside the JSON + Schema. rc.5 uses `collection.merge`, a map from field to strategy, + because mdbase semantics live outside the JSON Schema payload (Chapter 06) + and multi-type composition already works there. `x-merge` has no portable + meaning. +2. **The merge unit is the top-level key.** `collection.merge` keys, `add`, + and `remove` address top-level fields only. A nested object merges as one + value. +3. **Which types supply strategies.** The types that the first version matches + at its path. +4. **Default precedence.** A declaration wins; otherwise `max` for a field any + matched type's lifecycle assigns with `now` or `today`, then `union` for + `tags` or a top-level `uniqueItems` array in any matched type, then + `conflict`. Different declarations for one field across matched types are + a `type_conflict`. +5. **The value at a conflict.** The merged version holds the first + (earlier-ordered) value. The second value must not be dropped silently; + how it is kept and shown is implementation behavior. +6. **Ordering for `max` and `min`.** Numbers numerically, RFC 3339 date-times + as instants, full dates in calendar order, other strings in code-point + order. A present value beats a removed key under both `max` and `min`. + Incomparable values, including null against a value, are a conflict. +7. **`union` edge cases.** A missing key or null reads as an empty list; a + `tags` string reads as a one-item list; any other non-list is a conflict. + An empty result where one side removed the key leaves the key missing. +8. **Body alignment.** rc.5 does not pin which longest common + subsequence aligns lines when several exist; implementations SHOULD use + Myers' algorithm, and the fixtures avoid ambiguous alignments. +9. **Append-append scope.** Only appends at the end of the body union; two + insertions at the same interior position are a conflict. When the first + append lacks a final line break, a `\n` is inserted. +10. **Paths and non-mapping frontmatter in a merge.** A path merges with the + `conflict` strategy (one move wins over none; two moves conflict). + Frontmatter that is not a mapping merges as one unit. + +11. **Comments between entries come from the first version.** The merged + frontmatter starts from the first version's source, so comment or blank + lines between entries that only the second version changed are lost + unless the first version is unchanged from the base. + +### Validation and uniqueness + +12. **Warnings on successful writes.** Cross-record issues on a successful + write are `warning` diagnostics at every level, so the result is + `valid: true`; `read` and `validate` report them at the level's severity. +13. **Data contract view failures are single-record**, although a view built + from projections can depend on other records or the time. +14. **`enforce: write` applies at every validation level, including `off`.** +15. **Default uniqueness scope is `type`.** A rule governs records of its + declaring type; `scope` selects the comparison set. Raw values compare by + deep JSON equality without coercion. +16. **`enforce: write` and existing duplicates.** Only writes that introduce a + duplicate fail; edits that leave covered values unchanged succeed. An + implementation that cannot order writes for a collection must reject + writes to covered fields. + +### Paths + +17. **Case folding is full Unicode default folding** (`ß` collides with `ss`), + computed as NFC, fold, NFC. This flags more collisions than some file + systems would, never fewer. +18. **Order without an order.** When nothing orders two colliding records, the + smaller path in code-point order keeps the path. Suffixes go before the + final extension, and existing suffixes are not parsed. +19. **Discovered collisions are not renamed by readers.** Core Read keeps + both files as records and reports `path_collision`; the suffix rule + applies when a tool creates records or copies them onto another file + system. +20. **Path-pattern values** are also rejected when empty, when they begin with + `.` (which would hide the file), or when they are lists or objects. Values + are not trimmed. + +### Matching and expressions + +21. **The bindings `match.expr` should not use.** Besides `now()` and + `today()`, the list includes `file.mtime`, `file.ctime`, and helpers + that read other records (`asFile()`, `file.backlinks`, + `file.hasLink()`). +22. **`match.where` stays.** It is a structured predicate in YAML, not an + expression language, so it does not conflict with SC8. Whether to retire + it in favor of `match.expr` is left open. + +### Identity and moves + +23. **Move detection parameters.** The window is at least 5 seconds, and the + 0.5 and 0.8 similarity thresholds come from the prototype. They may be + tuned. +24. **The identity hint is `settings.id_field`**, rather than a new setting. + Equal non-empty string values pair; different values never pair. +25. **Tie-breaking.** Candidate pairs are chosen greedily by rule rank, + similarity, then the two paths in code-point order. + +### Regular expressions + +26. **JSON Schema `pattern` follows the mdbase regex profile**, not the + ECMA-262 dialect that JSON Schema recommends, so that one engine serves + schema validation, matching, and CEL. The two differ on `\s` and on + look-around and backreferences. +27. **The supported syntax is pinned to `regex-lite`.** The list in Chapter + 10 names the constructs fixtures rely on; anything `regex-lite` rejects is + invalid. A later candidate may enumerate the grammar fully. + +### Body edits + +28. **Offsets count Unicode scalar values**, the unit regex `.` matches. + Byte offsets could split a character; UTF-16 code units are specific to + JavaScript editors, which convert at the adapter boundary. +29. **Offsets refer to the base body**, edits are sorted and non-overlapping, + and two insertions at one offset are invalid rather than ordered. +30. **`body_base` is a SHA-256 digest of the body bytes**, defined by the spec + rather than the opaque `revision`, so a client can compute it without the + engine. +31. **Conflicts fail the update.** A rebase that conflicts fails with + `concurrent_modification` (`body_conflict`) and writes nothing, rather than + writing the current body and holding the edit; the caller has the base and + can retry. +32. **Retaining bases is optional.** A stale base without `body_base_text` and + without a retained copy fails with `body_base_unavailable`. + +### Writers and links + +33. **Format fidelity covers YAML document records too**, replacing the rc.4 + text that let structured writes to `.base` files re-emit the mapping. + Keeping a scalar's quoting and an entry's trailing comment is SHOULD; the + rest is MUST. +34. **New diagnostic codes** `link_target_type_mismatch` (previously + unnamed), `path_collision`, `path_value_invalid`, and + `nondeterministic_match`. Filename tiebreakers + are now required. +35. **`if_revision` is opt-in.** Transports and SDKs must not add it on a + caller's behalf; this is stated as an operation contract, not as a + replication rule. + +## Not in this release + +- Replication, logs, sequencers, ordering services, and conflict envelopes. +- Delete-versus-update policy. A merge of two versions is defined; what + happens when one side deleted the record is implementation behavior. +- Schema epochs for records written under an older type, and revalidation + after a type change. +- Rename alias resolution for links written concurrently with a rename. From e6fb2defbda0fa34ac5790a5a90764f6465c45bf Mon Sep 17 00:00:00 2001 From: callumalpass Date: Sun, 4 Oct 2026 15:12:25 +1100 Subject: [PATCH 5/5] Pin filename tiebreakers and rename reference updating for rc.5 Reconciles rc.5 with the open link PRs #49 and #51. - Chapter 08: the second tiebreaker counts path segments, and no tool may choose among duplicate filename matches by scan, index, or storage order. A tiebreaker-resolved link is not ambiguous. - Chapter 12: a rename with reference updating rewrites the links that resolved to the renamed record, including tiebreaker-selected filename matches, and never rewrites ambiguous links or links to other records. The result lists rewrites in references_updated. A rename keeps the record's identity, as a detected move does in Chapter 12A. - Chapter 14, Chapter 13, the changelog, and the release notes follow. - Conformance: links.filename_tiebreakers and core_write.rename_references, with four tiebreaker tests in core/link-ambiguity.yaml and four rename tests in the new core/rename-references.yaml (155 rc.5 tests in all). --- 08-links.md | 11 +- 12-operations.md | 23 ++++- 13-migrations-and-compatibility.md | 5 +- 14-conformance.md | 6 ++ CHANGELOG.md | 8 +- docs/releases/0.3.0-rc.5.md | 27 ++++- tests/v0.3/README.md | 9 ++ tests/v0.3/core/link-ambiguity.yaml | 134 ++++++++++++++++++++++++- tests/v0.3/core/rename-references.yaml | 121 ++++++++++++++++++++++ tests/v0.3/manifest.yaml | 3 + 10 files changed, 333 insertions(+), 14 deletions(-) create mode 100644 tests/v0.3/core/rename-references.yaml diff --git a/08-links.md b/08-links.md index b0dc2cd..bae639e 100644 --- a/08-links.md +++ b/08-links.md @@ -62,12 +62,19 @@ If filename resolution finds multiple candidates, tools MUST apply these tiebreakers in order: 1. same directory as referring file -2. shortest collection path +2. fewest path segments, that is, the shallowest collection path 3. smallest path in Unicode code-point order +Tools MUST NOT choose among candidates by any other criterion, such as +discovery or scan order, the order in which files were indexed, or a +provider's storage order, so every conforming tool resolves a simple link to +the same record. The path segments of `docs/readme.md` are `docs` and +`readme.md`; its depth is 2 whatever the length of its names. + Filename candidates whose paths are equivalent under Chapter 02 path keys are ambiguous with each other even after the tiebreakers. If ambiguity remains, -the link resolves to null. +the link resolves to null. A link that the tiebreakers resolve is not +ambiguous and reports no `ambiguous_link`. An ambiguous link reports an `ambiguous_link` cross-record issue on the referring record, with `details.candidates` listing the candidate paths in diff --git a/12-operations.md b/12-operations.md index ef003bc..a9d83d9 100644 --- a/12-operations.md +++ b/12-operations.md @@ -274,9 +274,26 @@ path key equals that of a different existing record fails with `path_conflict`. A target whose path key equals the source's own path key only changes the path's spelling, such as its case, and is not a conflict. -If reference updating is enabled, link updates SHOULD preserve link style, -alias, and anchor where possible. ID-based links SHOULD not be rewritten if the -target ID did not change. +If reference updating is enabled, the rename updates the links that +reference the renamed record: the links that, before the rename, resolve to it +under Chapter 08, including a simple link whose filename match was selected by +the tiebreakers. Each such link that would no longer resolve to the record +after the move MUST be rewritten so that it resolves to the record at its new +path. An ID-based link whose target ID did not change still resolves and +SHOULD NOT be rewritten. A rename MUST NOT rewrite links that resolve to +another record, are unresolved, or are ambiguous; in particular, a link whose +configured ID is ambiguous resolves to no record and is not rewritten, even +when it matches the renamed record's filename. Link updates SHOULD preserve +link style, alias, and anchor where possible. + +The rename result lists each rewritten link in `references_updated`, with the +referring record's `path`, the `field` that held the link (or +`location: body`), `old_value`, and `new_value`, ordered by referring path. + +A rename keeps the record's identity: internal identity, history, and merge +bases follow the record to its new path, as for a move detected under Chapter +12A, and a tool that reports watch notifications reports it as +`record_renamed`. An implementation MAY commit a rename and its reference updates as one atomic batch. Otherwise reference updates are applied after the rename, and failed diff --git a/13-migrations-and-compatibility.md b/13-migrations-and-compatibility.md index 8b6b099..9c010ff 100644 --- a/13-migrations-and-compatibility.md +++ b/13-migrations-and-compatibility.md @@ -35,7 +35,8 @@ worked under rc.4 stops loading. | derived path already taken | `path_conflict` | the first free suffixed path | supply an explicit path to get an error instead | | structured writes to Markdown frontmatter | only array and object structure SHOULD be preserved | MUST re-emit only changed entries | nothing | | structured writes to YAML document records | could re-emit the whole mapping | follow writer format fidelity | nothing | -| filename link tiebreakers | SHOULD | MUST, ending in code-point order | nothing | +| filename link tiebreakers | SHOULD; "shortest path" unmeasured | MUST: referring directory, then fewest path segments, then code-point order; never scan or storage order | nothing | +| rename with reference updating | which links are rewritten was unspecified | only links that resolved to the renamed record, including tiebreaker-selected matches; ambiguous links and links to other records unchanged; listed in `references_updated` | nothing | ### Tightenings reported as diagnostics @@ -76,6 +77,8 @@ An engine moving from rc.4 to rc.5: claims the `merge` profile - implements move detection if it reports renames of externally moved files - reports `nondeterministic_match`, `path_collision`, and `ambiguous_link` +- applies the filename tiebreakers exactly, and rewrites on rename only the + links that resolved to the renamed record - removes any `id` default for `settings.id_field` Engines that conform to rc.4 keep claiming `0.3.0-rc.4` until they pass the diff --git a/14-conformance.md b/14-conformance.md index 5fca9a0..e75003f 100644 --- a/14-conformance.md +++ b/14-conformance.md @@ -360,6 +360,8 @@ Links implementations MUST: `settings.id_field` is configured - resolve ambiguous IDs to null without filename fallback, and report `ambiguous_link` +- choose among duplicate filename matches only with the Chapter 08 + tiebreakers, ending in code-point order - expose `file.links`, `file.embeds`, `file.tags`, and `file.backlinks` - provide the CEL link helpers from Chapter 10 - bound `asFile()` traversal @@ -377,6 +379,10 @@ Core Write implementations MUST: with the Chapter 12A body merge when the base body is available, and report `body_conflict` and `body_base_unavailable` otherwise - reject paths that escape the collection root +- when a rename updates references, rewrite the links that resolved to the + renamed record and would no longer resolve to it, including + tiebreaker-selected filename matches, and never rewrite ambiguous links or + links to other records - reject an explicit path that collides under path equivalence with `path_conflict`, give a colliding derived path the first free suffix, and reject invalid path-pattern values with `path_value_invalid` diff --git a/CHANGELOG.md b/CHANGELOG.md index 5220515..4449ccf 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -50,7 +50,7 @@ Bases, seed type upgrades, and saved-view identification. - Diagnostic codes `path_collision`, `path_value_invalid`, `ambiguous_link`, `link_target_type_mismatch`, `nondeterministic_match`, `invalid_pattern`, and `duplicate_value` are listed as core codes. -- The v0.3 suite gains 147 tests marked `since: 0.3.0-rc.5`, including the 13 +- The v0.3 suite gains 155 tests marked `since: 0.3.0-rc.5`, including the 13 merge fixtures of the mdbase-next prototype in a new `merge_records` format and new `merge` and `watch` fixture sets. `scripts/check_v03_tests.py` runs the 95 pure-function fixtures against an executable model @@ -79,8 +79,12 @@ Bases, seed type upgrades, and saved-view identification. - `collection.path.pattern` values containing `/` or `\`, beginning with `.`, empty, or not scalars fail with `path_value_invalid`. - An ambiguous configured ID resolves to null with `ambiguous_link` and no - filename fallback. Filename tiebreakers are required and end in code-point + filename fallback. Filename tiebreakers are required: referring directory, + then fewest path segments, then code-point order, and never scan or storage order. +- A rename with reference updating rewrites only the links that resolved + to the renamed record, including tiebreaker-selected filename matches, and + reports them in `references_updated`; ambiguous links are not rewritten. - `if_revision` is opt-in; transports and SDKs must not add it on a caller's behalf. - Lifecycle states that it has no sequence provider; implementations may diff --git a/docs/releases/0.3.0-rc.5.md b/docs/releases/0.3.0-rc.5.md index 6c0a5e0..e571d56 100644 --- a/docs/releases/0.3.0-rc.5.md +++ b/docs/releases/0.3.0-rc.5.md @@ -55,6 +55,14 @@ not contain `/`, and an ambiguous ID resolves to null with `ambiguous_link` and no filename fallback. rc.4 already said all three; engines did not follow it. +Duplicate filenames resolve deterministically. The filename tiebreakers +(referring directory, then fewest path segments, then code-point order) are +required, and no tool may choose by scan or storage order, so every +conforming tool binds a simple link to the same record. A rename with +reference updating rewrites only the links that resolved to the renamed +record, including tiebreaker-selected matches, and leaves ambiguous links +alone. + ## Behavior changes since rc.4 - **Uniqueness reports by default.** At level `error`, rc.4 rejected a write @@ -103,7 +111,7 @@ no change, and the engine checklist. ## Conformance The rc.5 additions live in the v0.3 suite (`tests/v0.3/`) and carry -`since: 0.3.0-rc.5`. There are 147 new tests; 95 are pure functions that CI +`since: 0.3.0-rc.5`. There are 155 new tests; 95 are pure functions that CI executes against an executable model (`scripts/concurrent_edits_model.py`) through `scripts/check_v03_tests.py`: @@ -119,9 +127,11 @@ through `scripts/check_v03_tests.py`: - `merge/body-edits.yaml`: 12 pure `apply_body_edits` cases (direct, rebased, conflict, base unavailable, invalid ranges); `core/body-edits-update.yaml` covers `update` with `body_edits` +- `core/link-ambiguity.yaml`: ID ambiguity, the `id_field` default, and the + filename tiebreakers; `core/rename-references.yaml`: which links a rename + rewrites - adapter-target suites `core/validation-tiers.yaml`, - `core/list-ops-and-fidelity.yaml`, `core/link-ambiguity.yaml`, and - `cel/match-determinism.yaml` + `core/list-ops-and-fidelity.yaml`, and `cel/match-determinism.yaml` `tests/v0.3/README.md` documents the `since`/`changed` markers and the new operations, including the `merge_records` format. @@ -157,6 +167,11 @@ engine fails them even though rc.4 had no test for the case: | `paths.derived_slash_rejected`, `paths.derive_slash` | create a subfolder | | `links.no_default_id_field` | some engines default `id_field` to `id` | | `links.ambiguous_id_validate` | no `ambiguous_link` diagnostic | +| `links.tiebreak_depth` | an engine that measured "shortest path" in characters picks `a/b/target.md` | +| `links.tiebreak_code_point` | an engine that sorted alphabetically ignoring case, or by locale, picks `alpha/target.md` | +| `links.tiebreak_path_key_ambiguous` | resolves to one of the case-different paths | +| `core_write.rename_tiebreak_winner`, `core_write.rename_tiebreak_same_directory` | treat any duplicate filename as ambiguous and rewrite nothing | +| `core_write.rename_no_filename_fallback` | rewrites the link after falling back from an ambiguous ID to the filename | | `cel_match.nondeterministic_match_warning` | no diagnostic | | `core_write.patch_minimal_bytes`, `core_write.new_key_appended`, `core_write.yaml_document_fidelity`, `core_write.add_flow_style`, `core_write.add_block_style` | may re-emit the frontmatter | | every `core_write` list-operation test | `add` and `remove` are unknown request members | @@ -324,7 +339,7 @@ declared stable. 34. **New diagnostic codes** `link_target_type_mismatch` (previously unnamed), `path_collision`, `path_value_invalid`, and `nondeterministic_match`. Filename tiebreakers - are now required. + are now required, with depth measured in path segments. 35. **`if_revision` is opt-in.** Transports and SDKs must not add it on a caller's behalf; this is stated as an operation contract, not as a replication rule. @@ -337,3 +352,7 @@ declared stable. - Schema epochs for records written under an older type, and revalidation after a type change. - Rename alias resolution for links written concurrently with a rename. +- Protecting links whose tiebreaker outcome changes because a create or + rename adds a closer filename candidate. Such a link follows the + tiebreakers to the new record; authors who need a fixed target use a path + or a unique `id_field`. diff --git a/tests/v0.3/README.md b/tests/v0.3/README.md index 460cc0d..af292ba 100644 --- a/tests/v0.3/README.md +++ b/tests/v0.3/README.md @@ -96,6 +96,7 @@ Future v0.3 adapters should support these operations: - `evaluate_cel` - `create` - `update` +- `rename` - `runtime_load_contracts` - `runtime_compose_registry` - `runtime_preflight_workflows` @@ -166,6 +167,14 @@ diagnostics. Tests that create two files differing only in case, such as `paths.discovered_collision`, need a case-sensitive file system; an adapter on a case-insensitive one reports them as skipped. +### `rename` with `update_refs` (since rc.5) + +`core/rename-references.yaml` renames with `input.update_refs: true` and +expects `references_updated`: every rewritten link, ordered by referring path, +with `path`, `field` (or `location: body`), `old_value`, and `new_value`. A +record the list does not name is unchanged, so an empty list asserts that no +link was rewritten. + ### `merge_records` The extension for merge cases: base + two edits + declarations -> expected diff --git a/tests/v0.3/core/link-ambiguity.yaml b/tests/v0.3/core/link-ambiguity.yaml index cf35802..986061e 100644 --- a/tests/v0.3/core/link-ambiguity.yaml +++ b/tests/v0.3/core/link-ambiguity.yaml @@ -1,6 +1,6 @@ # Link clarifications in 0.3.0-rc.5 (Chapters 04 and 08): settings.id_field has -# no default, and an ambiguous ID resolves to null with ambiguous_link and no -# filename fallback. +# no default, an ambiguous ID resolves to null with ambiguous_link and no +# filename fallback, and the filename tiebreakers are required. name: "link resolution clarifications" spec_version: "0.3.0" fixture_set: core_collection @@ -82,3 +82,133 @@ groups: - code: ambiguous_link severity: warning field: parent + + # Filename tiebreakers (Chapter 08) are required in 0.3.0-rc.5. They order + # by referring directory, then path depth in segments, then code-point + # order, never by scan or storage order. + - name: "same directory wins over a shallower path" + setup: + config: | + spec_version: "0.3.0" + files: + target.md: | + --- + title: Root target + --- + tasks/target.md: | + --- + title: Same-directory target + --- + tasks/source.md: | + --- + ref: "[[target]]" + --- + tests: + - name: "a candidate in the referring directory wins" + id: links.tiebreak_same_directory + since: 0.3.0-rc.5 + covers: [links.filename_tiebreakers] + operation: resolve_link + input: + path: "tasks/source.md" + field: ref + expect: + valid: true + resolved: "tasks/target.md" + diagnostics: [] + + - name: "depth counts path segments" + setup: + config: | + spec_version: "0.3.0" + files: + a/b/target.md: | + --- + title: Deeper but shorter path + --- + a-much-longer-folder-name/target.md: | + --- + title: Shallower but longer path + --- + source.md: | + --- + ref: "[[target]]" + --- + tests: + - name: "the shallowest candidate wins, whatever the length of its names" + id: links.tiebreak_depth + since: 0.3.0-rc.5 + covers: [links.filename_tiebreakers] + operation: resolve_link + input: + path: "source.md" + field: ref + expect: + valid: true + resolved: "a-much-longer-folder-name/target.md" + diagnostics: [] + + - name: "code-point order is the final tiebreaker" + setup: + config: | + spec_version: "0.3.0" + files: + alpha/target.md: | + --- + title: Lowercase folder + --- + Beta/target.md: | + --- + title: Uppercase folder + --- + source.md: | + --- + ref: "[[target]]" + --- + tests: + - name: "equal-depth candidates resolve to the smallest path in code-point order" + id: links.tiebreak_code_point + since: 0.3.0-rc.5 + covers: [links.filename_tiebreakers] + operation: resolve_link + input: + path: "source.md" + field: ref + expect: + valid: true + resolved: "Beta/target.md" + diagnostics: [] + + - name: "path-key-equivalent candidates" + setup: + config: | + spec_version: "0.3.0" + files: + notes/plan.md: | + --- + title: lower + --- + Notes/plan.md: | + --- + title: upper + --- + tasks/t.md: | + --- + ref: "[[plan]]" + --- + tests: + - name: "candidates with one path key stay ambiguous after the tiebreakers" + id: links.tiebreak_path_key_ambiguous + since: 0.3.0-rc.5 + covers: [links.filename_tiebreakers] + operation: resolve_link + input: + path: "tasks/t.md" + field: ref + expect: + valid: true + resolved: null + diagnostics: + - code: ambiguous_link + field: ref + details: { candidates: ["Notes/plan.md", "notes/plan.md"] } diff --git a/tests/v0.3/core/rename-references.yaml b/tests/v0.3/core/rename-references.yaml new file mode 100644 index 0000000..3d1e5df --- /dev/null +++ b/tests/v0.3/core/rename-references.yaml @@ -0,0 +1,121 @@ +# Rename reference updating in 0.3.0-rc.5 (Chapter 12): a rename rewrites +# only the links that resolved to the renamed record under Chapter 08, +# including tiebreaker-selected filename matches, and leaves ambiguous links +# unchanged. `references_updated` lists every rewritten link; records it does +# not list are unchanged. +name: "rename reference updating" +spec_version: "0.3.0" +fixture_set: core_collection +category: core_write +spec_ref: "v0.3/08, v0.3/12" + +groups: + - name: "duplicate filenames resolved by the tiebreakers" + setup: + config: | + spec_version: "0.3.0" + files: + folder-a/shared-name.md: | + --- + title: A + --- + folder-b/shared-name.md: | + --- + title: B + --- + notes/source.md: | + --- + ref: "[[shared-name]]" + --- + folder-b/local.md: | + --- + ref: "[[shared-name]]" + --- + tests: + - name: "a link selected by code-point order is rewritten" + id: core_write.rename_tiebreak_winner + since: 0.3.0-rc.5 + covers: [core_write.rename_references, links.filename_tiebreakers] + operation: rename + input: + from: "folder-a/shared-name.md" + to: "folder-a/renamed.md" + update_refs: true + expect: + valid: true + path: "folder-a/renamed.md" + references_updated: + - path: "notes/source.md" + field: ref + old_value: "[[shared-name]]" + new_value: "[[renamed]]" + + - name: "a link selected by the referring directory is rewritten, and only that link" + id: core_write.rename_tiebreak_same_directory + since: 0.3.0-rc.5 + covers: [core_write.rename_references, links.filename_tiebreakers] + operation: rename + input: + from: "folder-b/shared-name.md" + to: "folder-b/other.md" + update_refs: true + expect: + valid: true + path: "folder-b/other.md" + references_updated: + - path: "folder-b/local.md" + field: ref + old_value: "[[shared-name]]" + new_value: "[[other]]" + + - name: "ambiguous configured ids" + setup: + config: | + spec_version: "0.3.0" + settings: + id_field: id + files: + tasks/a.md: | + --- + id: shared + --- + tasks/b.md: | + --- + id: shared + --- + tasks/shared.md: | + --- + title: Filename match + --- + tasks/child.md: | + --- + parent: "[[shared]]" + --- + tests: + - name: "renaming one record with an ambiguous id rewrites no link" + id: core_write.rename_ambiguous_id_unchanged + since: 0.3.0-rc.5 + covers: [core_write.rename_references, links.ambiguous_id_link] + operation: rename + input: + from: "tasks/a.md" + to: "tasks/a2.md" + update_refs: true + expect: + valid: true + path: "tasks/a2.md" + references_updated: [] + + - name: "renaming the filename match of an ambiguous id rewrites no link" + id: core_write.rename_no_filename_fallback + since: 0.3.0-rc.5 + covers: [core_write.rename_references, links.ambiguous_id_link] + operation: rename + input: + from: "tasks/shared.md" + to: "tasks/renamed.md" + update_refs: true + expect: + valid: true + path: "tasks/renamed.md" + references_updated: [] diff --git a/tests/v0.3/manifest.yaml b/tests/v0.3/manifest.yaml index d35d748..41689c3 100644 --- a/tests/v0.3/manifest.yaml +++ b/tests/v0.3/manifest.yaml @@ -103,6 +103,7 @@ claim_profiles: - backlinks - id_field_no_default # since 0.3.0-rc.5 - ambiguous_id_link # since 0.3.0-rc.5 + - filename_tiebreakers # since 0.3.0-rc.5 - id: core_write status: draft requires: [collection_semantics] @@ -115,6 +116,7 @@ claim_profiles: - writer_format_fidelity # since 0.3.0-rc.5 - path_collisions # since 0.3.0-rc.5 - path_pattern_values # since 0.3.0-rc.5 + - rename_references # since 0.3.0-rc.5 - body_edits # since 0.3.0-rc.5 - id: merge # since 0.3.0-rc.5 status: coverage_complete @@ -225,6 +227,7 @@ fixture_sets: - core/list-ops-and-fidelity.yaml - core/paths.yaml - core/link-ambiguity.yaml + - core/rename-references.yaml - core/body-edits-update.yaml - id: data_contracts description: Data contracts, type implementations, projections, stable digests, and conflicts.