Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 63 additions & 2 deletions 02-collection-layout.md
Original file line number Diff line number Diff line change
Expand Up @@ -148,8 +148,55 @@ This includes `..` traversal, symlink traversal where the implementation follows
symlinks, and absolute paths supplied where a collection-relative path is
required.

Implementations MAY reject platform-reserved filenames or characters when a
write operation targets a filesystem where those paths cannot be represented.
### Portable Paths

A collection is read and written on Linux, macOS, and Windows file systems,
and inside applications such as Obsidian. A path that one platform reads
differently from another can escape the collection, reach a tool's private
or executable state, or fail to be written at all. Every tool MUST therefore
reject, for every record, resource, and file it creates, renames, moves,
replicates, or applies from another tool, a collection-relative path that:

- is empty, or longer than 1,024 bytes of UTF-8;
- begins with `/` (an absolute or UNC path), or contains `\`;
- has an empty segment (`a//b`, a trailing `/`), a `.` or `..` segment, or a
segment longer than 255 bytes;
- contains any of `<`, `>`, `:`, `"`, `|`, `?`, `*` (`:` also covers drive
prefixes such as `C:x.md` and alternate data streams);
- contains a control character (U+0000–U+001F, U+007F–U+009F);
- contains a character that is invisible or that HFS+ ignores in names:
U+00AD, U+034F, U+115F, U+1160, U+17B4, U+17B5, U+180B–U+180F,
U+200B–U+200F, U+202A–U+202E, U+2060–U+206F, U+3164, U+FE00–U+FE0F,
U+FEFF, U+FFA0, U+FFF0–U+FFF8, U+1BCA0–U+1BCA3, U+1D173–U+1D17A, and
U+E0000–U+E0FFF (so `\u200C.git` cannot name `.git`);
- has a segment ending in `.` or a space, which Windows strips;
- has a segment that is a Windows device name, compared without ASCII case
on the part before its first `.` with trailing spaces removed: `CON`,
`PRN`, `AUX`, `NUL`, `CONIN$`, `CONOUT$`, `CLOCK$`, and `COM` or `LPT`
followed by one of `0`–`9`, `¹`, `²`, `³`;
- has a segment shaped like an NTFS 8.3 short name, which can alias another
name such as `.git`: one to six characters, `~`, a decimal number without a
leading zero, and optionally `.` and up to three characters (`GIT~1`,
`MDBASE~2.TXT`);
- has a segment beginning with `.`: hidden files and tool state such as
`.obsidian`, `.git`, `.vscode`, and `.mdbase` are never collection content
(see Record Discovery);
- has a segment equal to `node_modules`.

The names `.mdbase` and `node_modules` compare without ASCII case, with U+017F
LATIN SMALL LETTER LONG S read as `s` and U+212A KELVIN SIGN as `k`: the only
characters whose case folding gives a letter of those names. This comparison
is fixed and does not depend on a Unicode version, so the rule gives the same
verdict in every release.

A tool that reads a collection skips files at such paths as it skips
excluded paths. A write that would create one fails with `invalid_request`
(`details.reason` names the rule) before any write.

**Provisional (rc.5).** The rule set comes from the
review of the first implementation. Earlier drafts let implementations
reject such paths optionally; replicating a collection across file systems
needs every tool to agree on them.

## Path Equivalence

Expand Down Expand Up @@ -181,12 +228,26 @@ collides, whatever file system it runs on:
Read reports a `path_collision` warning on every record of the group, with
`details.paths` listing the group in code-point order.

**Provisional (rc.5).** A tool that keeps records in a log or replicates
them across file systems cannot hold two records with one path key. It
MAY resolve a discovered group instead, by the collision rule below: with
no other order, the path that is smaller in code-point order keeps it, and
each other record moves to its first free suffixed path. It reports the
moves as `record_renamed` and no `path_collision` warning remains.

Path globs (above) remain case-sensitive and match paths as written.

**Provisional (rc.5).** Case folding uses the full mappings, so `ß` and
`ss` collide. This flags more collisions than some file systems would, never
fewer.

**Provisional (rc.5).** NFC and case folding use the data of **Unicode
17.0.0** (`UnicodeData.txt`, `CompositionExclusions.txt`, `CaseFolding.txt`).
Unicode's stability policies keep both stable for assigned characters, but a
character assigned in a later version can fold or compose differently once a
tool knows it. Pinning the version makes every tool compute the same path
keys. A later release names its own version.

## Path Collisions

When a new record would take a path whose path key is already in use, the
Expand Down
28 changes: 28 additions & 0 deletions 03-records-and-frontmatter.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,3 +210,31 @@ data model before JSON Schema validation. Non-JSON YAML values such as NaN,
Infinity, binary values, and timestamps with parser-specific objects MUST be
handled by the mdbase YAML profile before schema validation or rejected with a
clear diagnostic.

### Scalar Resolution

Plain (unquoted) scalars resolve with the YAML 1.2 core schema. A tool MUST
read a plain scalar as follows, and MUST NOT apply YAML 1.1 resolution:

| Plain scalar | Value |
| --- | --- |
| empty, `~`, `null`, `Null`, `NULL` | null |
| `true`, `True`, `TRUE`, `false`, `False`, `FALSE` | boolean |
| `[-+]?[0-9]+`, `0o[0-7]+`, `0x[0-9a-fA-F]+` | integer |
| `[-+]?(\.[0-9]+\|[0-9]+(\.[0-9]*)?)([eE][-+]?[0-9]+)?` | number |
| anything else | string |

So `yes`, `no`, `on`, and `off` are strings, `0777` is the integer 777,
`1_000` and `1:20` are strings, timestamps such as `2026-10-01` and
`2026-10-01T09:00:00Z` are strings, and `<<` is an ordinary key. Quoted and
block scalars are always strings. `.inf`, `.nan`, and decimal numbers whose
value is not a finite double are outside the JSON data model and are read as
the string as written.

A writer that has no existing style to keep SHOULD quote a string that a YAML
1.1 parser would read as another type (for example `"yes"` or
`"2026-10-01"`), so tools that still use YAML 1.1 read the same value.

**Provisional (rc.5).** This table resolves an ambiguity of earlier release
candidates, which asked for "a safe YAML parser" without naming a schema. It
matches Obsidian and the current engines.
6 changes: 6 additions & 0 deletions 07-collection-semantics.md
Original file line number Diff line number Diff line change
Expand Up @@ -333,6 +333,12 @@ fields. Values are converted to strings without expression evaluation: a
string is used as written, and a number or boolean uses its JSON
representation. A missing or null value produces `path_value_missing`.

**Provisional (rc.5).** The JSON representation of a number is the one RFC
8785 (JSON Canonicalization Scheme) specifies, which is ECMAScript's
`Number.prototype.toString`: an integer-valued number has no fraction
(`42.0` gives `42`), and exponents are written `1e+21` and `1e-7`. CEL's
`string(double)` uses the same text (Chapter 10).

A placeholder value always stays within one path component. A converted
value is invalid, and the operation fails with `path_value_invalid` naming the
field, when it:
Expand Down
37 changes: 35 additions & 2 deletions 10-cel-profile.md
Original file line number Diff line number Diff line change
Expand Up @@ -175,6 +175,13 @@ The durable runtime profile defines when each workflow expression is evaluated.

## Record Values

**Provisional (rc.5).** CEL maps are unordered, but comprehensions over a
map (`all`, `exists`, `exists_one`, `map`, `filter`) visit its keys, and
`map` and `filter` return lists in that order. Maps iterate in insertion
order: frontmatter maps in the order of their keys in the source, map
literals in the order written. `string(double)` writes the RFC 8785 number
text of Chapter 07 (`1e+21`, `0.3333333333333333`, and `0` for negative zero).

Frontmatter is converted to CEL values after the JSON data-model conversion in
Chapter 06:

Expand Down Expand Up @@ -291,7 +298,9 @@ profile adds two string methods:
| `s.lower()` | `s` with every character mapped to lowercase |
| `s.upper()` | `s` with every character mapped to uppercase |

Both use the Unicode default full case mappings without locale tailoring, so
Both use the Unicode default full case mappings without locale tailoring,
including the context-dependent final sigma rule, from the same Unicode
version as path keys (Chapter 02, provisionally 17.0.0), so
`"Éclair".lower()` is `"éclair"` and `"Straße".upper()` is `"STRASSE"`. A
case-insensitive search lowercases the text it searches:

Expand Down Expand Up @@ -330,7 +339,31 @@ profile keeps that syntax and removes Unicode classes from it.
- the flags `i`, `m`, `s`, and `x`, set as `(?flags)` or `(?flags:...)`

A pattern that uses anything else is invalid. In particular, Unicode classes
such as `\p{L}` and `\pN`, backreferences, and look-around are invalid.
such as `\p{L}` and `\pN`, backreferences, and look-around are invalid, and
so are constructs that `regex-lite` accepts beyond this list: `\b{start}` and
the other `\b{...}` boundaries, `(?<name>...)` (use `(?P<name>...)`), the
`U` and `R` flags, escapes of letters or digits other than those listed (such
as `\a` and `\u0041`), nested bracket classes, and class set operations
(`&&`, `--`, `~~`). A `{` that does not begin a valid repetition, and a `]` or
`}` outside a class, are invalid too; escape them. `\<` and `\>` are
literal `<` and `>`, as the escape rule above says, although `regex-lite`
itself reads them as word boundaries.

**Limits.** Every tool accepts and rejects the same patterns:

- a pattern longer than 8,192 bytes is invalid;
- groups nested deeper than 64 are invalid;
- a repetition count above 1,000 in `{n}`, `{n,}`, or `{n,m}` is invalid, as
in RE2;
- a pattern whose **program size** exceeds 100,000 is invalid. The size is: 1
for a literal, `.`, a bracket class, or a Perl class; 0 for an anchor or
boundary; the sum for a concatenation; the sum plus 1 per `|` for an
alternation; the inner size plus 1 for a group; twice the operand for `*`,
`+`, and `?`; and the operand times `max(n, m, 1) + 1` for `{n}`, `{n,}`,
and `{n,m}`.

`regex-lite`'s own size limit depends on the platform's pointer width, so it
cannot be the portable limit.

**Semantics.** Matching runs over Unicode scalar values, and a match is
unanchored unless the pattern anchors it, as in RE2 and JSON Schema. Among
Expand Down
67 changes: 53 additions & 14 deletions 12a-concurrent-edits.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,11 @@ The result is a **merged version** and a possibly empty list of
**conflicts**. A conflict has a `kind`: `field` for one top-level frontmatter
field, which it names in `field`; `frontmatter` for a whole frontmatter
block; `body`; or `path`. It carries the base, first, and second values.

Conflicts are listed in this order: field conflicts in key order (the first
version's keys in its order, then keys only the second version has, in its
order, then keys only the base has, in its order), then a `frontmatter`
conflict, then a `body` conflict, then a `path` conflict.
Where a conflict exists, the merged version holds the first version's value.
A tool that writes a merged version with conflicts MUST NOT discard the second
version's conflicting values silently: how it keeps and surfaces them is
Expand Down Expand Up @@ -159,7 +164,10 @@ standing for that key's state in each:
3. Otherwise, if `F` equals `B`, the result is `S`: only the second side
changed the key.
4. Otherwise both sides changed the key differently, and the key's strategy
decides.
decides. When the matched types declare different strategies for the key
(a `type_conflict`, Chapter 05), the key uses `conflict`: a merge never
fails and never drops a value it cannot decide. Reading the merged record
still reports the `type_conflict`.

| Strategy | Result when both sides changed the key differently |
| --- | --- |
Expand Down Expand Up @@ -223,10 +231,21 @@ standing for the three bodies:
the result is `S`.
2. **Append-append.** If `F` and `S` both begin with all of `B`, so that both
sides only appended text at the end, the result is `B`, then `F`'s
appended text, then `S`'s appended text. When `F`'s appended text is not
empty and does not end with a line terminator, a `\n` is inserted between
the two. Journals, logs, and checklists grow this way, and appending to
them concurrently is not a conflict.
appended text, then `S`'s appended text, joined as follows:
- when `B` is not empty and does not end with a line terminator, and both
appended texts begin with one, the line terminator at the start of
`S`'s text is dropped: `F`'s text already ended `B`'s last line;
- then, when `F`'s appended text is not empty and does not end with a line
terminator, and `S`'s remaining text does not begin with one, a line
terminator is inserted between the two, in the body's line-ending style:
the style of `B`'s first line terminator, or when `B` has none, of the
first line terminator in `F`'s and then `S`'s appended text, and `\n`
when there is none at all. A merge therefore never mixes `\n` and
`\r\n` in a body that used one style.

So no empty line appears between the two appends. Journals, logs, and
checklists grow this way, and appending to them concurrently is not a
conflict.
3. Otherwise the bodies merge as a three-way line merge (diff3). The lines of
`B` that are aligned with unchanged lines in both `F` and `S` divide the
bodies into stable regions and changed chunks. For each changed chunk, the
Expand All @@ -236,13 +255,29 @@ standing for the three bodies:

When the body is in conflict, the merged body is `F` as a whole and the
conflict carries the three bodies. The alignment of `B` with each side is a
longest common subsequence of lines.

**Provisional (rc.5).** When several longest common subsequences exist, this
release candidate does not fix which one is used, so two implementations can
split some changed chunks differently. Implementations SHOULD use the Myers
difference algorithm. Conformance fixtures use bodies whose alignment is
unique.
longest common subsequence of lines, chosen as follows, so that every
implementation splits chunks the same way:

1. Lines common to the start of both sequences are aligned, then lines
common to their end.
2. The remaining middle parts are split at the **middle snake** of Myers'
linear-space algorithm (E. Myers, "An O(ND) Difference Algorithm and Its
Variations", 1986, section 4b), in the formulation of diff-match-patch's
`bisect`: for each edit distance `d` the forward search runs before the
reverse search; on each diagonal a search continues from the neighbouring
diagonal with the larger furthest-reaching value, and from the lower
diagonal (a deletion) when they are equal; the split point is the forward
search's furthest-reaching point on the diagonal where the searches first
overlap.
3. Each part is aligned recursively with these rules.

The executable model (`scripts/concurrent_edits_model.py`, `_lcs_pairs`)
is the reference for this procedure.

**Provisional (rc.5).** Earlier drafts left the choice among several longest
common subsequences open. The procedure above is the one the reference model
and the first engine implement. A later release may name a simpler canonical
alignment if one proves as fast.

**Provisional (rc.5).** Append-append applies only to appends at the end
of the body. Two insertions at the same place inside the body remain a
Expand Down Expand Up @@ -285,8 +320,12 @@ document records alike.

A frontmatter source consists of **top-level entries** and the lines between
them. An entry is a line that begins a top-level key at column 0, together
with every following line that belongs to that key's value. Blank lines and
comment lines at column 0 between entries are not part of any entry.
with every following line up to the last line of that key's value, and then
any directly following blank and indented comment lines up to the last
indented comment line. Lines inside the value belong to the entry even when
they are blank or column-0 comments, such as a comment between two `- ` items
of a block sequence at column 0. Blank lines and comment lines at column 0
after the value are not part of any entry.

1. An entry whose key the write does not change MUST stay byte-identical,
including its comments, quoting, indentation, and position.
Expand Down
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,23 @@ Bases, seed type upgrades, and saved-view identification.

#### Clarified

- Errata from the first implementation (provisional; details in the release
notes):
- YAML 1.2 core scalar resolution;
- which lines belong to a frontmatter entry;
- the pinned Myers body alignment;
- a `type_conflict` during a merge merges as `conflict`;
- conflict order;
- append-append onto a base without a final line break, with separators in
the body's line-ending style;
- portable paths: a normative path-safety rule set (Chapter 02);
- regex `\<`/`\>`, `regex-lite`-only syntax and portable limits;
- Unicode 17.0.0 for path keys and case mappings;
- RFC 8785 number text;
- CEL map iteration order.

The executable model follows each, and all `cel/cel-profile.yaml` tests now
have ids.
- `settings.id_field` has no default; engines that resolve through `id`
without configuration are non-conforming. It also serves as the
move-detection identity hint.
Expand Down
Loading
Loading