Skip to content

LibRed: audit and whole-file ACE-parity fixes, collation-true text, faster engine, Access SQL additions (11.0.0-alpha.4) - #303

Open
ChrisJollyAU wants to merge 87 commits into
CirrusRedOrg:masterfrom
ChrisJollyAU:alpha4
Open

ChrisJollyAU wants to merge 87 commits into
CirrusRedOrg:masterfrom
ChrisJollyAU:alpha4

Conversation

@ChrisJollyAU

Copy link
Copy Markdown
Member

This branch closes what two audits of LibRed and a new whole-file comparison against ACE turned up, makes LibRed's query engine substantially faster, compares text in the database's own collation, and adds a handful of Access and standard SQL forms. Nearly every behaviour change was measured against ACE first, and the file-format changes are pinned by byte comparisons against ACE's own output: a 139-statement run now leaves a file byte-identical to ACE's after 138 of its statements. Version bumped to 11.0.0-alpha.4.

EF Core providers (Jet and shared EFCore.Jet.Common)

  • Migrations no longer drop and recreate indexes around an ALTER COLUMN. That is SQL Server's workaround for refusing to alter an indexed column. Jet's ALTER COLUMN rebuilds the indexes over the column itself, keeping each one's name, columns, order and flags, the primary key included. The shared generator and LibRed's copy both lose the index handling.
  • A char converts to a number by its code point (AscW), not by parsing its text, in both SQL generators.

LibRed EF Core provider

  • UseLibRed accepts any DbConnection.
  • Extended mode writes an inline collection as SQL Server does: (VALUES (0, CLNG(1)), (1, 2)) AS v(_ord, Value). Before, it used EF's fallback: a leading SELECT naming the columns, then UNION ALL VALUES.

Packaging

  • LibRed.Core, LibRed.Sql, LibRed.Engine and LibRed.Ado target net10.0 as well as net11.0, since none of them depends on an EF Core package. LibRed.EFCore and the Jet provider stay net11.0-only. The test projects stay single-target; the net10.0 leg is compiled, not run.
  • The MyGet feeds are gone. CI builds stay available as the nupkgs artifact.

LibRed SQL engine

  • Text compares in the database's collation, everywhere, by LibRed's own keys.
    • Comparison, equality, hashing and sorting all take the collation on page 0, never a column's, as ACE does. That covers the evaluator, GROUP BY and DISTINCT keys, set operations, joins, IN sets, MIN/MAX, windows, list aggregates and referential-integrity keys.
    • Nothing consults the runtime's culture any more, so an answer no longer depends on ICU against Windows NLS. GROUP BY folds what the collation folds: ß with ss, case and trailing spaces, but not accents.
    • The keys cover the whole BMP, including the 66 controls, DEL and U+FEFF, verified against ACE.
  • LIKE folds case by ACE's own table. ACE was asked about 26,440 character pairs. It folds 973 case pairs, a narrower set than any runtime's casing, and spells out ß, æ, œ and þ. Before, LibRed matched 442 pairs ACE does not and missed 12 it does. The same pairs hold in all 405 creatable orders.
  • Values are typed as ACE types them.
    • A parameter takes the type of the column it is compared with: a number against a text column compares as text. Since alpha.3 it read the column as a number instead, and threw on the first non-numeric value.
    • CVar and Variants: a Variant keeps its type inside an expression but is written out as text. A choice mixing text with anything else, or a Variant with a non-Variant, is text. A date beside a number is a date. SWITCH and CHOOSE are typed as IIF is.
    • A value holding a choice anywhere is converted to the column's declared type. linq2db's 1000 - CASE WHEN s IS NULL THEN 0 ELSE s END over an empty group came back as the literal's Int32 under a Decimal column.
    • A scalar subquery declares its column's type, so a choice over it widens with it.
    • GROUP BY puts a Long 1 and a Double 1.0 in one group, as ACE does. Byte-array keys, compared by reference, gave every row its own group.
  • An index seek answers only in its column's own kind. S = @p, with a numeric parameter against an indexed text column, crashed with a cast error where the unindexed query returned ACE's 0 rows. Other cross-kind comparisons now scan.
  • New syntax:
    • Access's Nz, with the results Access itself gives: always a Variant, so Nz(K, 0) sorts as text, and one argument over a Null gives VBA's Empty.
    • WITH OWNERACCESS OPTION, requested by linq2db. It is accepted wherever ACE accepts it, changes nothing, and is kept in a stored query.
    • TOP n WITH TIES and FETCH … WITH TIES. ACE's own TOP always keeps ties; LibRed's plain TOP stays exact, so WITH TIES is how to ask for ACE's rows.
    • A derived table's column list, (query) AS t(a, b).
    • IS [NOT] TRUE/FALSE and IS [NOT] DISTINCT FROM.
    • ACE's {guid {…}} literal, and GUID text against an indexed GUID column.
  • Access's naming forms: [Table]![Column] bang notation, with an unaliased one named as ACE names it; [Forms]![f]![c] as one parameter name; Yes/On/No/Off as True and False; any word as a member after . or !; and unbracketed names in any script. Every stored query in the example corpus that uses a bang now parses.
  • Stored queries:
    • A stored query's DISTINCT, TOP and PERCENT are written on ACE's single option row. Three things were wrong there:
      • A stored TOP 50 PERCENT came back as TOP 50 and returned every row.
      • DISTINCT and TOP went on two rows where ACE writes one.
      • A make-table or append query lost its DISTINCT and TOP altogether, so a stored SELECT DISTINCT … INTO copied every row.
    • A bracketed parameter name is stored and read as ACE stores it. An ACE-written [@firstName] rebuilt as [[@firstName]] and did not parse.
  • Names: a name only has to be free within its own MSysObjects container, as ACE checks it. A query's name used as a table's is refused as "Table 'X' already exists." rather than as an index violation.
  • ADO.NET: every reader accessor returns the OLE epoch as 1899-12-30. GetDateTime gave default(DateTime) for a serial of 0.

LibRed performance

Measured with the benchmark harness against ACE at 10,000 rows.

  • Decode only what a statement reads. A pass prunes each scan and seek to the columns the plan names, subqueries and writes included, and a join builds its rows from those columns alone. Scan 2,707 → 1,652 µs, GROUP BY 8,372 → 4,436 µs, hash join 2,988 → 1,681 µs.
  • Text, currency and rows decode without general-purpose detours: plain UTF-16 is copied, Currency is built without dividing, and rows are walked as arrays.
  • One evaluator per query, rebound per row, instead of a new one per row. A projected column is located once, not by name per row.
  • GROUP BY aggregates in one pass, feeding each row to its group's accumulators rather than buffering the rows. A group's text key is built once.
  • Collation weights are looked up by code point, not binary search, and text keys are built without per-character allocation. A text GROUP BY over 10,000 rows went 28.5 → 13.7 ms across the changes.
  • Seeks, inserts and keyed lookups:
    • Referential-integrity and complex-value lookups seek by their key's index: 2,922 → 17 µs for one row of 10,000.
    • A transaction's own page parses are kept and survive its commit.
    • INSERT … VALUES has one grammar derivation, where the parser used to spend about 60% of an insert settling an ambiguity. A SQL insert in a transaction went 300 → 68 µs.
    • Opening a file keys it by its full path alone: 329 → 75 µs per open and close.

LibRed file format — writes that now match ACE

  • A read-only audit of LibRed.Core, then its completion:
    • Corruption:
      • CREATE INDEX after DROP INDEX could take another object's usage-map row.
      • A root split left other open handles on the old root.
      • The page cache keyed files on a lowercased path, merging two files on a case-sensitive system.
      • The password operations rewrote the file in place; they now stream through a sibling copy, exclusively opened, as ACE requires.
    • Wrong data:
      • A range seek on a descending index returned a fraction of its rows.
      • A Yes/No index key is flag plus value.
      • Text length was checked against compressed bytes.
      • Dates before year 100 were silently changed.
      • Negative-zero Decimal lost its sign.
      • Retyping a column on a General (v1) database wrote it back as v0.
    • Carried across rather than refused: ALTER COLUMN on a table owning a complex column carries all five of its links.
    • Relationships: an unenforced relationship no longer demands an index, and foreign-key write skew is re-checked at commit.
    • Readers stop papering over bad files: a walked index's keys must not go backwards, a page may declare at most 255 rows, and a missing index column or data block is reported.
  • A whole-file comparison with ACE runs the same statements through both engines on copies of one DAO-made database and compares the files byte for byte. What it found is fixed:
    • A table's property blob is inline up to 64 bytes.
    • A new object's MSysACEs rows are its container's inheritable grants.
    • A statement's deletes run in the order the rows were found.
    • An index built over rows writes each leaf at the prefix its filling reached.
    • A long-value page leaves the free map at 257 bytes free.
    • A row's chained long values are written before its single-page ones.
    • A compressed value is placed by its uncompressed size, which is left behind in the page's free space.
    • An UPDATE frees a replaced long value only after writing the new one.
  • Index B-trees split and build as ACE does:
    • A splitting root keeps its page, and both halves move.
    • A full page is compressed in place before it splits.
    • The cut follows the page's byte midpoint.
    • Nodes fill, compress and link as leaves do.
    • CREATE INDEX over existing rows writes the tree that sequential inserts would leave.
    • An emptied leaf leaves the tree.
  • Emptied pages: a data page whose last live row is deleted is stamped 0x09 and released. A table never gives back its first data page.
  • ALTER COLUMN is always an edit in place. To or from Memo/OLE it no longer drops and recreates the table. A TEXT(n) length change takes a full retype, as ACE's does. Several indexes over a column are rebuilt in ACE's order and slots.
  • DROP COLUMN of an append-only memo takes its version-history column, flat table and template with it when it was the last one.
  • Type 0x11 is BigBinary: up to 4,000 bytes, always inline, and treated as OLE in queries and indexes.
  • The catalog:
    • A relationship cascades by its index blocks' action bytes, as ACE does.
    • grbit 0x01 means one-to-one.
    • GetSchema lists names in collation order.
    • Every property type reads and writes, and a rewritten blob keeps its name pool: all 3,046 blobs in a corpus of Access files now rewrite unchanged.
    • The Name AutoCorrect maps read and write byte-identically.
    • The catalog column flags 0x10/0x20 are written only where Access writes them.
    • MSysObjects.Flags 0x00040000 marks a table that owns a complex column.
  • Usage-map rows follow CREATE TABLE's declaration order. A constraint written after a long-value column, the shape a migration emits, took a row ACE gives the column.
  • Collations:
    • The Chinese, Japanese and Korean sort orders are encoded, each measured over the whole BMP.
    • Page 0 carries each order's code page, where 1252 was always written.
  • The page type is the 16-bit word at offset 0.
  • Page 0's header mask is RC4, and SIDs use a per-file keystream. Both are recorded in the spec.
  • Left unmatched deliberately, and recorded: page placement for large inserts. ACE reserves session extents, four aligned 8-page groups at a time, and keeps data and index pages in separate groups; LibRed takes the lowest free page. A probe (PageGroupProbeTest) measures the behaviour, but what triggers an extent is not fully pinned. It is the one statement the parity test lists as different.

Tests and CI

  • The step-by-step whole-file parity test asserts. Every statement must leave a file byte-identical to ACE's, except a named list, and a listed statement that becomes identical fails too.
  • An ACE older than Access 2311 (build 16.0.17029) refuses unused ANSI-92 reserved words as names. The query-name cases skip there with that reason, as on CI's ACE 2016.
  • The lock-manager case test runs per platform: paths differing only in case are one file on Windows and macOS, and two on Linux.
  • Benchmarks append to their history only when run with --record.
  • New suites cover collation queries, LIKE folding against ACE, column pruning, index seek kinds, the parity fixes, stored query options and OWNERACCESS.

Docs

  • Format spec (src/LibRed/docs/format): updated with everything above that was measured, including released data pages (page-09, renamed for the page type), the version-history structure, the Name AutoCorrect maps and the object kinds.
  • docs/functions.md: Nz, Variants and the new predicates.
  • CLAUDE.md/AGENTS.md: the net10.0 targets, and hooks that deny reading or editing files through the shell and a whole unfiltered ACE access suite in the foreground.

🤖 Generated with Claude Code

ChrisJollyAU and others added 30 commits September 23, 2026 17:15
- Bump to alpha.4 after the alpha.3 release
- Remove the MyGet feeds; CI builds stay available as the nupkgs artifact
- Convert a char to a number by its code point (AscW), not by parsing its text
- Accept any DbConnection in UseLibRed

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A parameter takes the type of what it is compared with, as in ACE: a number
against a text column is compared as text, text against a number column is
read as a number. Since alpha.3 a number parameter against a text column read
the column as a number instead, and threw on the first value that was not one
('' or 'abc'), where ACE and alpha.2 answered.

Measured against ACE for =, IN and BETWEEN, with number, double and Boolean
parameters. A number literal is unchanged: ACE refuses it outright, LibRed
still reads the column as a number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GetValue returned 1899-12-30 for a serial of 0, while GetDateTime and
GetFieldValue<DateTime> turned it into default(DateTime). All now return the
stored date, as ACE's OLE DB reader does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The index key encoder reads a GUID from text, as the row codec already did;
  comparing or inserting a string GUID on an indexed column threw
- Both read ACE's {guid {…}} form, in text and as a literal; the grammar now
  lexes it (parser regenerated)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Standard SQL predicates that ACE rejects, added as extensions. Both are never
Null: a truth test treats Null as neither True nor False, and DISTINCT FROM
compares with Null taken as a value. Parser regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ACE stores a declared parameter's name as written, brackets included
('[@firstname]'). LibRed bracketed it again on read, so an ACE-written
parameterised query rebuilt as PARAMETERS [[@firstname]] and would not parse,
and reported a name that could not be bound.

- Read: keep a bracketed name as it stands; report and bind by the name inside
  the brackets
- Write: store the name as ACE does, brackets kept and a bare @ dropped

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0x11 is ACE's BigBinary, which its DataTypes list names but no one had
tied to a code: BIGBINARY(n) in DDL, up to 4000 bytes, bare BIGBINARY
taking the maximum. Measured against ACE:

- The descriptor is VARBINARY's apart from the type byte; no format raise.
- The value is always inline, never a long value, so it counts against
  the 4060-byte record cap: a binary bounded by a single data page.
- In queries and indexes ACE treats it exactly as an OLE Object: no sort,
  group, distinct, aggregate, join, union or index, same messages.
- MSysAccessObjects.Data, the only place it had been seen, is a fixed
  3992-byte column of this type.

LibRed now creates, writes and reports it as a binary type with a 4000
cap, and refuses it in an index as it refuses OLE. CreateFormat is filled
in for every DataTypes row, BigBinary included.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A scalar subquery declares its one column's type, found by describing its
  plan, so IIF/CASE/COALESCE over it widen with it.
- Only a NULL literal is skipped when unifying a choice's arms; any other
  untyped arm leaves the choice untyped, rather than letting the typed arms
  declare alone and converting every value to that.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each was measured against ACE first where ACE has an answer, and the
format spec corrected where the measurement disagreed with it.

Corruption:
- CREATE INDEX after DROP INDEX could take another object's usage-map
  row, computing it from a layout formula a drop invalidates. ACE never
  reuses an orphaned row, and neither do we now.
- An index root split left every other open handle on the old root, so
  seeks missed rows and a duplicate PK could get in.
- The page cache and lock manager keyed on the lowercased full path,
  merging two files on a case-sensitive system and splitting one file
  reached two ways. One FileIdentity.Key now serves both registries.
- A self-cascading UPDATE wrote stale snapshot values back over a key
  its own cascade had already moved.
- DROP TABLE and DROP COLUMN ignored complex columns; ALTER COLUMN on a
  table owning one is refused rather than run, since the rebuild carries
  none of its three links.
- The password operations loaded the whole file, truncated it and wrote
  it back. They now take a database the caller opened exclusively, as
  ACE requires of ALTER DATABASE PASSWORD, stream pages through a single
  page-sized buffer into a sibling copy and replace the database with
  it, so a change is atomic and never writes plaintext to disk.

Wrong data:
- A range seek on a descending index returned a fraction of its rows.
- Deleting from an ACE-written leaf whose shared prefix reaches into the
  row pointer threw.
- A Yes/No index key is flag + value, not a bare byte, and -1 keys true.
- Negative-zero Decimal lost its sign on rewrite, orphaning its entry.
- Text length was checked against compressed bytes, so a TEXT(5) took
  eight characters where ACE refuses six.
- A null PK or DISALLOW NULL key was accepted, and an AutoNumber could
  be updated; ACE refuses every update of one.
- Dates before year 100 were changed silently. 0100-01-01 is ACE's
  floor; below it is now refused, naming the column.
- Owner/ACE SIDs were hard-coded with two different files' masks. The
  mask is recoverable from MSysObjects' own owner, so both writers now
  derive the accounts per database.

Leaks, parity and transactions:
- DROP INDEX freed only the root page and left the usage-map record
  behind; both drops now retire it, byte-identical to ACE.
- Space freed mid-table never set the page's free-map bit, so it was
  never offered to another insert.
- A freed page outside a map's range vanished instead of being reported.
- The row's leading count is the TDEF 0x29 id high-water, a dead id's
  null-bitmap bit is clear in an inserted row, and the variable trailer
  survives the last variable column being dropped.
- A failed statement inside a nested transaction left its savepoint
  frame open, after which every COMMIT threw.

Tests live in files named for their subject rather than for the audit,
and the legacy Jet 4 fixture is a database DatabaseCreator builds rather
than a header followed by random bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The rest of the audit's findings, and then a pass over the shape they
share rather than the instances.

Findings:
- ALTER COLUMN on a table owning a complex column carries it across
  instead of refusing. Five links, not the three the audit named: the
  descriptor's 0x0B, the flat tables, the catalog rows, the in-row
  complex ids, and the 0x1C counter a recreated table restarts at zero.
  The property blob and MSysACEs grants come across verbatim too.
- LibRed demanded an indexed parent key for every relationship, where
  ACE needs one only to enforce. Measured over three files: every
  enforced relationship has an index on both sides, no unenforced one
  does. An unenforced relationship is now catalog rows and nothing else.
- Foreign-key write skew. A transaction records the conditions its
  writes depend on and they are re-checked under the publication gate.
  Semantic, not page-based, so an unrelated row on the same page cannot
  produce a false conflict.
- A relocated row that outgrows its new page moves again rather than
  throwing, byte-identical to ACE, tombstone included. Both write paths
  had their own copy of the limitation.
- ADD COLUMN measured the fixed region as the sum of the live columns'
  lengths, ignoring the hole a dropped one leaves; DROP INDEX could
  renumber a data block another logical index still named; a memo-
  indexed UPDATE threw; SET NULL onto a row the same DELETE was removing
  left the delete holding a stale key; an update that freed a memo then
  failed had no undo outside a transaction; a rename left a complex
  column's catalog row naming a column that no longer existed; ACE moves
  DateUpdate on an ALTER and LibRed never did; an unknown version byte
  threw on every rollback; a handle kept a stale Format after another
  raised it; a failed close left the connection pointing at a disposed
  database.

Measured and found not to be defects: the LvProp owner record's second
field is always zero; a leading U+FEFF is indistinguishable from the
compression marker and Access loses it the same way; ToOADate quantises
to milliseconds on the way in, so a DateTime key round-trips exactly.

Readers no longer paper over what only this engine could have written: a
walked index checks its keys never go backwards, a page declaring more
than 255 rows is refused, the TDEF reader stops silently dropping an
index whose column or data block is missing, an unresolvable relocation
pointer is reported, and an AutoNumber increment of 0 — which ACE
refuses outright — is refused on write and reported on read.

Then the shape behind the memo cluster. A row had two representations in
one object?[] — the values a caller holds and the descriptors the record
carries — and materialising in place made the caller's array change
meaning halfway through a write. Every memo defect was the wrong one
reaching a consumer that wanted the other. Insert and Update now keep
the logical row and encode from a separate storage copy, which also
removes the two conditional clones that existed to work around the
mutation; RowDecoder refuses to decode a long value without the reader
that resolves it, and the descriptor-only operations are static so there
is no mode to get wrong.

The same shape turned up a live bug: TableCreator and TdefBuilder both
defaulted a missing collating order to General-Legacy, and only 2 of 18
call sites passed one. On a General (v1) database, retyping a column
wrote it back as v0 — a different key encoding in a file whose other
columns use v1. It survived because every fixture here is v0, so the
wrong default was the right answer everywhere it was exercised. The
order is now required, which turned the defect into 16 compiler errors.

Also folded together: five bespoke catalog-row updaters into one beside
DeleteCatalogRows, twelve copies of the distinct-real-index expression
into TableDef.RealIndexes, and the transaction-state clear repeated in
three places.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The four items the audit left needing a measurement, and the parity gaps
chasing them turned up. Each was measured against ACE before anything
changed, and the spec updated where the measurement went past it.

The probes:
- An emptied index leaf leaves the tree. ACE closes the leaf chain over
  it, drops its separator from the parent node and releases the page;
  LibRed kept it linked with a live separator pointing at no keys. The
  node is left with no entries and only its child-tail, which is a valid
  node rather than one to collapse. A leaf that IS the root, or that is
  its parent's only remaining child, stays: an empty leaf is valid and
  neither shape has anywhere to go.
- A node page's prev/next are 0 in both, as LibRed already wrote them.
- Currency rounds half-to-even, on both routes into a Currency - the
  CCur function and storing into a CURRENCY column - which is what
  LibRed already did. The only doubles that can settle a 4-decimal
  midpoint are the odd multiples of 1/32; anything else is already off
  the midpoint and rounds the same way under every rule.
- An empty Memo is the full 12-byte inline descriptor, 00 00 00 80 then
  eight zero bytes, and an empty LONGBINARY likewise. Identical to
  LibRed's, so the zero-width slot its reader refuses is a form ACE
  never writes and the reader's strictness is right.

Reclaiming an emptied page, which the empty-leaf byte diff exposed:
- A data page whose last live row is deleted is stamped 0x09, cleared
  from both of the table's maps and released. So 0x09 is not confined to
  packed long-value pages - which is exactly what page-09 recorded as an
  unaccounted-for case - and the two are told apart by the owner at
  0x04, an LVAL signature against the owning TDEF. The spec file is
  renamed for the page type rather than for one of its causes.
- A table never gives back its FIRST data page. Deleting every row of a
  35-page table releases 34 and leaves that one at 0x01, empty, still in
  both maps; it is kept even while later pages hold live rows.
- A page emptied by reclaiming a hidden relocation target is not
  released - ACE keeps it as an ordinary data page. Only deleting a live
  row releases the page it empties.
- A released index leaf keeps its 0x04 type byte and is never stamped,
  so a 0x09 page is always a data page.

And the usage-map row order, the open half of the audit's first finding.
Rows after the table's own two follow the order the CREATE TABLE
statement declares them: a long-value column takes two where its column
is written, an index one where its constraint is. LibRed always used the
inline order, so a constraint written after a long-value column - the
shape a migration emits - took a row ACE gives to the column. The
constraint's position now travels from the parser, which already tracked
the token offsets for self-reference resolution, through to TableCreator.
Index data blocks are unaffected and keep their PK-then-unique-then-FK
order, so a block's ordinal and its map row are independent.

The other half of that finding - CREATE INDEX reusing an orphaned
usage-map row - was already closed; the audit's note had gone stale.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CVar was a pass-through. Measured against ACE over OLE DB, a Variant
keeps its own type while an expression uses it and through a derived
table, but is written out as text (as CStr writes it) by a result, a
scalar subquery, a union and a make-table query. A choice whose values
differ in kind - text beside anything else, or a Variant beside a
non-Variant - is mixed: text wherever it is written out, derived tables
included.

- Both sort, group and take Min/Max/First/Last as their text.
- As an operand either counts as a Double; only + beside a Variant, a
  mixed value or text keeps a Variant.
- In a choice, a date beside a number is a date, a Boolean beside one a
  Long; SWITCH and CHOOSE are typed as IIF is.
- A value holding a choice anywhere is converted to its declared type.
- CVar(Null) stays untyped, so EFCore.Jet's projected Nulls leave a
  union typed from its other arm, where ACE makes it text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…found

Sixteen findings: where the same value is derived twice, and where one
job is written out along several paths. Every "redundant" claim was
checked against docs/format/ first, because the format sometimes
requires the second pass - one finding turned out that way.

Two were live bugs the copies had hidden:

- GROUP BY compared its keys with CLR equality for anything that was not
  text, so a LONG 1 and a DOUBLE 1.0 arriving under one key were two
  groups where ACE returns one, and byte-array keys - compared by
  reference - gave every row a group of its own. Measured against ACE
  through an IIF with differently-typed arms, a UNION of a LONG and a
  DOUBLE column, and a DOUBLE against a CURRENCY. The fold is within a
  kind, because the hash partitions by kind: text keeps the invariant
  comparison its hash is built on rather than taking CompareForSort's
  database collation, which disagrees with it over ss and where the
  collation key is far too heavy to build per row.
- IndexKeyDecoder's FixedKeySize knew neither Complex nor FixedPoint and
  returned -1, which its caller reads as a lossy key to stop at, so a
  DECIMAL key column ended the decode and nulled itself and every later
  column of a composite key. The encoder's copy knew both. One table in
  Formats now, so they cannot drift again.

The walks and writers, each written once:

- TdefRegions.Of replaces five hand-written walks over the TDEF's
  variable-length regions. Only IndexWriter's bounded its steps; the
  three in TableCreator advanced raw offsets, so a file-sourced count or
  name length could carry a DDL write past the buffer - or, on an
  overflow, back inside it at the wrong index block. They are bounded
  now.
- UsageMapBits.Append replaces four copies of the bitmap-to-page-number
  loop. The validation difference between them is deliberate and stays:
  a read feeds its numbers straight into page reads, so one outside the
  file is corruption, while a rewrite must keep bits it cannot currently
  represent and widen the record instead. Rejecting there had already
  failed an ordinary DROP TABLE on a real file.
- CatalogWriter writes an object's MSysObjects row and its MSysACEs pair
  once, for tables, queries and relationships alike; each caller still
  owns the type, container, flags and masks that genuinely vary.
  DatabaseCreator keeps its own, populating both tables before either
  has an index or a catalog to resolve a name through.
- Rc4Cipher.Apply serves the legacy Jet codec and the Office-Standard
  one, and Agile's spin derivation is shared by its open and create
  paths.
- JetCatalog.RequireTable and TableDef.RequireColumn replace twenty
  copies of the system-table lookup-or-throw and thirteen column ones,
  along with four private aliases for the same lookup. ObjectProperties
  also returns the Id index now, so five LvProp read-modify-write
  methods lost four identical opening lines each.

The recomputes:

- INSERT read a data page on every row to re-derive the fixed-region
  length. The region may never shrink below what existing rows carry,
  and a retired column id leaves a hole the live descriptors cannot see
  - but only then, so the read is skipped when 0x29 says no id was ever
  retired.
- A row's long values were freed by re-reading the whole definition per
  column; the long-value map pointers cannot move under an inserter, so
  they are read once.
- UPDATE parsed one row's layout twice, a column reference was resolved
  by name for every row, and Format() re-parsed its format string on
  every call.
- Two callers materialised every page a table owns to look at one of
  them, which made a bulk load quadratic; both ends of the owned-pages
  bitmap now share one scan.
- The property blob was parsed once per column, for every table, on
  catalog load; its five accessors take the parsed list instead.

AggregateSurfaceTests binds the three places that answer the aggregate
surface, and AggregateResultType's fall-through returns null rather than
the argument's type, so a new aggregate reaching it fails the test
instead of being declared as whatever it was fed.

Left open: there is still no parse or plan cache, the one finding with
real cost behind it. Plans embed live IndexDefs and Invalidate with
markChanged false breaks the obvious cache key, so it needs its own
change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both were already covered in prose and enforced by nothing.

- Reading repository files through git. Read-only git is pre-allowed,
  so it was the way to spell a read the read rule did not match and
  skip its prompt. Reviewing history is untouched.
- LibRed.Core.AccessTests and LibRed.Engine.AccessTests are about 7 and
  1.5 minutes, and running either whole and in the foreground is dead
  session time. They now need a --filter naming the tests that cover
  the change, or run_in_background. The other suites are left alone -
  Ado and EFCore are two seconds each, and friction on a rule with no
  payoff is what creates pressure to work around it.

The first pattern's own first draft allowed arbitrary text between the
two words, which made the commit describing it match itself; only
option-like tokens may sit in the gap now.

Not added: a block on the functional suites, which CLAUDE.md also
forbids in prose. That rule has never been broken, because it is
binary - a project name is either on the list or it is not. The rule
that does get broken asks for a judgement about which tests cover a
change, and no hook reaches that. Requiring a --filter is the closest
proxy, since it forces the claim to be named.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A comparison gave a different answer depending on whether the column
happened to be indexed. `S = 1` on a text column compares as a number
(ComparisonOperatorTests), so 'abc' is a type mismatch - but with an
index on S the seek reached IndexKeyEncoder, which casts the value by
the column's type and raised "Unable to cast Int32 to String" out of the
storage layer instead. The encoder is right to demand a typed value; it
is shared with the row-writing path, where the type is guaranteed. The
engine was handing it a value of another kind.

Worse on the parameter path, which is the one EF uses: `S = @p` with a
numeric parameter answers 0 rows unindexed - a parameter compared with
text takes the text's type, verified against ACE, which answers 0 too -
and crashed with the same cast error when S was indexed. So the index
changed the result, not just the error.

TryGetSeekKey settles what a seek may be keyed by. An index answers only
in its column's own kind, because that is how its keys are encoded and
ordered, so a comparison happening in another kind has no key range to
seek: `S = 1` matches ' 1 ', '1.0' and '+1' as well as '1', which sit
nowhere near each other in a text index. The one cross-kind case that IS
seekable is a parameter against text, which CompareAsKinds converts to
text before comparing; the seek converts it the same way and so asks the
index the same question. Anything else falls back to a scan - safe
because the single-table seek keeps the FilterNode it was planned under
and the index-nested-loop join keeps its ON whole as the residual, so
only the reading strategy changes, never the answer.

Measured against ACE before and after: the indexed and unindexed forms
now agree everywhere, and `S = @p` returns the 0 rows ACE returns. The
literal coercion policy is untouched - text against a number still reads
as the number it is, as SQL Server does and as ComparisonOperatorTests
records, where ACE refuses the comparison outright.

IndexSeekKindTests pins indexed against unindexed for both literals and
parameters. Reverting TryGetSeekKey to its old behaviour fails five of
them and leaves the same-kind and NULL controls passing, which is how
they were checked. Two of the five were not in the original report: a
numeric indexed column against a text literal crashed the same way, and
so did a cross-kind key in an index-nested-loop join.

Also drops the last reference to a probe that should not have been
committed with the audit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LibRed.Core, LibRed.Sql, LibRed.Engine and LibRed.Ado depend on no EF
Core package, so nothing about them requires the version the rest of the
repository is pinned to. They now multi-target LibRedTargetFrameworks,
net10.0 plus whatever JetTargetFramework is, which lets a consumer on the
current LTS use the format reader/writer, the SQL front end, the engine
and the ADO.NET surface without moving first.

LibRed.EFCore is deliberately not in that list - it sits on EF Core 11,
which is net11.0-only - and neither is anything under src/EFCore.Jet.

Every test project stays single-target net11.0, which is a choice rather
than an oversight and is written down in CLAUDE.md so it does not get
helpfully undone. The net10.0 leg is compiled and never run: multi-
targeting the suites would double every run locally and across the five
CI platforms, and would need the .NET 10 runtime on each runner, which
the SDK global.json pins does not carry. Compiling it needs only the
reference pack, which restore fetches, so CI needs no change at all.

Checked rather than assumed, once: the whole solution builds both legs
with no warnings, so nothing in the code was relying on a C# 14 feature
that net10.0 cannot have despite LangVersion being preview; and with
LibRed.Core.Tests and LibRed.Engine.Tests temporarily multi-targeted,
both passed identically on net10.0 - 660 and 4137. Those two csproj
changes were then reverted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- The 128-byte page-0 header mask is the RC4 keystream of the fixed key
  C7 DA 39 6B, verified over all 128 bytes.
- An on-disk SID is its workgroup SID XOR'd with a per-file keystream from
  its first byte; the 2-byte short-SID mask is that keystream's first two
  bytes, and the 102-byte SIDs Access adds use the same keystream.
- How the keystream derives from the creation date is recorded as not
  known, with the derivations ruled out, replacing the shared-PRNG guess.
  The workgroup Admins SID is 102 bytes, not 98.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
[Table]![Column] is Table.Column in every clause, and an unaliased one is
named as ACE names it, text as written. A longer chain such as
[Forms]![frmMenu]![txtCity] is one name, reachable as a parameter the query
declares; a period between parts counts as a bang, as ACE binds them. A
stored parameter declared that way now reads back as that name rather than
with only its outer brackets stripped.

Yes/On and No/Off are True and False, winning over a column of that name
unless it is qualified. Any word, reserved or not, names a column after a
period or bang, and an unbracketed name may use any script's letters.

Every stored query in the example corpus that uses a bang now parses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… blob

The flag byte is a bit field, not a DDL boolean: Access writes 0x80 on its
own account, and every stored query of one example database carries it.
Reading refused anything but 0x00/0x01, which would stop a table carrying
such a property from loading; the byte is now kept whole and written back.

A value block's type says what owns it, and an index's block (0x0002, by
mdbtools' account) is named for the index, usually its column's name. Every
block was read as a column's and written back as one, so a column's Required
or DefaultValue could come from an index, and a DROP, RENAME or ALTER of the
column moved the index's properties with it. Each property now keeps its
block type, and the column and table accessors match only their own block.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Column flag 0x10 is set on every column of MSysObjects, MSysACEs,
MSysQueries, MSysRelationships and MSysComplexColumns and on no other, and
0x20 on the two security-identifier columns, MSysObjects.Owner and
MSysACEs.SID. The spec listed neither and said the unlisted bits were zero in
every file. The extended flags gain the attachment value column's 0x10, and
the 0x04/0x08 a complex column's flat table sets.

A new database gave the MSysComplexType_* templates the catalog flag, which
Access never does; the flag was also what zeroed their 0x09. That is now its
own switch, and the attachment template carries its extended flag, so every
system column a new database writes matches Access's byte for byte.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…file diffing

A new probe runs the same statements through ACE and LibRed on copies of one
DAO-created database and compares the files byte for byte, both at the end of
one session and statement by statement. What it found, each now verified and
fixed:

- A table's property blob is stored like any long value: inline up to 64
  bytes, on a page above that. LibRed always used a page.
- A new object's MSysACEs rows are its container's inheritable grants, the
  Creator's becoming the owner's, so the masks follow the database rather than
  the object class the three constant pairs assumed.
- A foreign key's index root is allocated after its relationship's rows.
- A statement's deletes run in the order the rows were found, which decides
  the bytes left in the space they free.
- An index built over rows writes each leaf at the prefix its filling reached,
  not the largest its keys share.
- An UPDATE moving an entry lowers the index's total and holds its unique count
  to it, on an index built over rows.
- A long-value page leaves the free map with 257 bytes free or fewer.

Both measurements now come out identical: all 104 statements, and the whole
run's 74 pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… pages

The whole-file probe now also runs a thousand-row INSERT ... SELECT, a
253-column table and long values chained across pages. What it found about
index splits, measured against ACE and now written the same way:

- A root that splits keeps its page: both halves move to new pages, left
  first, and the root becomes the node over them.
- A full page is compressed in place before it splits, and the split is made
  over that image; the left half keeps its prefix unless the new entry became
  its first, when it is written whole.
- The cut keeps every old entry that starts before the page's byte midpoint,
  the new entry joining its side; a new first entry halves the entries by count.
- Nodes fill, compress and split as leaves do, link to their siblings, and split
  at the right edge when the new separator is last.
- Bytes past each half's live end match ACE's, including the promoted entry a
  node split leaves behind.

CREATE INDEX over existing rows writes the tree sequential inserts leave, so
the bulk build now follows the right edge of that tree instead of building node
levels bottom up, and no longer moves the index's root.

ADD COLUMN writes descriptor 0x09 as ACE does: it is DAO's OrdinalPosition,
which ADD COLUMN compacts to ranks, and an added fixed column's variable-table
index counts dropped variable columns too.

Page 0's 0x30 and 0x34 are bounded by the largest possible page, not by the
file's length, as the spec had said.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A profile of a 10,000-row scan put most of its time in three places, none
of them the format:

- Plain UTF-16 text went through Encoding.Unicode, which counts and then
  converts with full validation - about a fifth of the scan. Without a
  surrogate code unit and at an even length the bytes ARE the string's
  chars, so they are copied; anything the decoder could treat differently
  (a lone surrogate, an odd trailing byte) still goes through it.
  Compressed text with no switch byte is one Latin-1 run, which is what
  the per-character loop produced for it; the mixed form decodes into a
  stack buffer rather than a StringBuilder.
- CURRENCY was divided by 10000m per value. Division returns the smallest
  scale that holds the quotient, so the same decimal - value and scale,
  which shows in its text - is built from the integer by dropping the
  four places' trailing zeros. CurrencyDecodeTests holds the two equal
  bit for bit across 600,000 values and the edges.
- RowDecoder enumerated its columns through the interface, and
  HasVariableSection re-scanned them, for every row. It walks an array
  now and answers the variable-section question from the lowest variable
  column id; Boolean values come from two shared boxes.

The raw scan goes from 4,084 to 2,895 us and 5,385 to 4,761 KB.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fbbed31 memoised where each column reference resolves inside EvalScope,
on the premise that only the row moves between rows. That held for the
join paths, which rebind one scope, but projection, WHERE, sort keys,
group keys and aggregates built a fresh scope and evaluator per row - so
each row also built a dictionary, used it once and dropped it. A scan
allocated 11,577 KB where it had allocated 9,388, and scan, GROUP BY and
the unindexed sort all slowed; found by bisecting on allocation, which is
exact where timings are not.

Those paths now make one evaluator and rebind it per row, as ExecuteJoin
already did. Filter and projection make theirs inside the iterator, so
each enumeration of a result has its own. The memo then pays off
everywhere, and the per-row scope and evaluator go too: a scan now
allocates 5,871 KB.

A projection item that is just a column of its input is located once
when the projection is planned and read from the row, rather than
resolved by name on every row; it still takes the item's conversions. A
name that is not exactly one input column (outer, niladic, ambiguous)
stays with the evaluator, so its error still comes from the row that
raises it. EvalScope.Locate is the one matching rule for both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… fixes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decoding was most of a scan's cost, and every read decoded every column: a
three-column SELECT over a ten-column table built all ten, and a GROUP BY
held all ten of every row alive. A row decoder now takes a mask, and a
column outside it reads as null - so the mask is only ever set where
nothing can look at that column.

Queries: ColumnPruning runs after index selection on every plan that is
executed - top-level statements, INSERT ... SELECT and SELECT INTO
sources, derived and lateral sources, subqueries - and tells each scan
and seek which names to decode. The names come from every column
reference anywhere in the plan, subqueries included (their correlated
references read the outer tables), by name alone, so a table decodes
anything a reference could mean. They are found by walking every public
property of the plan and AST records, so a clause added to a node later
cannot be missed; a value the walk does not know turns pruning off. A
read is pruned only below a projection or aggregate without a star,
because SELECT *, set operations and DISTINCT pass whole rows to the
output, and DISTINCTROW dedupes on the underlying rows - which is also
why nothing below a node the pass does not know is pruned.

UPDATE and DELETE: the read that chooses the rows decodes what the
statement names anywhere, and CompleteRows then reads each row about to
be written in full, into the array its joined rows share, before any of
it is used - the rewrite, its constraint and cascade checks and its index
moves all see every value, as before. The parent check every child
INSERT makes decodes the key alone; a cascade finds its children by key
and reads the matches whole; complex-value cleanup reads the link and the
index keys.

ColumnPruningTests covers each place a column can be named, each shape
that passes a row through whole, and each write. Breaking the write-side
completion fails five of them, and pruning below DISTINCTROW two (it
collapses 12 rows to 3).

Against ACE at 10,000 rows: scan 2,707 -> 1,652 us, GROUP BY 8,372 ->
4,436, hash join 2,988 -> 1,681, unindexed TOP 10 4,637 -> 3,317. The
walk reads properties through compiled getters: through PropertyInfo it
cost a primary-key lookup 4 us of its 22.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…or a match

The reads that find a row by one or two columns decoded every column of
every row they passed: an index back-fill decoded whole rows to build
keys from a few, and each MSysObjects lookup resolved every object's
LvProp blob - a long value - on the way to the one it wanted.

Table gains DecodeOnly, a mask for a read that looks at nothing else, and
RowsWhere, which searches on the key alone and reads each match whole,
because a row looked up by key is usually about to be rewritten or
deleted. The index back-fill and its duplicate check, the narrowing
ALTER COLUMN check, and the object-name, permission and next-id scans
decode only what they compare. The LvProp edits, relationship renames
and deletes, catalog-row updates and deletes, permission reads and
MarkAsSystemTable find their rows through RowsWhere (most via the
existing RowsKeyed), as do a cascade's child search and
ReadComplexValues, which no longer decodes other records' attachments
to find one record's.

A back-fill over 2,000 rows allocates 2,045 KB where it allocated 2,662;
the catalog scans are too small to show on the benchmark corpus's three
tables. ALTER COLUMN retypes, rebuilds and an AutoNumber ADD COLUMN
rewrite every row and still read them whole. LibRed.Core.AccessTests
pass, so the DDL still writes what ACE writes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ChrisJollyAU and others added 20 commits September 28, 2026 15:22
A name only has to be free among the objects of its MSysObjects container -
the rule the unique (ParentId, Name) index states. A table or query clashes
with a table, query or linked table (the Tables container) and with nothing
else; measured against ACE for every object kind.

CREATE TABLE and SELECT INTO now share the rename check, which tests the
container rather than the table and query types, so a query's or linked
table's name is refused as "Table 'X' already exists." instead of surfacing
as an index violation. CREATE VIEW no longer refuses the names of
relationships, forms, reports, macros, modules, database documents and
containers, and reports a clash in ACE's words.

The spec gains a table of the MSysObjects object kinds - their types,
containers and filled columns - and the name rule, and drops the claim that
tables and queries sit in different containers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… word

Access keeps each object's Name AutoCorrect map twice: in the NameMap column
of its MSysNameMap row and in a NameMap property in its LvProp, in two
layouts. NameMap decodes and encodes both, carrying every field - the
uninitialised bytes Access leaves, a missing terminator, the property's end
record - so a map read and written back is byte-identical; every map in the
corpus round-trips. JetDatabase reads the rows and a named object's property,
and replaces either; a row is never added, its Id not being understood. The
engine maintains neither copy, through ACE or LibRed.

MSysNameMap.NameMap has an owned-pages map and no free-pages map, as Access
writes it (and MSysAccessXML.LValue); every value on such a column takes a
page of its own. The long-value writer refused the column outright, so no
map over 64 bytes could be stored; it now writes one the way Access does, and
ACE reads it back and keeps it through a compact.

The page type is the 16-bit word at offset 0, not a type byte followed by a
constant flags byte: every page Jet and ACE write carries 0x01 in the high
byte, and ACE tests the whole word - it reads a table's owned page as rows
only for 0x0101, 0x0103, 0x0104, 0x0106, 0x0107 and 0x0109, and skips
everything else, 0x0001 and 0x0201 included. PageType is now that word and
every check compares it whole; what LibRed writes is unchanged.

The spec gains MSysNameMap, both map layouts, the owned-only long-value
column, and the page type as a word, with the field at 0x02 named for what it
is: the page's free-space count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A row whose relocation pointer names a page the file does not hold still
counts - COUNT(*) counts slots - but reading it fails "Unrecognized database
format", and ACE marks the reading user's commit slot 01 00, after which
every open fails until a repair. Unlike a page missing from a table's
owned-pages map, which the scan skips silently. LibRed refuses the read too,
and leaves the commit slot alone, having none of its own.

Page 0's commit-byte section now lists the two measured occasions for 01 00,
and three page types the move to 16-bit types missed are corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A table keeps the history of all its append-only memos in one hidden complex
column, VersionHistory_F5F8918F-0A3F-4DA9-AE71-184EE5012880, whose template
and flat table carry a value column per memo. LibRed dropped the memo and
left all of it behind. It now does what ACE does: while another append-only
memo remains, only the dropped memo's value column goes, from the template
and the flat table; when it was the last, the hidden column and its index,
its MSysComplexColumns row, the flat table and the template go too - their
MSysACEs rows left behind, as ACE leaves them - with the table-level
AppendOnly property, and the table's complex-column flag unless another
complex column remains. The table-level property leaves in the same rewrite
as the memo's own, as ACE writes it once.

The spec gains the version-history structure and both drop rules.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ALTER COLUMN to or from Memo/OLE fell back to a logical rebuild: the table
dropped and recreated, every row re-inserted, every index rebuilt - the
primary key moved to the front - where ACE edits the column in place and
rebuilds only the indexes over it. It now takes the same in-place edit and
re-lay as every other retype, plus the long-value side of ADD and DROP
COLUMN: a column becoming a long value gets its map entry and records as ADD
COLUMN places them, and each value is stored as an insert stores it; a
column ceasing to be one has its maps retired and its pages released as DROP
COLUMN does. The rebuild and the helpers only it used are gone.

The re-lay shared by every retype also differed from ACE on rows whose old
value was NULL: the dead id's null bit is now carried over from the old row
rather than set, and a NULL fixed target's new slot keeps the old record's
bytes rather than zeros. Text to Memo, Memo to Text, Memo to Memo and a
fixed retype over a NULL now leave the file ACE leaves, bar MSysObjects'
DateUpdate and, for the Memo re-declaration, stale bytes below the live
records on the usage-map page.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With every ALTER now an edit in place, nothing builds a column from an
existing descriptor any more. The raw-descriptor passthrough on ColumnSpec,
BuildColumnDescriptor's start-from-the-original path, the two flag masks
it needed, and TdefBuilder.Build's complex-counter parameter had the
rebuild as their only caller, and go. The comments and the spec sentence
that described unmodelled descriptor bytes surviving "through a rebuild"
now say why they survive: the descriptor is never re-emitted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A TEXT(n) to TEXT(m) change edited the descriptor's length and nothing
else. ACE has no such path: widening or narrowing, indexed or not, it burns
a fresh column id, re-lays every row and rebuilds the indexes over the
column, as for any other retype - which the spec already said. The length
change now takes that retype, and a value too wide for the new declaration
is refused by the re-lay's own width check, with ACE's message.

With more than one index over the retyped column, LibRed rebuilt each in its
own slot. ACE rebuilds them in logical-block (name) order, the primary key
included, allocating their roots in that order, and hands them back the
real-index slots they held between them in that order; the logical blocks
over them hand their numbers round the same way, which is distinct from the
data ordinal once two logical indexes share a real one. A column becoming a
Memo gets its map records ahead of the rebuilds' on the usage-map page.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The migrations SQL generators wrapped a column's type change, narrowing or
NOT NULL in DROP INDEX / CREATE INDEX for every index over it - SQL Server's
workaround, since SQL Server refuses to alter an indexed column. Jet does
not: its ALTER COLUMN rebuilds the indexes over the column itself, keeping
each one's name, columns, order and flags, the primary key included, and a
same-type NOT NULL leaves them untouched. A computed-column change needs no
index handling either, since Jet cannot usefully index a calculated column.
That leaves nothing for GetIndexesToRebuild, DropIndexes, CreateIndexes or
the operation list they read, and all of it goes, from the shared generator
and LibRed's copy alike.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
OLE DB runs ACE in ANSI-92 mode, and before Access version 2311 (build
16.0.17029) ACE refused an ANSI-92 reserved word that Access's own SQL does
not use as a name - even qualified, K.Except or K!Except - with reserved error
-1001, which has no message. 2311 fixed it, so the current engine takes Full,
Then, Cross, End, Fetch, Next, Rows, Only, Intersect, Except, Restrict,
Temporary and Language like any other name, and LibRed follows it. The ACE
2016 redistributable CI installs and ACE 2010 both predate the fix, so there
those cases now skip with that reason; any other refusal still fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nz is the Access application's, not the expression service's - ACE's OLE DB
provider answers "Undefined function" - yet queries written in Access use it
everywhere, and whoever wrote them expects what Access returns. Measured
against Access itself over the same rows: the result is always a Variant, so
Nz(K, 0) is written out as the text "0", ORDER BY Nz(K, 0) sorts 10 before 2,
and Nz(K, 0) + 1 still adds. LibRed already models that Variant for CVar, and
Nz takes the same path. With one argument a Null gives VBA's Empty, written
out as "" but read as 0 by arithmetic, so Nz(Null) + 2 is 2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
WITH TIES returns the n rows and every further row whose ORDER BY keys equal
the last one's, PERCENT included. It is the rule ACE's own TOP always
follows - a plain TOP 2 there returns every row tied at the boundary, which
LibRed's exact-n TOP deliberately does not - so WITH TIES chooses the rows ACE
chooses (checked against ACE's plain TOP as a set, since ACE orders tied rows
unstably). The standard's FETCH … ROWS WITH TIES takes the same meaning.

The node that orders the rows makes the cut, since only it holds their keys:
the sort, which then stays above the joins, or the aggregate over a grouped
query. It needs an ORDER BY, as in SQL Server, and is refused beside
DISTINCT, whose rows collapse above the sort. A view refuses it: a stored TOP
has no ties flag, and ACE reads one back with ties where LibRed reads it
without.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…at way

A derived table now takes the standard's column list, (query) AS t(a, b),
naming its columns in order - a VALUES table's included, whose columns have
no names of their own. SQL Server and PostgreSQL take the same syntax; ACE
has none, so a view refuses one (Access stores a derived source as its query
text alone), and a derived table with one is only ever read, never written
through. The list must name every column, each once.

With that, extended mode writes an inline collection the way EF Core's
SQL Server provider does, (VALUES (0, CLNG(1)), (1, 2)) AS v(_ord, Value),
instead of EF's fallback for databases without a column list, which names
the columns on a leading SELECT and puts the other rows after UNION ALL
VALUES.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The registry test acquired C:\dir\db.accdb and C:\dir\DB.accdb and expected
one manager. That held while the file key folded case everywhere, but the
key now folds it only where file names are case-insensitive, Windows and
macOS: on Linux those are two files, and two managers is right. So the test
failed on every Linux leg. The refcount check now uses one path, and the
case rule has a test of its own that expects one manager on Windows and
macOS and two on Linux.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The whole-file parity diff showed ACE placing the pages of a large
INSERT ... SELECT in 8-page groups where LibRed takes the lowest free page.
This probe measures why, on a thousand-row insert and variations of it:
row by row, and with the file read while ACE's session is still open, where
its reservations are visible.

ACE reserves extents in the global free map, four aligned 8-page groups at a
time, holds them for the session and frees the rest at close; outside them
it keeps data and index pages to separate groups, and grows the file at the
end or the next 8-page boundary depending on what the session has used.
What triggers an extent is not fully pinned. Every test is explicit: the
reservation samples take up to twenty minutes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The step-by-step comparison only reported. 136 of its 139 statements leave
a file byte-identical to ACE's, so it now asserts that, with the three that
differ named and the reason for each: the bulk INSERT ... SELECT (ACE's
extent allocation), a row's long values written in a different order, and
an UPDATE that reuses the pages of the value it replaces. A named statement
must still differ, and differ only in pages, so the list cannot outlive a
fix. The whole-run comparison stays a report: one early placement
difference misaligns every page after it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An UPDATE that replaced a chained Memo/OLE value freed the old pages before
writing the new value, so the new value landed on them. ACE frees them in
the same statement but only after the new value is written: the UPDATE that
frees them never reuses them, the statements after it do. Freeing afterwards
makes LibRed's file byte-identical to ACE's for such an UPDATE, so the
whole-file parity comparison no longer lists it as a known difference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two differences in how a row's Memo/OLE values reach their pages, both
measured against ACE with a new probe:

- ACE writes all of a row's chained values first and its single-page values
  after, each in column order. LibRed wrote them in plain column order, so a
  single-page value in an earlier column took the page before the chains.

- ACE places a compressed single-page value by its uncompressed size: it
  writes the uncompressed bytes where they would go, then the compressed row
  over their upper end, leaving the rest in the page's free space. A page
  with room for the compressed row but not the uncompressed bytes is passed
  over for a new one. LibRed placed and wrote only the compressed row.

With both, the whole-file parity comparison's long-value INSERT is
byte-identical to ACE's and leaves the known differences.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
linq2db reported 1000 - CASE WHEN s IS NULL THEN 0 ELSE s END over an
empty group coming back as the 0 literal's Int32 under a column declared
Decimal. Converting a value that holds a choice anywhere to its declared
type already fixed it; nothing tested that shape, where the choice is
nested in arithmetic and takes the literal arm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A stored query's DISTINCT, TOP and PERCENT are bits of the one MSysQueries
option row (Attribute 3), the TOP count in its Name1: ACE writes DISTINCT
TOP 2 as Flag 18 and DISTINCT TOP 25 PERCENT as 50. LibRed had three things
wrong there:

- PERCENT was never written, so a stored TOP 50 PERCENT came back as TOP 50
  and returned every row instead of half.
- DISTINCT and TOP went on two rows where ACE writes one.
- A make-table or append query's own SELECT lost its DISTINCT and TOP
  altogether, written and read, so a stored SELECT DISTINCT ... INTO copied
  every row.

Every kind of stored query now writes ACE's single row, and a make-table or
append query reads its DISTINCT, DISTINCTROW and TOP back into the SELECT.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
linq2db asked for Access SQL's WITH OWNERACCESS OPTION, which runs a query
with its owner's permissions. LibRed has no users to act for, so a
statement with it does exactly what it does without it. It is taken where
ACE takes it: at the end of every SELECT, a UNION's arms and subqueries
included, after a query's ORDER BY, and at the end of an INSERT, UPDATE or
DELETE. It is refused where ACE refuses it: before the query's ORDER BY,
twice, incomplete, or on DDL. Neither word is reserved.

A stored query keeps it, as ACE does: CREATE VIEW or CREATE PROCEDURE with
it sets bit 0x04 of the query's option row, on the DISTINCT or TOP row when
there is one, and otherwise on a row of its own. LibRed reads it back as
the clause at the end of the rebuilt statement.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ChrisJollyAU
ChrisJollyAU requested a review from a team as a code owner September 29, 2026 17:38
ChrisJollyAU and others added 9 commits September 30, 2026 01:57
The ReservedWords collection lists every keyword the grammar lexes,
reserved or not, and WITH OWNERACCESS OPTION added two.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
LibRed.Core.AccessTests has been ending partway through on CI with an
xUnit TestPipelineException, reporting only the tests that finished
before it, and it never does locally. The job kept nothing to say which
test was running. Its four test steps now pass --blame-crash as well as
--blame-hang-timeout, and a failed run uploads their TestResults folders:
the Sequence_*.xml naming the test in progress, and any dump.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
EF translates a cast from char to an integer type to ASCW, which returns
Access's signed 16-bit Integer, so a character from U+8000 up came back
negative: cast to uint, U+8000 and U+FFFF read 4294934528 and 4294967295,
in both SQL modes. The code is now widened to a Long and masked to 16 bits,
(CLNG(ASCW(x)) BAND 65535), which is the character's UTF-16 value. The
widening comes first because ACE sign-extends a 16-bit operand of BAND.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Column pruning lets a table read decode only the columns its statement
names, which is wrong under anything that looks at a whole row. DISTINCT
compares every column, so SELECT Grp FROM (SELECT DISTINCT * FROM T)
collapsed twelve rows to three; UNION, INTERSECT and EXCEPT did the same.
A set operation also matches its inputs' columns by position, so a
SELECT * FROM U under a UNION ALL read none of U's columns when the outer
query named them by the other arm's names. The inputs of both now read
every column.

A derived table with a column list, AS t(x, y), renames its input's columns
by position, and the names the outer query uses say nothing about the
names inside it. It is left unpruned.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A child INSERT or UPDATE inside a transaction checks that its parent
exists, and reading the parent writes no page, so the commit's page
conflict check could not see it: another connection could delete the
parent and commit, and both transactions would succeed, leaving a child
referencing nothing. The check is now registered with the transaction and
made again when it commits. It only holds the commit to the parent while
the parent is still needed: while the relationship survives and a child row
still holds the key, so deleting or moving the child, or dropping the
relationship, releases it. The relationship is identified by its tables'
definition pages and its columns' ids, so renaming a table or column in the
same transaction does not lose it. The transaction holds one check per
relationship and key however many rows relied on it, and a savepoint
rollback frees the keys of the checks it discards.

ALTER TABLE ... ADD FOREIGN KEY never looked at the rows already in the
table. It is now refused where ACE refuses it, with ACE's message: a child
with no parent, or a composite key partly null, whatever the ON DELETE
action. ACE then holds both tables exclusively until the transaction ends;
LibRed takes no table locks, so the same check is made again at commit.

A violation raised by an UPDATE now says so, not "INSERT into".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TableCreator.AddForeignKey had a branch for a self-reference that repeated
the rest of the method. What differs is only where the parent's definition
page, key index and incoming block number come from: the child itself, with
the incoming block taking the next free number above the outgoing one.
Those three are chosen first and everything after is shared, with the same
calls in the same order, so the bytes written are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A cascade rewrote a child with its own partial copy of the UPDATE checks:
no null-key check, so ON DELETE SET NULL could null a primary-key column,
and none of the child's own relationships, so a cascade stopped one level
down and ON DELETE SET NULL onto a key other rows referenced left them
pointing at nothing. The UPDATE loop's per-row write is now UpdateRow, and
SetChildKey hands the rewritten child to it: a cascaded row is checked
exactly as a directly updated one, and a key it changes applies its own ON
UPDATE rules, cascading on down or refused. The parent's cascade runs once
the parent row is written, so the child's foreign-key check finds the new
key.

The commit-time re-check only ran from the child's end. A DELETE, or a
parent key change, that found no child writes nothing another connection's
child insert conflicts with, so both could commit. It now registers the
same condition for the key it removed, held once with the child side's.

SELECT INTO writes through InsertNewRow like any insert. A deleted row, a
cascade's and a complex column's values go through one DeleteRow; the
MATCH FULL key extraction exists once, shared with ADD FOREIGN KEY's check
of existing rows; the "related records" refusal is one throw. CHECK
expressions, relationships in both directions, their tables and column
positions, and an outer-join UPDATE's defaults are resolved once per
statement rather than once per row.

PageChannel.PageCount inside a transaction is never below the file's own
page count, so a transaction can read what another connection committed on
pages appended since it began, which the commit-time check has to read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CLAUDE.md and AGENTS.md: LibRed.Ado.Tests and LibRed.EFCore.Tests use no
driver and only run in the Windows ACE job; four functional-suite shards,
each run up to three times; pass-lists exist for the two ACE 2010 x86 legs
only, and auto_commit extends them after a successful PR run; both
connection env vars; fixed order and the en-US culture lock apply to the
EF functional suites; two test projects are MSTest; dotnet test builds
only net11.0; the Northwind.accdb and LibRed.Shared consumers; page locks
are process-local. AGENTS.md also catches up on LibRed.Core.AccessTests,
the AccessTests split, the deferred-write overlay and the Core type guard.

The LibRed README's "Not yet" lists only open work: stored action queries
other than crosstab, pass-through and UNION-kind are written and run,
HAVING views work, the 1:1 flag and ValidationRule are read, a column-level
CHECK is dropped, RENAME INDEX throws. The referential-integrity and SQL
statement descriptions, LOG(base, x), the SQL-mode readers and the Jet3Format
stub are corrected. transactions.md gains the parent-side commit check and
marks the undo-log design as superseded; the package readme, format index,
benchmark and JetLockTrace docs and copilot-instructions are corrected too.

LibRed.Engine.Tests no longer suppresses CA1416: it has no ACE code, and the
analyzer is what keeps a cross-platform suite that way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… changes

The SID keystream is folded from page 0's header (the 0x42 field and the
creation date), so a database is no longer created from one baked
date/SID pair: it is dated now and its SIDs are the default workgroup
accounts XOR the stream that header gives. The accounts carry their real
roles - admin owns MSysDb and new objects, Engine the system tables,
Users holds the grants, Creator the inheritable placeholder.

Anything that changes the field changes the stream, so the stored SIDs
are re-masked with it: the legacy Jet password on an .mdb, and an
.accdb's whole-file encryption, whose 0x42 field Access fills with the
database key's low byte. The .accdb re-mask runs in a transaction the
page copy reads through and then rolls back, so the original file is
untouched unless the replacement lands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant