LibRed: audit and whole-file ACE-parity fixes, collation-true text, faster engine, Access SQL additions (11.0.0-alpha.4) - #303
Open
ChrisJollyAU wants to merge 87 commits into
Open
ChrisJollyAU wants to merge 87 commits into
ChrisJollyAU wants to merge 87 commits into
Conversation
- Bump to alpha.4 after the alpha.3 release - Remove the MyGet feeds; CI builds stay available as the nupkgs artifact - Convert a char to a number by its code point (AscW), not by parsing its text - Accept any DbConnection in UseLibRed Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A parameter takes the type of what it is compared with, as in ACE: a number
against a text column is compared as text, text against a number column is
read as a number. Since alpha.3 a number parameter against a text column read
the column as a number instead, and threw on the first value that was not one
('' or 'abc'), where ACE and alpha.2 answered.
Measured against ACE for =, IN and BETWEEN, with number, double and Boolean
parameters. A number literal is unchanged: ACE refuses it outright, LibRed
still reads the column as a number.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GetValue returned 1899-12-30 for a serial of 0, while GetDateTime and GetFieldValue<DateTime> turned it into default(DateTime). All now return the stored date, as ACE's OLE DB reader does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The index key encoder reads a GUID from text, as the row codec already did;
comparing or inserting a string GUID on an indexed column threw
- Both read ACE's {guid {…}} form, in text and as a literal; the grammar now
lexes it (parser regenerated)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Standard SQL predicates that ACE rejects, added as extensions. Both are never Null: a truth test treats Null as neither True nor False, and DISTINCT FROM compares with Null taken as a value. Parser regenerated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ACE stores a declared parameter's name as written, brackets included
('[@firstname]'). LibRed bracketed it again on read, so an ACE-written
parameterised query rebuilt as PARAMETERS [[@firstname]] and would not parse,
and reported a name that could not be bound.
- Read: keep a bracketed name as it stands; report and bind by the name inside
the brackets
- Write: store the name as ACE does, brackets kept and a bare @ dropped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0x11 is ACE's BigBinary, which its DataTypes list names but no one had tied to a code: BIGBINARY(n) in DDL, up to 4000 bytes, bare BIGBINARY taking the maximum. Measured against ACE: - The descriptor is VARBINARY's apart from the type byte; no format raise. - The value is always inline, never a long value, so it counts against the 4060-byte record cap: a binary bounded by a single data page. - In queries and indexes ACE treats it exactly as an OLE Object: no sort, group, distinct, aggregate, join, union or index, same messages. - MSysAccessObjects.Data, the only place it had been seen, is a fixed 3992-byte column of this type. LibRed now creates, writes and reports it as a binary type with a 4000 cap, and refuses it in an index as it refuses OLE. CreateFormat is filled in for every DataTypes row, BigBinary included. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A scalar subquery declares its one column's type, found by describing its plan, so IIF/CASE/COALESCE over it widen with it. - Only a NULL literal is skipped when unifying a choice's arms; any other untyped arm leaves the choice untyped, rather than letting the typed arms declare alone and converting every value to that. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each was measured against ACE first where ACE has an answer, and the format spec corrected where the measurement disagreed with it. Corruption: - CREATE INDEX after DROP INDEX could take another object's usage-map row, computing it from a layout formula a drop invalidates. ACE never reuses an orphaned row, and neither do we now. - An index root split left every other open handle on the old root, so seeks missed rows and a duplicate PK could get in. - The page cache and lock manager keyed on the lowercased full path, merging two files on a case-sensitive system and splitting one file reached two ways. One FileIdentity.Key now serves both registries. - A self-cascading UPDATE wrote stale snapshot values back over a key its own cascade had already moved. - DROP TABLE and DROP COLUMN ignored complex columns; ALTER COLUMN on a table owning one is refused rather than run, since the rebuild carries none of its three links. - The password operations loaded the whole file, truncated it and wrote it back. They now take a database the caller opened exclusively, as ACE requires of ALTER DATABASE PASSWORD, stream pages through a single page-sized buffer into a sibling copy and replace the database with it, so a change is atomic and never writes plaintext to disk. Wrong data: - A range seek on a descending index returned a fraction of its rows. - Deleting from an ACE-written leaf whose shared prefix reaches into the row pointer threw. - A Yes/No index key is flag + value, not a bare byte, and -1 keys true. - Negative-zero Decimal lost its sign on rewrite, orphaning its entry. - Text length was checked against compressed bytes, so a TEXT(5) took eight characters where ACE refuses six. - A null PK or DISALLOW NULL key was accepted, and an AutoNumber could be updated; ACE refuses every update of one. - Dates before year 100 were changed silently. 0100-01-01 is ACE's floor; below it is now refused, naming the column. - Owner/ACE SIDs were hard-coded with two different files' masks. The mask is recoverable from MSysObjects' own owner, so both writers now derive the accounts per database. Leaks, parity and transactions: - DROP INDEX freed only the root page and left the usage-map record behind; both drops now retire it, byte-identical to ACE. - Space freed mid-table never set the page's free-map bit, so it was never offered to another insert. - A freed page outside a map's range vanished instead of being reported. - The row's leading count is the TDEF 0x29 id high-water, a dead id's null-bitmap bit is clear in an inserted row, and the variable trailer survives the last variable column being dropped. - A failed statement inside a nested transaction left its savepoint frame open, after which every COMMIT threw. Tests live in files named for their subject rather than for the audit, and the legacy Jet 4 fixture is a database DatabaseCreator builds rather than a header followed by random bytes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The rest of the audit's findings, and then a pass over the shape they share rather than the instances. Findings: - ALTER COLUMN on a table owning a complex column carries it across instead of refusing. Five links, not the three the audit named: the descriptor's 0x0B, the flat tables, the catalog rows, the in-row complex ids, and the 0x1C counter a recreated table restarts at zero. The property blob and MSysACEs grants come across verbatim too. - LibRed demanded an indexed parent key for every relationship, where ACE needs one only to enforce. Measured over three files: every enforced relationship has an index on both sides, no unenforced one does. An unenforced relationship is now catalog rows and nothing else. - Foreign-key write skew. A transaction records the conditions its writes depend on and they are re-checked under the publication gate. Semantic, not page-based, so an unrelated row on the same page cannot produce a false conflict. - A relocated row that outgrows its new page moves again rather than throwing, byte-identical to ACE, tombstone included. Both write paths had their own copy of the limitation. - ADD COLUMN measured the fixed region as the sum of the live columns' lengths, ignoring the hole a dropped one leaves; DROP INDEX could renumber a data block another logical index still named; a memo- indexed UPDATE threw; SET NULL onto a row the same DELETE was removing left the delete holding a stale key; an update that freed a memo then failed had no undo outside a transaction; a rename left a complex column's catalog row naming a column that no longer existed; ACE moves DateUpdate on an ALTER and LibRed never did; an unknown version byte threw on every rollback; a handle kept a stale Format after another raised it; a failed close left the connection pointing at a disposed database. Measured and found not to be defects: the LvProp owner record's second field is always zero; a leading U+FEFF is indistinguishable from the compression marker and Access loses it the same way; ToOADate quantises to milliseconds on the way in, so a DateTime key round-trips exactly. Readers no longer paper over what only this engine could have written: a walked index checks its keys never go backwards, a page declaring more than 255 rows is refused, the TDEF reader stops silently dropping an index whose column or data block is missing, an unresolvable relocation pointer is reported, and an AutoNumber increment of 0 — which ACE refuses outright — is refused on write and reported on read. Then the shape behind the memo cluster. A row had two representations in one object?[] — the values a caller holds and the descriptors the record carries — and materialising in place made the caller's array change meaning halfway through a write. Every memo defect was the wrong one reaching a consumer that wanted the other. Insert and Update now keep the logical row and encode from a separate storage copy, which also removes the two conditional clones that existed to work around the mutation; RowDecoder refuses to decode a long value without the reader that resolves it, and the descriptor-only operations are static so there is no mode to get wrong. The same shape turned up a live bug: TableCreator and TdefBuilder both defaulted a missing collating order to General-Legacy, and only 2 of 18 call sites passed one. On a General (v1) database, retyping a column wrote it back as v0 — a different key encoding in a file whose other columns use v1. It survived because every fixture here is v0, so the wrong default was the right answer everywhere it was exercised. The order is now required, which turned the defect into 16 compiler errors. Also folded together: five bespoke catalog-row updaters into one beside DeleteCatalogRows, twelve copies of the distinct-real-index expression into TableDef.RealIndexes, and the transaction-state clear repeated in three places. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The four items the audit left needing a measurement, and the parity gaps chasing them turned up. Each was measured against ACE before anything changed, and the spec updated where the measurement went past it. The probes: - An emptied index leaf leaves the tree. ACE closes the leaf chain over it, drops its separator from the parent node and releases the page; LibRed kept it linked with a live separator pointing at no keys. The node is left with no entries and only its child-tail, which is a valid node rather than one to collapse. A leaf that IS the root, or that is its parent's only remaining child, stays: an empty leaf is valid and neither shape has anywhere to go. - A node page's prev/next are 0 in both, as LibRed already wrote them. - Currency rounds half-to-even, on both routes into a Currency - the CCur function and storing into a CURRENCY column - which is what LibRed already did. The only doubles that can settle a 4-decimal midpoint are the odd multiples of 1/32; anything else is already off the midpoint and rounds the same way under every rule. - An empty Memo is the full 12-byte inline descriptor, 00 00 00 80 then eight zero bytes, and an empty LONGBINARY likewise. Identical to LibRed's, so the zero-width slot its reader refuses is a form ACE never writes and the reader's strictness is right. Reclaiming an emptied page, which the empty-leaf byte diff exposed: - A data page whose last live row is deleted is stamped 0x09, cleared from both of the table's maps and released. So 0x09 is not confined to packed long-value pages - which is exactly what page-09 recorded as an unaccounted-for case - and the two are told apart by the owner at 0x04, an LVAL signature against the owning TDEF. The spec file is renamed for the page type rather than for one of its causes. - A table never gives back its FIRST data page. Deleting every row of a 35-page table releases 34 and leaves that one at 0x01, empty, still in both maps; it is kept even while later pages hold live rows. - A page emptied by reclaiming a hidden relocation target is not released - ACE keeps it as an ordinary data page. Only deleting a live row releases the page it empties. - A released index leaf keeps its 0x04 type byte and is never stamped, so a 0x09 page is always a data page. And the usage-map row order, the open half of the audit's first finding. Rows after the table's own two follow the order the CREATE TABLE statement declares them: a long-value column takes two where its column is written, an index one where its constraint is. LibRed always used the inline order, so a constraint written after a long-value column - the shape a migration emits - took a row ACE gives to the column. The constraint's position now travels from the parser, which already tracked the token offsets for self-reference resolution, through to TableCreator. Index data blocks are unaffected and keep their PK-then-unique-then-FK order, so a block's ordinal and its map row are independent. The other half of that finding - CREATE INDEX reusing an orphaned usage-map row - was already closed; the audit's note had gone stale. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CVar was a pass-through. Measured against ACE over OLE DB, a Variant keeps its own type while an expression uses it and through a derived table, but is written out as text (as CStr writes it) by a result, a scalar subquery, a union and a make-table query. A choice whose values differ in kind - text beside anything else, or a Variant beside a non-Variant - is mixed: text wherever it is written out, derived tables included. - Both sort, group and take Min/Max/First/Last as their text. - As an operand either counts as a Double; only + beside a Variant, a mixed value or text keeps a Variant. - In a choice, a date beside a number is a date, a Boolean beside one a Long; SWITCH and CHOOSE are typed as IIF is. - A value holding a choice anywhere is converted to its declared type. - CVar(Null) stays untyped, so EFCore.Jet's projected Nulls leave a union typed from its other arm, where ACE makes it text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…found Sixteen findings: where the same value is derived twice, and where one job is written out along several paths. Every "redundant" claim was checked against docs/format/ first, because the format sometimes requires the second pass - one finding turned out that way. Two were live bugs the copies had hidden: - GROUP BY compared its keys with CLR equality for anything that was not text, so a LONG 1 and a DOUBLE 1.0 arriving under one key were two groups where ACE returns one, and byte-array keys - compared by reference - gave every row a group of its own. Measured against ACE through an IIF with differently-typed arms, a UNION of a LONG and a DOUBLE column, and a DOUBLE against a CURRENCY. The fold is within a kind, because the hash partitions by kind: text keeps the invariant comparison its hash is built on rather than taking CompareForSort's database collation, which disagrees with it over ss and where the collation key is far too heavy to build per row. - IndexKeyDecoder's FixedKeySize knew neither Complex nor FixedPoint and returned -1, which its caller reads as a lossy key to stop at, so a DECIMAL key column ended the decode and nulled itself and every later column of a composite key. The encoder's copy knew both. One table in Formats now, so they cannot drift again. The walks and writers, each written once: - TdefRegions.Of replaces five hand-written walks over the TDEF's variable-length regions. Only IndexWriter's bounded its steps; the three in TableCreator advanced raw offsets, so a file-sourced count or name length could carry a DDL write past the buffer - or, on an overflow, back inside it at the wrong index block. They are bounded now. - UsageMapBits.Append replaces four copies of the bitmap-to-page-number loop. The validation difference between them is deliberate and stays: a read feeds its numbers straight into page reads, so one outside the file is corruption, while a rewrite must keep bits it cannot currently represent and widen the record instead. Rejecting there had already failed an ordinary DROP TABLE on a real file. - CatalogWriter writes an object's MSysObjects row and its MSysACEs pair once, for tables, queries and relationships alike; each caller still owns the type, container, flags and masks that genuinely vary. DatabaseCreator keeps its own, populating both tables before either has an index or a catalog to resolve a name through. - Rc4Cipher.Apply serves the legacy Jet codec and the Office-Standard one, and Agile's spin derivation is shared by its open and create paths. - JetCatalog.RequireTable and TableDef.RequireColumn replace twenty copies of the system-table lookup-or-throw and thirteen column ones, along with four private aliases for the same lookup. ObjectProperties also returns the Id index now, so five LvProp read-modify-write methods lost four identical opening lines each. The recomputes: - INSERT read a data page on every row to re-derive the fixed-region length. The region may never shrink below what existing rows carry, and a retired column id leaves a hole the live descriptors cannot see - but only then, so the read is skipped when 0x29 says no id was ever retired. - A row's long values were freed by re-reading the whole definition per column; the long-value map pointers cannot move under an inserter, so they are read once. - UPDATE parsed one row's layout twice, a column reference was resolved by name for every row, and Format() re-parsed its format string on every call. - Two callers materialised every page a table owns to look at one of them, which made a bulk load quadratic; both ends of the owned-pages bitmap now share one scan. - The property blob was parsed once per column, for every table, on catalog load; its five accessors take the parsed list instead. AggregateSurfaceTests binds the three places that answer the aggregate surface, and AggregateResultType's fall-through returns null rather than the argument's type, so a new aggregate reaching it fails the test instead of being declared as whatever it was fed. Left open: there is still no parse or plan cache, the one finding with real cost behind it. Plans embed live IndexDefs and Invalidate with markChanged false breaks the obvious cache key, so it needs its own change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both were already covered in prose and enforced by nothing. - Reading repository files through git. Read-only git is pre-allowed, so it was the way to spell a read the read rule did not match and skip its prompt. Reviewing history is untouched. - LibRed.Core.AccessTests and LibRed.Engine.AccessTests are about 7 and 1.5 minutes, and running either whole and in the foreground is dead session time. They now need a --filter naming the tests that cover the change, or run_in_background. The other suites are left alone - Ado and EFCore are two seconds each, and friction on a rule with no payoff is what creates pressure to work around it. The first pattern's own first draft allowed arbitrary text between the two words, which made the commit describing it match itself; only option-like tokens may sit in the gap now. Not added: a block on the functional suites, which CLAUDE.md also forbids in prose. That rule has never been broken, because it is binary - a project name is either on the list or it is not. The rule that does get broken asks for a judgement about which tests cover a change, and no hook reaches that. Requiring a --filter is the closest proxy, since it forces the claim to be named. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A comparison gave a different answer depending on whether the column happened to be indexed. `S = 1` on a text column compares as a number (ComparisonOperatorTests), so 'abc' is a type mismatch - but with an index on S the seek reached IndexKeyEncoder, which casts the value by the column's type and raised "Unable to cast Int32 to String" out of the storage layer instead. The encoder is right to demand a typed value; it is shared with the row-writing path, where the type is guaranteed. The engine was handing it a value of another kind. Worse on the parameter path, which is the one EF uses: `S = @p` with a numeric parameter answers 0 rows unindexed - a parameter compared with text takes the text's type, verified against ACE, which answers 0 too - and crashed with the same cast error when S was indexed. So the index changed the result, not just the error. TryGetSeekKey settles what a seek may be keyed by. An index answers only in its column's own kind, because that is how its keys are encoded and ordered, so a comparison happening in another kind has no key range to seek: `S = 1` matches ' 1 ', '1.0' and '+1' as well as '1', which sit nowhere near each other in a text index. The one cross-kind case that IS seekable is a parameter against text, which CompareAsKinds converts to text before comparing; the seek converts it the same way and so asks the index the same question. Anything else falls back to a scan - safe because the single-table seek keeps the FilterNode it was planned under and the index-nested-loop join keeps its ON whole as the residual, so only the reading strategy changes, never the answer. Measured against ACE before and after: the indexed and unindexed forms now agree everywhere, and `S = @p` returns the 0 rows ACE returns. The literal coercion policy is untouched - text against a number still reads as the number it is, as SQL Server does and as ComparisonOperatorTests records, where ACE refuses the comparison outright. IndexSeekKindTests pins indexed against unindexed for both literals and parameters. Reverting TryGetSeekKey to its old behaviour fails five of them and leaves the same-kind and NULL controls passing, which is how they were checked. Two of the five were not in the original report: a numeric indexed column against a text literal crashed the same way, and so did a cross-kind key in an index-nested-loop join. Also drops the last reference to a probe that should not have been committed with the audit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
LibRed.Core, LibRed.Sql, LibRed.Engine and LibRed.Ado depend on no EF Core package, so nothing about them requires the version the rest of the repository is pinned to. They now multi-target LibRedTargetFrameworks, net10.0 plus whatever JetTargetFramework is, which lets a consumer on the current LTS use the format reader/writer, the SQL front end, the engine and the ADO.NET surface without moving first. LibRed.EFCore is deliberately not in that list - it sits on EF Core 11, which is net11.0-only - and neither is anything under src/EFCore.Jet. Every test project stays single-target net11.0, which is a choice rather than an oversight and is written down in CLAUDE.md so it does not get helpfully undone. The net10.0 leg is compiled and never run: multi- targeting the suites would double every run locally and across the five CI platforms, and would need the .NET 10 runtime on each runner, which the SDK global.json pins does not carry. Compiling it needs only the reference pack, which restore fetches, so CI needs no change at all. Checked rather than assumed, once: the whole solution builds both legs with no warnings, so nothing in the code was relying on a C# 14 feature that net10.0 cannot have despite LangVersion being preview; and with LibRed.Core.Tests and LibRed.Engine.Tests temporarily multi-targeted, both passed identically on net10.0 - 660 and 4137. Those two csproj changes were then reverted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- The 128-byte page-0 header mask is the RC4 keystream of the fixed key C7 DA 39 6B, verified over all 128 bytes. - An on-disk SID is its workgroup SID XOR'd with a per-file keystream from its first byte; the 2-byte short-SID mask is that keystream's first two bytes, and the 102-byte SIDs Access adds use the same keystream. - How the keystream derives from the creation date is recorded as not known, with the derivations ruled out, replacing the shared-PRNG guess. The workgroup Admins SID is 102 bytes, not 98. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
[Table]![Column] is Table.Column in every clause, and an unaliased one is named as ACE names it, text as written. A longer chain such as [Forms]![frmMenu]![txtCity] is one name, reachable as a parameter the query declares; a period between parts counts as a bang, as ACE binds them. A stored parameter declared that way now reads back as that name rather than with only its outer brackets stripped. Yes/On and No/Off are True and False, winning over a column of that name unless it is qualified. Any word, reserved or not, names a column after a period or bang, and an unbracketed name may use any script's letters. Every stored query in the example corpus that uses a bang now parses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… blob The flag byte is a bit field, not a DDL boolean: Access writes 0x80 on its own account, and every stored query of one example database carries it. Reading refused anything but 0x00/0x01, which would stop a table carrying such a property from loading; the byte is now kept whole and written back. A value block's type says what owns it, and an index's block (0x0002, by mdbtools' account) is named for the index, usually its column's name. Every block was read as a column's and written back as one, so a column's Required or DefaultValue could come from an index, and a DROP, RENAME or ALTER of the column moved the index's properties with it. Each property now keeps its block type, and the column and table accessors match only their own block. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Column flag 0x10 is set on every column of MSysObjects, MSysACEs, MSysQueries, MSysRelationships and MSysComplexColumns and on no other, and 0x20 on the two security-identifier columns, MSysObjects.Owner and MSysACEs.SID. The spec listed neither and said the unlisted bits were zero in every file. The extended flags gain the attachment value column's 0x10, and the 0x04/0x08 a complex column's flat table sets. A new database gave the MSysComplexType_* templates the catalog flag, which Access never does; the flag was also what zeroed their 0x09. That is now its own switch, and the attachment template carries its extended flag, so every system column a new database writes matches Access's byte for byte. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…file diffing A new probe runs the same statements through ACE and LibRed on copies of one DAO-created database and compares the files byte for byte, both at the end of one session and statement by statement. What it found, each now verified and fixed: - A table's property blob is stored like any long value: inline up to 64 bytes, on a page above that. LibRed always used a page. - A new object's MSysACEs rows are its container's inheritable grants, the Creator's becoming the owner's, so the masks follow the database rather than the object class the three constant pairs assumed. - A foreign key's index root is allocated after its relationship's rows. - A statement's deletes run in the order the rows were found, which decides the bytes left in the space they free. - An index built over rows writes each leaf at the prefix its filling reached, not the largest its keys share. - An UPDATE moving an entry lowers the index's total and holds its unique count to it, on an index built over rows. - A long-value page leaves the free map with 257 bytes free or fewer. Both measurements now come out identical: all 104 statements, and the whole run's 74 pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… pages The whole-file probe now also runs a thousand-row INSERT ... SELECT, a 253-column table and long values chained across pages. What it found about index splits, measured against ACE and now written the same way: - A root that splits keeps its page: both halves move to new pages, left first, and the root becomes the node over them. - A full page is compressed in place before it splits, and the split is made over that image; the left half keeps its prefix unless the new entry became its first, when it is written whole. - The cut keeps every old entry that starts before the page's byte midpoint, the new entry joining its side; a new first entry halves the entries by count. - Nodes fill, compress and split as leaves do, link to their siblings, and split at the right edge when the new separator is last. - Bytes past each half's live end match ACE's, including the promoted entry a node split leaves behind. CREATE INDEX over existing rows writes the tree sequential inserts leave, so the bulk build now follows the right edge of that tree instead of building node levels bottom up, and no longer moves the index's root. ADD COLUMN writes descriptor 0x09 as ACE does: it is DAO's OrdinalPosition, which ADD COLUMN compacts to ranks, and an added fixed column's variable-table index counts dropped variable columns too. Page 0's 0x30 and 0x34 are bounded by the largest possible page, not by the file's length, as the spec had said. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A profile of a 10,000-row scan put most of its time in three places, none of them the format: - Plain UTF-16 text went through Encoding.Unicode, which counts and then converts with full validation - about a fifth of the scan. Without a surrogate code unit and at an even length the bytes ARE the string's chars, so they are copied; anything the decoder could treat differently (a lone surrogate, an odd trailing byte) still goes through it. Compressed text with no switch byte is one Latin-1 run, which is what the per-character loop produced for it; the mixed form decodes into a stack buffer rather than a StringBuilder. - CURRENCY was divided by 10000m per value. Division returns the smallest scale that holds the quotient, so the same decimal - value and scale, which shows in its text - is built from the integer by dropping the four places' trailing zeros. CurrencyDecodeTests holds the two equal bit for bit across 600,000 values and the edges. - RowDecoder enumerated its columns through the interface, and HasVariableSection re-scanned them, for every row. It walks an array now and answers the variable-section question from the lowest variable column id; Boolean values come from two shared boxes. The raw scan goes from 4,084 to 2,895 us and 5,385 to 4,761 KB. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fbbed31 memoised where each column reference resolves inside EvalScope, on the premise that only the row moves between rows. That held for the join paths, which rebind one scope, but projection, WHERE, sort keys, group keys and aggregates built a fresh scope and evaluator per row - so each row also built a dictionary, used it once and dropped it. A scan allocated 11,577 KB where it had allocated 9,388, and scan, GROUP BY and the unindexed sort all slowed; found by bisecting on allocation, which is exact where timings are not. Those paths now make one evaluator and rebind it per row, as ExecuteJoin already did. Filter and projection make theirs inside the iterator, so each enumeration of a result has its own. The memo then pays off everywhere, and the per-row scope and evaluator go too: a scan now allocates 5,871 KB. A projection item that is just a column of its input is located once when the projection is planned and read from the row, rather than resolved by name on every row; it still takes the item's conversions. A name that is not exactly one input column (outer, niladic, ambiguous) stays with the evaluator, so its error still comes from the row that raises it. EvalScope.Locate is the one matching rule for both. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… fixes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decoding was most of a scan's cost, and every read decoded every column: a three-column SELECT over a ten-column table built all ten, and a GROUP BY held all ten of every row alive. A row decoder now takes a mask, and a column outside it reads as null - so the mask is only ever set where nothing can look at that column. Queries: ColumnPruning runs after index selection on every plan that is executed - top-level statements, INSERT ... SELECT and SELECT INTO sources, derived and lateral sources, subqueries - and tells each scan and seek which names to decode. The names come from every column reference anywhere in the plan, subqueries included (their correlated references read the outer tables), by name alone, so a table decodes anything a reference could mean. They are found by walking every public property of the plan and AST records, so a clause added to a node later cannot be missed; a value the walk does not know turns pruning off. A read is pruned only below a projection or aggregate without a star, because SELECT *, set operations and DISTINCT pass whole rows to the output, and DISTINCTROW dedupes on the underlying rows - which is also why nothing below a node the pass does not know is pruned. UPDATE and DELETE: the read that chooses the rows decodes what the statement names anywhere, and CompleteRows then reads each row about to be written in full, into the array its joined rows share, before any of it is used - the rewrite, its constraint and cascade checks and its index moves all see every value, as before. The parent check every child INSERT makes decodes the key alone; a cascade finds its children by key and reads the matches whole; complex-value cleanup reads the link and the index keys. ColumnPruningTests covers each place a column can be named, each shape that passes a row through whole, and each write. Breaking the write-side completion fails five of them, and pruning below DISTINCTROW two (it collapses 12 rows to 3). Against ACE at 10,000 rows: scan 2,707 -> 1,652 us, GROUP BY 8,372 -> 4,436, hash join 2,988 -> 1,681, unindexed TOP 10 4,637 -> 3,317. The walk reads properties through compiled getters: through PropertyInfo it cost a primary-key lookup 4 us of its 22. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…or a match The reads that find a row by one or two columns decoded every column of every row they passed: an index back-fill decoded whole rows to build keys from a few, and each MSysObjects lookup resolved every object's LvProp blob - a long value - on the way to the one it wanted. Table gains DecodeOnly, a mask for a read that looks at nothing else, and RowsWhere, which searches on the key alone and reads each match whole, because a row looked up by key is usually about to be rewritten or deleted. The index back-fill and its duplicate check, the narrowing ALTER COLUMN check, and the object-name, permission and next-id scans decode only what they compare. The LvProp edits, relationship renames and deletes, catalog-row updates and deletes, permission reads and MarkAsSystemTable find their rows through RowsWhere (most via the existing RowsKeyed), as do a cascade's child search and ReadComplexValues, which no longer decodes other records' attachments to find one record's. A back-fill over 2,000 rows allocates 2,045 KB where it allocated 2,662; the catalog scans are too small to show on the benchmark corpus's three tables. ALTER COLUMN retypes, rebuilds and an AutoNumber ADD COLUMN rewrite every row and still read them whole. LibRed.Core.AccessTests pass, so the DDL still writes what ACE writes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A name only has to be free among the objects of its MSysObjects container - the rule the unique (ParentId, Name) index states. A table or query clashes with a table, query or linked table (the Tables container) and with nothing else; measured against ACE for every object kind. CREATE TABLE and SELECT INTO now share the rename check, which tests the container rather than the table and query types, so a query's or linked table's name is refused as "Table 'X' already exists." instead of surfacing as an index violation. CREATE VIEW no longer refuses the names of relationships, forms, reports, macros, modules, database documents and containers, and reports a clash in ACE's words. The spec gains a table of the MSysObjects object kinds - their types, containers and filled columns - and the name rule, and drops the claim that tables and queries sit in different containers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… word Access keeps each object's Name AutoCorrect map twice: in the NameMap column of its MSysNameMap row and in a NameMap property in its LvProp, in two layouts. NameMap decodes and encodes both, carrying every field - the uninitialised bytes Access leaves, a missing terminator, the property's end record - so a map read and written back is byte-identical; every map in the corpus round-trips. JetDatabase reads the rows and a named object's property, and replaces either; a row is never added, its Id not being understood. The engine maintains neither copy, through ACE or LibRed. MSysNameMap.NameMap has an owned-pages map and no free-pages map, as Access writes it (and MSysAccessXML.LValue); every value on such a column takes a page of its own. The long-value writer refused the column outright, so no map over 64 bytes could be stored; it now writes one the way Access does, and ACE reads it back and keeps it through a compact. The page type is the 16-bit word at offset 0, not a type byte followed by a constant flags byte: every page Jet and ACE write carries 0x01 in the high byte, and ACE tests the whole word - it reads a table's owned page as rows only for 0x0101, 0x0103, 0x0104, 0x0106, 0x0107 and 0x0109, and skips everything else, 0x0001 and 0x0201 included. PageType is now that word and every check compares it whole; what LibRed writes is unchanged. The spec gains MSysNameMap, both map layouts, the owned-only long-value column, and the page type as a word, with the field at 0x02 named for what it is: the page's free-space count. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A row whose relocation pointer names a page the file does not hold still counts - COUNT(*) counts slots - but reading it fails "Unrecognized database format", and ACE marks the reading user's commit slot 01 00, after which every open fails until a repair. Unlike a page missing from a table's owned-pages map, which the scan skips silently. LibRed refuses the read too, and leaves the commit slot alone, having none of its own. Page 0's commit-byte section now lists the two measured occasions for 01 00, and three page types the move to 16-bit types missed are corrected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A table keeps the history of all its append-only memos in one hidden complex column, VersionHistory_F5F8918F-0A3F-4DA9-AE71-184EE5012880, whose template and flat table carry a value column per memo. LibRed dropped the memo and left all of it behind. It now does what ACE does: while another append-only memo remains, only the dropped memo's value column goes, from the template and the flat table; when it was the last, the hidden column and its index, its MSysComplexColumns row, the flat table and the template go too - their MSysACEs rows left behind, as ACE leaves them - with the table-level AppendOnly property, and the table's complex-column flag unless another complex column remains. The table-level property leaves in the same rewrite as the memo's own, as ACE writes it once. The spec gains the version-history structure and both drop rules. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ALTER COLUMN to or from Memo/OLE fell back to a logical rebuild: the table dropped and recreated, every row re-inserted, every index rebuilt - the primary key moved to the front - where ACE edits the column in place and rebuilds only the indexes over it. It now takes the same in-place edit and re-lay as every other retype, plus the long-value side of ADD and DROP COLUMN: a column becoming a long value gets its map entry and records as ADD COLUMN places them, and each value is stored as an insert stores it; a column ceasing to be one has its maps retired and its pages released as DROP COLUMN does. The rebuild and the helpers only it used are gone. The re-lay shared by every retype also differed from ACE on rows whose old value was NULL: the dead id's null bit is now carried over from the old row rather than set, and a NULL fixed target's new slot keeps the old record's bytes rather than zeros. Text to Memo, Memo to Text, Memo to Memo and a fixed retype over a NULL now leave the file ACE leaves, bar MSysObjects' DateUpdate and, for the Memo re-declaration, stale bytes below the live records on the usage-map page. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With every ALTER now an edit in place, nothing builds a column from an existing descriptor any more. The raw-descriptor passthrough on ColumnSpec, BuildColumnDescriptor's start-from-the-original path, the two flag masks it needed, and TdefBuilder.Build's complex-counter parameter had the rebuild as their only caller, and go. The comments and the spec sentence that described unmodelled descriptor bytes surviving "through a rebuild" now say why they survive: the descriptor is never re-emitted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A TEXT(n) to TEXT(m) change edited the descriptor's length and nothing else. ACE has no such path: widening or narrowing, indexed or not, it burns a fresh column id, re-lays every row and rebuilds the indexes over the column, as for any other retype - which the spec already said. The length change now takes that retype, and a value too wide for the new declaration is refused by the re-lay's own width check, with ACE's message. With more than one index over the retyped column, LibRed rebuilt each in its own slot. ACE rebuilds them in logical-block (name) order, the primary key included, allocating their roots in that order, and hands them back the real-index slots they held between them in that order; the logical blocks over them hand their numbers round the same way, which is distinct from the data ordinal once two logical indexes share a real one. A column becoming a Memo gets its map records ahead of the rebuilds' on the usage-map page. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The migrations SQL generators wrapped a column's type change, narrowing or NOT NULL in DROP INDEX / CREATE INDEX for every index over it - SQL Server's workaround, since SQL Server refuses to alter an indexed column. Jet does not: its ALTER COLUMN rebuilds the indexes over the column itself, keeping each one's name, columns, order and flags, the primary key included, and a same-type NOT NULL leaves them untouched. A computed-column change needs no index handling either, since Jet cannot usefully index a calculated column. That leaves nothing for GetIndexesToRebuild, DropIndexes, CreateIndexes or the operation list they read, and all of it goes, from the shared generator and LibRed's copy alike. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
OLE DB runs ACE in ANSI-92 mode, and before Access version 2311 (build 16.0.17029) ACE refused an ANSI-92 reserved word that Access's own SQL does not use as a name - even qualified, K.Except or K!Except - with reserved error -1001, which has no message. 2311 fixed it, so the current engine takes Full, Then, Cross, End, Fetch, Next, Rows, Only, Intersect, Except, Restrict, Temporary and Language like any other name, and LibRed follows it. The ACE 2016 redistributable CI installs and ACE 2010 both predate the fix, so there those cases now skip with that reason; any other refusal still fails. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nz is the Access application's, not the expression service's - ACE's OLE DB provider answers "Undefined function" - yet queries written in Access use it everywhere, and whoever wrote them expects what Access returns. Measured against Access itself over the same rows: the result is always a Variant, so Nz(K, 0) is written out as the text "0", ORDER BY Nz(K, 0) sorts 10 before 2, and Nz(K, 0) + 1 still adds. LibRed already models that Variant for CVar, and Nz takes the same path. With one argument a Null gives VBA's Empty, written out as "" but read as 0 by arithmetic, so Nz(Null) + 2 is 2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
WITH TIES returns the n rows and every further row whose ORDER BY keys equal the last one's, PERCENT included. It is the rule ACE's own TOP always follows - a plain TOP 2 there returns every row tied at the boundary, which LibRed's exact-n TOP deliberately does not - so WITH TIES chooses the rows ACE chooses (checked against ACE's plain TOP as a set, since ACE orders tied rows unstably). The standard's FETCH … ROWS WITH TIES takes the same meaning. The node that orders the rows makes the cut, since only it holds their keys: the sort, which then stays above the joins, or the aggregate over a grouped query. It needs an ORDER BY, as in SQL Server, and is refused beside DISTINCT, whose rows collapse above the sort. A view refuses it: a stored TOP has no ties flag, and ACE reads one back with ties where LibRed reads it without. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…at way A derived table now takes the standard's column list, (query) AS t(a, b), naming its columns in order - a VALUES table's included, whose columns have no names of their own. SQL Server and PostgreSQL take the same syntax; ACE has none, so a view refuses one (Access stores a derived source as its query text alone), and a derived table with one is only ever read, never written through. The list must name every column, each once. With that, extended mode writes an inline collection the way EF Core's SQL Server provider does, (VALUES (0, CLNG(1)), (1, 2)) AS v(_ord, Value), instead of EF's fallback for databases without a column list, which names the columns on a leading SELECT and puts the other rows after UNION ALL VALUES. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The registry test acquired C:\dir\db.accdb and C:\dir\DB.accdb and expected one manager. That held while the file key folded case everywhere, but the key now folds it only where file names are case-insensitive, Windows and macOS: on Linux those are two files, and two managers is right. So the test failed on every Linux leg. The refcount check now uses one path, and the case rule has a test of its own that expects one manager on Windows and macOS and two on Linux. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The whole-file parity diff showed ACE placing the pages of a large INSERT ... SELECT in 8-page groups where LibRed takes the lowest free page. This probe measures why, on a thousand-row insert and variations of it: row by row, and with the file read while ACE's session is still open, where its reservations are visible. ACE reserves extents in the global free map, four aligned 8-page groups at a time, holds them for the session and frees the rest at close; outside them it keeps data and index pages to separate groups, and grows the file at the end or the next 8-page boundary depending on what the session has used. What triggers an extent is not fully pinned. Every test is explicit: the reservation samples take up to twenty minutes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The step-by-step comparison only reported. 136 of its 139 statements leave a file byte-identical to ACE's, so it now asserts that, with the three that differ named and the reason for each: the bulk INSERT ... SELECT (ACE's extent allocation), a row's long values written in a different order, and an UPDATE that reuses the pages of the value it replaces. A named statement must still differ, and differ only in pages, so the list cannot outlive a fix. The whole-run comparison stays a report: one early placement difference misaligns every page after it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An UPDATE that replaced a chained Memo/OLE value freed the old pages before writing the new value, so the new value landed on them. ACE frees them in the same statement but only after the new value is written: the UPDATE that frees them never reuses them, the statements after it do. Freeing afterwards makes LibRed's file byte-identical to ACE's for such an UPDATE, so the whole-file parity comparison no longer lists it as a known difference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two differences in how a row's Memo/OLE values reach their pages, both measured against ACE with a new probe: - ACE writes all of a row's chained values first and its single-page values after, each in column order. LibRed wrote them in plain column order, so a single-page value in an earlier column took the page before the chains. - ACE places a compressed single-page value by its uncompressed size: it writes the uncompressed bytes where they would go, then the compressed row over their upper end, leaving the rest in the page's free space. A page with room for the compressed row but not the uncompressed bytes is passed over for a new one. LibRed placed and wrote only the compressed row. With both, the whole-file parity comparison's long-value INSERT is byte-identical to ACE's and leaves the known differences. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
linq2db reported 1000 - CASE WHEN s IS NULL THEN 0 ELSE s END over an empty group coming back as the 0 literal's Int32 under a column declared Decimal. Converting a value that holds a choice anywhere to its declared type already fixed it; nothing tested that shape, where the choice is nested in arithmetic and takes the literal arm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A stored query's DISTINCT, TOP and PERCENT are bits of the one MSysQueries option row (Attribute 3), the TOP count in its Name1: ACE writes DISTINCT TOP 2 as Flag 18 and DISTINCT TOP 25 PERCENT as 50. LibRed had three things wrong there: - PERCENT was never written, so a stored TOP 50 PERCENT came back as TOP 50 and returned every row instead of half. - DISTINCT and TOP went on two rows where ACE writes one. - A make-table or append query's own SELECT lost its DISTINCT and TOP altogether, written and read, so a stored SELECT DISTINCT ... INTO copied every row. Every kind of stored query now writes ACE's single row, and a make-table or append query reads its DISTINCT, DISTINCTROW and TOP back into the SELECT. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
linq2db asked for Access SQL's WITH OWNERACCESS OPTION, which runs a query with its owner's permissions. LibRed has no users to act for, so a statement with it does exactly what it does without it. It is taken where ACE takes it: at the end of every SELECT, a UNION's arms and subqueries included, after a query's ORDER BY, and at the end of an INSERT, UPDATE or DELETE. It is refused where ACE refuses it: before the query's ORDER BY, twice, incomplete, or on DDL. Neither word is reserved. A stored query keeps it, as ACE does: CREATE VIEW or CREATE PROCEDURE with it sets bit 0x04 of the query's option row, on the DISTINCT or TOP row when there is one, and otherwise on a row of its own. LibRed reads it back as the clause at the end of the rebuilt statement. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ReservedWords collection lists every keyword the grammar lexes, reserved or not, and WITH OWNERACCESS OPTION added two. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
LibRed.Core.AccessTests has been ending partway through on CI with an xUnit TestPipelineException, reporting only the tests that finished before it, and it never does locally. The job kept nothing to say which test was running. Its four test steps now pass --blame-crash as well as --blame-hang-timeout, and a failed run uploads their TestResults folders: the Sequence_*.xml naming the test in progress, and any dump. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
EF translates a cast from char to an integer type to ASCW, which returns Access's signed 16-bit Integer, so a character from U+8000 up came back negative: cast to uint, U+8000 and U+FFFF read 4294934528 and 4294967295, in both SQL modes. The code is now widened to a Long and masked to 16 bits, (CLNG(ASCW(x)) BAND 65535), which is the character's UTF-16 value. The widening comes first because ACE sign-extends a 16-bit operand of BAND. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Column pruning lets a table read decode only the columns its statement names, which is wrong under anything that looks at a whole row. DISTINCT compares every column, so SELECT Grp FROM (SELECT DISTINCT * FROM T) collapsed twelve rows to three; UNION, INTERSECT and EXCEPT did the same. A set operation also matches its inputs' columns by position, so a SELECT * FROM U under a UNION ALL read none of U's columns when the outer query named them by the other arm's names. The inputs of both now read every column. A derived table with a column list, AS t(x, y), renames its input's columns by position, and the names the outer query uses say nothing about the names inside it. It is left unpruned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A child INSERT or UPDATE inside a transaction checks that its parent exists, and reading the parent writes no page, so the commit's page conflict check could not see it: another connection could delete the parent and commit, and both transactions would succeed, leaving a child referencing nothing. The check is now registered with the transaction and made again when it commits. It only holds the commit to the parent while the parent is still needed: while the relationship survives and a child row still holds the key, so deleting or moving the child, or dropping the relationship, releases it. The relationship is identified by its tables' definition pages and its columns' ids, so renaming a table or column in the same transaction does not lose it. The transaction holds one check per relationship and key however many rows relied on it, and a savepoint rollback frees the keys of the checks it discards. ALTER TABLE ... ADD FOREIGN KEY never looked at the rows already in the table. It is now refused where ACE refuses it, with ACE's message: a child with no parent, or a composite key partly null, whatever the ON DELETE action. ACE then holds both tables exclusively until the transaction ends; LibRed takes no table locks, so the same check is made again at commit. A violation raised by an UPDATE now says so, not "INSERT into". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TableCreator.AddForeignKey had a branch for a self-reference that repeated the rest of the method. What differs is only where the parent's definition page, key index and incoming block number come from: the child itself, with the incoming block taking the next free number above the outgoing one. Those three are chosen first and everything after is shared, with the same calls in the same order, so the bytes written are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A cascade rewrote a child with its own partial copy of the UPDATE checks: no null-key check, so ON DELETE SET NULL could null a primary-key column, and none of the child's own relationships, so a cascade stopped one level down and ON DELETE SET NULL onto a key other rows referenced left them pointing at nothing. The UPDATE loop's per-row write is now UpdateRow, and SetChildKey hands the rewritten child to it: a cascaded row is checked exactly as a directly updated one, and a key it changes applies its own ON UPDATE rules, cascading on down or refused. The parent's cascade runs once the parent row is written, so the child's foreign-key check finds the new key. The commit-time re-check only ran from the child's end. A DELETE, or a parent key change, that found no child writes nothing another connection's child insert conflicts with, so both could commit. It now registers the same condition for the key it removed, held once with the child side's. SELECT INTO writes through InsertNewRow like any insert. A deleted row, a cascade's and a complex column's values go through one DeleteRow; the MATCH FULL key extraction exists once, shared with ADD FOREIGN KEY's check of existing rows; the "related records" refusal is one throw. CHECK expressions, relationships in both directions, their tables and column positions, and an outer-join UPDATE's defaults are resolved once per statement rather than once per row. PageChannel.PageCount inside a transaction is never below the file's own page count, so a transaction can read what another connection committed on pages appended since it began, which the commit-time check has to read. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CLAUDE.md and AGENTS.md: LibRed.Ado.Tests and LibRed.EFCore.Tests use no driver and only run in the Windows ACE job; four functional-suite shards, each run up to three times; pass-lists exist for the two ACE 2010 x86 legs only, and auto_commit extends them after a successful PR run; both connection env vars; fixed order and the en-US culture lock apply to the EF functional suites; two test projects are MSTest; dotnet test builds only net11.0; the Northwind.accdb and LibRed.Shared consumers; page locks are process-local. AGENTS.md also catches up on LibRed.Core.AccessTests, the AccessTests split, the deferred-write overlay and the Core type guard. The LibRed README's "Not yet" lists only open work: stored action queries other than crosstab, pass-through and UNION-kind are written and run, HAVING views work, the 1:1 flag and ValidationRule are read, a column-level CHECK is dropped, RENAME INDEX throws. The referential-integrity and SQL statement descriptions, LOG(base, x), the SQL-mode readers and the Jet3Format stub are corrected. transactions.md gains the parent-side commit check and marks the undo-log design as superseded; the package readme, format index, benchmark and JetLockTrace docs and copilot-instructions are corrected too. LibRed.Engine.Tests no longer suppresses CA1416: it has no ACE code, and the analyzer is what keeps a cross-platform suite that way. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… changes The SID keystream is folded from page 0's header (the 0x42 field and the creation date), so a database is no longer created from one baked date/SID pair: it is dated now and its SIDs are the default workgroup accounts XOR the stream that header gives. The accounts carry their real roles - admin owns MSysDb and new objects, Engine the system tables, Users holds the grants, Creator the inheritable placeholder. Anything that changes the field changes the stream, so the stored SIDs are re-masked with it: the legacy Jet password on an .mdb, and an .accdb's whole-file encryption, whose 0x42 field Access fills with the database key's low byte. The .accdb re-mask runs in a transaction the page copy reads through and then rolls back, so the original file is untouched unless the replacement lands. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This branch closes what two audits of LibRed and a new whole-file comparison against ACE turned up, makes LibRed's query engine substantially faster, compares text in the database's own collation, and adds a handful of Access and standard SQL forms. Nearly every behaviour change was measured against ACE first, and the file-format changes are pinned by byte comparisons against ACE's own output: a 139-statement run now leaves a file byte-identical to ACE's after 138 of its statements. Version bumped to 11.0.0-alpha.4.
EF Core providers (Jet and shared
EFCore.Jet.Common)ALTER COLUMN. That is SQL Server's workaround for refusing to alter an indexed column. Jet'sALTER COLUMNrebuilds the indexes over the column itself, keeping each one's name, columns, order and flags, the primary key included. The shared generator and LibRed's copy both lose the index handling.charconverts to a number by its code point (AscW), not by parsing its text, in both SQL generators.LibRed EF Core provider
UseLibRedaccepts anyDbConnection.(VALUES (0, CLNG(1)), (1, 2)) AS v(_ord, Value). Before, it used EF's fallback: a leadingSELECTnaming the columns, thenUNION ALL VALUES.Packaging
LibRed.Core,LibRed.Sql,LibRed.EngineandLibRed.Adotargetnet10.0as well asnet11.0, since none of them depends on an EF Core package.LibRed.EFCoreand the Jet provider staynet11.0-only. The test projects stay single-target; thenet10.0leg is compiled, not run.nupkgsartifact.LibRed SQL engine
GROUP BYandDISTINCTkeys, set operations, joins,INsets,MIN/MAX, windows, list aggregates and referential-integrity keys.GROUP BYfolds what the collation folds:ßwithss, case and trailing spaces, but not accents.LIKEfolds case by ACE's own table. ACE was asked about 26,440 character pairs. It folds 973 case pairs, a narrower set than any runtime's casing, and spells outß,æ,œandþ. Before, LibRed matched 442 pairs ACE does not and missed 12 it does. The same pairs hold in all 405 creatable orders.CVarand Variants: a Variant keeps its type inside an expression but is written out as text. A choice mixing text with anything else, or a Variant with a non-Variant, is text. A date beside a number is a date.SWITCHandCHOOSEare typed asIIFis.1000 - CASE WHEN s IS NULL THEN 0 ELSE s ENDover an empty group came back as the literal'sInt32under a Decimal column.GROUP BYputs a Long 1 and a Double 1.0 in one group, as ACE does. Byte-array keys, compared by reference, gave every row its own group.S = @p, with a numeric parameter against an indexed text column, crashed with a cast error where the unindexed query returned ACE's 0 rows. Other cross-kind comparisons now scan.Nz, with the results Access itself gives: always a Variant, soNz(K, 0)sorts as text, and one argument over a Null gives VBA's Empty.WITH OWNERACCESS OPTION, requested by linq2db. It is accepted wherever ACE accepts it, changes nothing, and is kept in a stored query.TOP n WITH TIESandFETCH … WITH TIES. ACE's ownTOPalways keeps ties; LibRed's plainTOPstays exact, soWITH TIESis how to ask for ACE's rows.(query) AS t(a, b).IS [NOT] TRUE/FALSEandIS [NOT] DISTINCT FROM.{guid {…}}literal, and GUID text against an indexed GUID column.[Table]![Column]bang notation, with an unaliased one named as ACE names it;[Forms]![f]![c]as one parameter name;Yes/On/No/Offas True and False; any word as a member after.or!; and unbracketed names in any script. Every stored query in the example corpus that uses a bang now parses.DISTINCT,TOPandPERCENTare written on ACE's single option row. Three things were wrong there:TOP 50 PERCENTcame back asTOP 50and returned every row.DISTINCTandTOPwent on two rows where ACE writes one.DISTINCTandTOPaltogether, so a storedSELECT DISTINCT … INTOcopied every row.[@firstName]rebuilt as[[@firstName]]and did not parse.MSysObjectscontainer, as ACE checks it. A query's name used as a table's is refused as "Table 'X' already exists." rather than as an index violation.GetDateTimegavedefault(DateTime)for a serial of 0.LibRed performance
Measured with the benchmark harness against ACE at 10,000 rows.
GROUP BY8,372 → 4,436 µs, hash join 2,988 → 1,681 µs.GROUP BYaggregates in one pass, feeding each row to its group's accumulators rather than buffering the rows. A group's text key is built once.GROUP BYover 10,000 rows went 28.5 → 13.7 ms across the changes.INSERT … VALUEShas one grammar derivation, where the parser used to spend about 60% of an insert settling an ambiguity. A SQL insert in a transaction went 300 → 68 µs.LibRed file format — writes that now match ACE
LibRed.Core, then its completion:CREATE INDEXafterDROP INDEXcould take another object's usage-map row.General(v1) database wrote it back as v0.ALTER COLUMNon a table owning a complex column carries all five of its links.MSysACEsrows are its container's inheritable grants.UPDATEfrees a replaced long value only after writing the new one.CREATE INDEXover existing rows writes the tree that sequential inserts would leave.0x09and released. A table never gives back its first data page.ALTER COLUMNis always an edit in place. To or from Memo/OLE it no longer drops and recreates the table. ATEXT(n)length change takes a full retype, as ACE's does. Several indexes over a column are rebuilt in ACE's order and slots.DROP COLUMNof an append-only memo takes its version-history column, flat table and template with it when it was the last one.0x11is BigBinary: up to 4,000 bytes, always inline, and treated as OLE in queries and indexes.grbit 0x01means one-to-one.GetSchemalists names in collation order.0x10/0x20are written only where Access writes them.MSysObjects.Flags 0x00040000marks a table that owns a complex column.CREATE TABLE's declaration order. A constraint written after a long-value column, the shape a migration emits, took a row ACE gives the column.PageGroupProbeTest) measures the behaviour, but what triggers an extent is not fully pinned. It is the one statement the parity test lists as different.Tests and CI
--record.LIKEfolding against ACE, column pruning, index seek kinds, the parity fixes, stored query options andOWNERACCESS.Docs
src/LibRed/docs/format): updated with everything above that was measured, including released data pages (page-09, renamed for the page type), the version-history structure, the Name AutoCorrect maps and the object kinds.docs/functions.md:Nz, Variants and the new predicates.CLAUDE.md/AGENTS.md: thenet10.0targets, and hooks that deny reading or editing files through the shell and a whole unfiltered ACE access suite in the foreground.🤖 Generated with Claude Code