Skip to content

Sign in for real, and decide who gets in - #67

Merged
davidmckayv merged 11 commits into
mainfrom
feat/i3-sign-in-for-real
Aug 21, 2026
Merged

Sign in for real, and decide who gets in#67
davidmckayv merged 11 commits into
mainfrom
feat/i3-sign-in-for-real

Conversation

@davidmckayv

Copy link
Copy Markdown
Contributor

What this changes

Sign in with what a company already has, and decide who gets in once they are here.

Any one provider turns sign-in on. Google, Microsoft and Okta come from the environment;
configure several and the sign-in screen offers each, on matching buttons carrying each provider's
own mark. Google and Entra are named providers Better Auth knows the endpoints of; Okta is not one
place, so it goes through the generic OAuth plugin against its issuer. They converge at the browser:
one signIn.social({ provider }) for all three, so the app never knows which kind it is asking for.

SAML and OpenID Connect are registered while running. A company's own identity provider cannot be
configured up front, because the deployment is built before it knows whose IdP it will trust. An
administrator pastes the metadata their identity team supplied under Admin → Identity providers, and
somebody signing in types their email address: the part after the @ decides which provider they are
handed to, so a company mid-merger can run two.

A People screen. /admin/people lists everybody who has signed in, the providers they arrived
through, and when they were last here. Promote, demote, or remove. Removing is both halves or it is
theatre: the deny list stops the next sign-in and deleting their sessions stops the current one,
because otherwise a removed person keeps working until their cookie expires, which can be days. It is
keyed on the email address, not the user id, because deleting the row is not removal — the next
sign-in through the provider recreates it with a fresh id.

Nothing configured is one administrator, without a flag, so a fresh clone reaches the product
without registering an OAuth client. The lock moved from a flag to NODE_ENV: somewhere other people
can reach, an unconfigured deployment refuses to start and names what to configure, because a public
URL where every visitor is an administrator is silent and looks like it works. OPENBOT_SINGLE_USER=true
is how somebody says they meant it.

INITIAL_ADMIN_EMAILS is a floor, and now required. An address it names is made an administrator
at every sign-in and cannot be demoted from the People screen, which is the way back in when the last
administrator demotes themselves by accident. Everybody else's role is the screen's to decide, and a
sign-in that rewrote it would make that screen lie the moment they came back.

Important

Breaking. An existing deployment with a provider configured and no INITIAL_ADMIN_EMAILS will
refuse to start. That state was silently adminless before: the role was written once at account
creation and no route anywhere changed one, so nobody could be promoted afterwards.

Five defects this found

Each came from driving it, not from reading it.

  1. Better Auth 1.7 requires an issuer on every account and this schema, written against 1.6,
    had no such column. The adapter rendered where ( = $1 ...) with an empty column name and the
    Google callback failed with an internal error. server/package.json also asked for ^1.6.27
    while 1.7.1 was what resolved, leaving three copies of the adapter installed.
  2. /sso/register only required a session. The upstream plugin guards it with
    sessionMiddleware, so any signed-in person could have registered an identity provider for a
    domain and signed in as anybody at it. Gated to administrators in front of the handler, and tested.
  3. There was no way to grant the administrator role after the first sign-in, and .env.example
    shipped the list commented out. Copying the example and adding a provider produced a deployment
    with no administrator and no route to make one.
  4. Entra's email claim is conditional and Better Auth has no fallback. Microsoft return it only
    when the profile carries an email attribute, and a multi-tenant application may receive no optional
    claims at all. Every authorization decision here is keyed on the address, so somebody would sign in,
    match no administrator, and land as a plain user with nothing explaining why. Now emailupn
    preferred_username, and a refused sign-in with a logged reason if none arrives.
  5. People who had never signed in sorted above people who just had, because Postgres puts nulls
    first on a descending order.

Where it runs

  • New state that outlives a request? Three tables, all in Postgres: sso_providers (owned by
    the plugin), revoked_access, and the issuer column on accounts. Nothing is held in a
    process.
  • What happens on the second replica? Identical. Roles are read per request with
    disableCookieCache, so a role changed on one replica applies to the next request on another.
    Sessions are already in the database.
  • Anything serialised? Setting a role deletes the rows that should not be there and inserts
    the one that should, inside one transaction: between the two, a request on another process
    would find no role and be refused with a 403 that reads as a permissions bug. Revoking is the
    same shape.
  • Anything fanned out to a browser? No.
  • New listener, port, or schedule? None.

Boundary and audit

  • Every people route is administrator-only, and so are the three SSO routes the upstream plugin
    guards with a session alone.
  • Role changes and access changes write person.role_changed, person.access_revoked and
    person.access_restored, carrying the address and the direction. The table holds the current
    answer; the trail is the only thing that can say who changed it.
  • Nothing new is trusted from the client. /api/capabilities projects provider names and a
    boolean, never a client secret, signing certificate, or the list of registered providers, which
    would tell anybody loading the sign-in page which companies use this deployment.
  • Three refusals stop a deployment reaching a state nobody can administer: no self-demotion, no
    removing your own access, and neither for an address the configuration names.

Migrations

0002 generated (nullable column, two tables) · 0003 custom (the backfill) · 0004 generated
(SET NOT NULL).

Only the middle one is written, through drizzle-kit generate --custom, which is Drizzle's own
mechanism for it. A generator diffs schema against schema, and "the rows whose provider is Google get
Google's issuer" is not in the schema. The generatable alternative is a column default, and it is
wrong rather than untidy: every existing Google account would take the placeholder, stop matching
https://accounts.google.com at the next sign-in, and Better Auth would create a second account for
the same person.

Changelog

  • Entries under Unreleased, in Added, Changed and Fixed.

Proof

Driven in Chrome against a database migrated from empty, after dropping and rebuilding it.

Google real sign-in completed, role: admin
Identity userId reaches CopilotKit and Intelligence; zero dev-local-user
SAML registered from metadata; someone@acme.com produces a signed SAMLRequest to the IdP, someone@nowhere.example gets 404
Admin gate same delete call: 403 signed out, 200 as an administrator
People switch and Remove clicked in the UI; row became padlock + "Access removed" + Restore
Roles same account admin → user → admin across sign-ins as the list changed
Audit all four person events on the trail with the address and direction
Sign-in screen three providers with their marks, light and dark; email box appears and disappears as the last provider comes and goes

drizzle-kit migrate run both ways: from empty, and against a database already holding Google,
credential and Microsoft accounts, where the rows come out with Google's real issuer and the
synthetic form for the rest.

799 tests, typecheck, format, lint and drizzle-kit check all clean.

Not proven: Microsoft and Okta have never run against a live provider. The flow is shared with
Google and the one real divergence is handled and tested, but the first person to configure either is
the first person to run it.

One identity provider was a decision somebody else already made. A company
running this has Google or Entra or Okta and is not going to acquire another,
so any one of the three turns sign-in on, several turn on several, and the
sign-in screen draws a button per provider in a fixed order.

Google and Entra are named providers Better Auth knows the endpoints of. Okta
is not one place, so it goes through the generic OAuth plugin against its
issuer, and the plugin is only registered when Okta is configured. They converge
at the browser: one `signIn.social({ provider })` for all three, so the app does
not know which kind it is asking for and a deployment can gain one without a
rebuild.

The provider list moved from the build to `/api/capabilities`. It used to be
compiled into the bundle from the build machine's environment, which was
survivable until the container: one image, built once, knowing nothing about the
deployment that runs it, would have offered a sign-in screen that had never
heard of the provider the operator configured.

Nothing configured now means one administrator without a flag, so a fresh clone
reaches the product without registering an OAuth client first. The lock moved
from a flag to `NODE_ENV`: somewhere other people can reach, an unconfigured
deployment refuses to start and names what to configure, because a public URL
where every visitor is an administrator is silent and looks like it works.
`OPENBOT_SINGLE_USER=true` is how somebody says they meant it.

Two defects found by signing in for real rather than reading the code.

Better Auth 1.7 requires an `issuer` on every account and this schema, written
against 1.6, had no such column. The adapter rendered `where ( = $1 ...)` with
an empty column name and the callback failed with an internal error. Migration
0002 adds it as three statements rather than the one Drizzle generates, because
`ADD COLUMN ... NOT NULL` with no default fails outright on a table that already
has rows, and Google's rows are backfilled with Google's real issuer so they
still match at the next sign-in.

`server/package.json` also asked for `^1.6.27` while 1.7.1 was what resolved,
leaving three copies of the adapter installed. Pinned to what actually runs.
Two ways a deployment could end up with nobody who can administer it, and no
way back from either.

`INITIAL_ADMIN_EMAILS` was optional. Configure sign-in without it and everybody
arrives as a plain user, nobody sees the admin screens, and nobody can promote
anyone, because the role is written from that list and no route anywhere changes
one. `.env.example` ships it commented out, so copying the example and adding a
provider was enough to do it. Sign-in now refuses to start without it.

The role was also written once, in the create hook. Adding yourself to the list
after you had already signed in did nothing at all: the row said `user`, for
ever. It is now reconciled on every sign-in, which also means an address taken
off the list loses `admin` next time it signs in. `user_roles` is a set and the
guard takes `admin` if any row says so, so reconciling deletes the rows that
should not be there rather than only inserting one, both inside a transaction:
between the two a request on another process would find no role at all and be
refused with a 403 that reads as a permissions bug.

Driven on the real path rather than reasoned about: the same account went
admin, then user with the address removed, then admin again with it restored.
The middle step is what the old hook could not do.
…its button

The list and an admin screen have to be able to disagree without one silently
undoing the other. So `INITIAL_ADMIN_EMAILS` is a floor: an address it names is
made an administrator at every sign-in and cannot be demoted, which is the way
back in when the last administrator demotes themselves by accident. Everybody
else is left exactly as they are, because their role is the admin screen's to
decide and a sign-in that rewrote it would make that screen lie the moment they
came back.

That is a change from an hour ago, when sign-in rewrote every role from the list
and would have reverted any promotion made in a screen that does not exist yet.

The buttons now carry each provider's own mark, drawn inline rather than
fetched: this is the one page somebody reaches before they have a session, so a
mark that arrives over the network is one that can be missing exactly when the
page has to look trustworthy, and it asks nothing of a third party from an
unauthenticated page.

Google's guidelines require the standard colour G at its own aspect ratio and
require their button be at least as prominent as any other sign-in option, so
all three are the same size and weight and none of them is the loud one. Okta's
is monochrome, which their guidelines allow: it is not a consumer button anybody
recognises by colour, it is whichever Okta the company uses, and it stays
legible in both themes without a second asset.
An environment variable was the only way to grant the administrator role, and
no route anywhere changed one. That is not how a company runs a deployment: the
people who need access arrive after the deployment does.

So a People screen. Everybody who has signed in, the providers they came
through, when they were last here, and two decisions per row.

Removing somebody is both halves or it is theatre. The deny list stops the next
sign-in and deleting their sessions stops the current one, because otherwise a
removed person keeps working until their cookie happens to expire, which can be
days. It is keyed on the email address rather than the user id: deleting the row
is not removal, since the next sign-in through the provider creates it again
with a fresh id and no memory of having been removed.

Three refusals, all enforced on the server and only mirrored in the browser.
Nobody may demote themselves or remove their own access, because either locks
them out of the screen that would undo it, and on a deployment with one
administrator that is the whole deployment. And somebody named in
INITIAL_ADMIN_EMAILS may be neither, because the floor promotes them again at
their next sign-in and the screen would be lying until then.

Every change writes a row. The table holds the current answer; the trail is the
only thing that can say who changed it and when.

Found by driving it: people who had never signed in sorted above people who just
had, because Postgres puts nulls first on a descending order. On a real
deployment that is the whole first screen given to people who have never used it.
The three configured providers cover a company that uses Google, Entra or Okta.
They do not cover a company that runs its own identity provider, which is most
of the ones that ask, and which cannot be configured up front because the
deployment is built before it knows whose IdP it will trust.

So they are registered while running. An administrator pastes the metadata their
identity team supplied and the provider is stored against an email domain.
Somebody signing in types their address, and the part after the @ decides which
provider they are handed to, so a company mid-merger can run two at once. No
password is asked for and none is checked here.

Registering, changing and removing one is administrator-only. Better Auth guards
those routes with `sessionMiddleware`, which asks only that somebody is signed
in, and that is the wrong bar: registering a provider for a domain means
anybody it vouches for can sign in, so a plain user reaching it could mint
themselves colleagues. The gate sits in front of the handler and is tested.

The sign-in screen grows the email box only when a provider is registered, and
the capability that says so is a boolean rather than a list: naming them would
tell anybody who loads the page which companies use this deployment.

Driven end to end. A registered SAML provider produces a real signed
SAMLRequest redirect for an address at its domain and a 404 for one that is not,
the same delete call answers 403 signed out and 200 as an administrator, and the
sign-in screen adds and drops the email box as the last provider comes and goes.
…s in

The sign-in flow really is the same for all three: authorization code with PKCE,
discovery, an ID token. Google and Entra run through the same function. The
claims inside that token are where they stop agreeing.

Entra does not always send `email`. Microsoft return it only when the profile
carries an email attribute, and a multi-tenant application may receive no
optional claims at all, because an external user's token is minted by their own
tenant and does not inherit this application's claim configuration. `common`,
the default tenant here, is multi-tenant. Better Auth maps `email` straight
through with no fallback, so on those deployments it arrives undefined.

That is worse here than in most products, because every authorization decision
OpenBot makes about a person is keyed on their address: INITIAL_ADMIN_EMAILS,
the role, the deny list and the People screen all read it. Somebody would sign
in successfully, match no administrator, and land as a plain user with nothing
on any screen explaining why.

So `upn` first, then `preferred_username`, and only if it looks like an address:
the OIDC spec explicitly does not promise that claim is one. If none of the
three is there, nothing is returned and Better Auth refuses the sign-in, which
is a better answer than quietly admitting somebody the deployment cannot
recognise. The reason is logged with the claims that did arrive.

Found by reading the provider Microsoft-side rather than by testing, since
there are no Entra credentials here yet.
The issuer migration was one file I had edited by hand after Drizzle generated
it, because the generated `ADD COLUMN ... NOT NULL` fails outright on a table
that already has rows. Editing a generated file is the wrong fix: it leaves a
file that no longer matches what the generator produced.

It is three steps instead, and only the middle one is written:

  0002  generated  the column, nullable, and the two new tables
  0003  custom     the backfill
  0004  generated  the column made required

`drizzle-kit generate --custom` is Drizzle's own mechanism for this, and their
documentation names data seeding as the reason it exists. A generator diffs
schema against schema, so "the rows whose provider is Google get Google's
issuer" cannot come out of one: it is not in the schema.

The generatable alternative is a column default, and it is wrong rather than
merely inelegant. Every existing Google account would take the placeholder, stop
matching `https://accounts.google.com` at that person's next sign-in, and Better
Auth would create them a second account.

Driven both ways with `drizzle-kit migrate` itself rather than by hand: from
empty, and against a database already holding Google, credential and Microsoft
accounts, where the three rows come out with Google's real issuer and the
synthetic form for the rest.

Worth knowing for the check that landed in #64: `drizzle-kit check` reports
"Everything's fine" when a journal entry names a migration file that does not
exist, which is a state a rebase can produce. It cost an hour here. The drift
probe does not catch it either, since both look at schemas rather than at
whether the journal and the directory agree.
The configuration reference still described Google as the only provider and
described `INITIAL_ADMIN_EMAILS` as optional, which is now a start-up failure.
It carries all three providers, what each needs, the callback URL to register,
and why the administrator list is required.

The architecture notes gain the parts a reader cannot infer from the code: that
one resolver answers both questions a run asks about a person, that the
configured list is a floor rather than a one-off, that registering an identity
provider is administrator-only where the upstream plugin asks only for a
session, and that removing somebody denies the address rather than deleting the
row, since deleting it is not removal.

Two lines in the README's feature list, because sign-in and deciding who gets in
are now things the product does rather than things it lacks.

The generated Drizzle snapshots are formatted, which is what the committed ones
already were: `drizzle-kit generate` writes them without a trailing newline and
the format check refuses that.
Audited every markdown file against everything that landed today, including the
work that was not mine.

`docs/coworkers.md` still told people to point `MANAGED_AGENT_AG_UI_URL` at
`4200`. #33 made `agent-langgraph` on `4201` the default precisely because the
proof-of-concept hand-writes the protocol and leaves the tool loop to whatever is
watching, so following that page produced the shape the change moved away from.

Three environment variables the server reads were in `.env.example` and nowhere
in the configuration reference: `AGENT_STALL_TIMEOUT_MS` from #19, which is the
only thing that notices a Bot's stream going silent; `AGENT_TOOL_TOKEN` from #34,
without which no framework Bot may call a granted tool back; and `APP_DIST_DIR`,
which the container sets so one process serves both halves.

Both documentation indexes had fallen behind their own directory and listed
neither `deployment.md` nor `releasing.md`.

`docs/development.md` gains the migration workflow the checks in #64 now enforce:
never hand-edit a generated migration, write a data step with `--custom`, and
what to do when `drizzle-kit migrate` hangs and exits non-zero with nothing
printed, which is the journal naming a file a rebase renamed. `drizzle-kit check`
calls that state fine, because it compares schemas rather than asking whether the
journal and the directory agree.

The README keeps its shape: what this is, how to run it, how to deploy it, and
where to read the rest.
The check boots the container with no identity provider, and the image sets
NODE_ENV=production, where that combination now refuses to start rather than
serve a deployment on which every visitor is an administrator. So the check has
to declare it, which is what the flag is for.

It was passing `OPENBOT_DEV_NO_AUTH=1`, which the code has never accepted:
both the old flag and the new one compare against the exact string "true". It
did nothing, and nothing noticed, because before this branch a deployment with
no provider still started and answered on an unauthenticated route. The refusal
turned a silent no-op into a visible failure, which is the check working.

Reproduced locally with the same command the job runs: answers on
/api/capabilities in four seconds, nothing respawning after fifteen, and the
`eventsource` import error that appeared in the failing log is absent, since it
was the crash loop rather than a fault of its own.
Two configurations that start today refuse to after this, and both were buried
mid-paragraph in Added and Changed. They are four lines at the top of Unreleased
now, saying what to set rather than what used to happen.

Rebased onto #68, which took the deployment's environment away from a Bot's
shell. Checked on the running computer rather than trusting the tests:
`GOOGLE_OAUTH_CLIENT_SECRET`, which this branch introduces, is absent from a
command's environment without anybody having added it to a list. That is the
allowlist earning its shape.
@davidmckayv
davidmckayv force-pushed the feat/i3-sign-in-for-real branch from 653b3be to 98593f4 Compare August 21, 2026 04:32
@davidmckayv
davidmckayv merged commit f725fb5 into main Aug 21, 2026
7 checks passed
@davidmckayv
davidmckayv deleted the feat/i3-sign-in-for-real branch August 21, 2026 04:41
davidmckayv added a commit that referenced this pull request Aug 21, 2026
* Close five ways the sign-in release would have embarrassed us

Found in review of #67, all in the auth work itself rather than in what it replaced.

Running with no sign-in was gated on NODE_ENV === "production", which is exactly backwards: NODE_ENV
is unset unless somebody sets it, so a container on a VM with a hand-written env file and no identity
provider served every visitor as an administrator, silently, because nothing looked wrong from the
outside. It now takes an explicit OPENBOT_SINGLE_USER=true and refuses to start without one.
.env.example ships that line switched on, so a clone still runs with no configuration at all, and the
line is greppable in a way a default never was. The most dangerous boolean in the codebase now has a
test file.

accounts.issuer no longer takes NOT NULL in the same release that adds the column. A rolling deploy
runs the migrations and then serves from old and new replicas at once, and an old replica inserts an
account without the column: under NOT NULL the release would have broken the first sign-in of
everybody who landed on a replica that had not been replaced yet. The constraint belongs to a later
release, once no replica predates the column.

Registered identity providers are facts about the deployment rather than about whichever
administrator pasted the metadata in. Better Auth answers GET /sso/providers with the ones the person
asking registered themselves and refuses a delete from anybody else, so a second administrator saw an
empty screen and registered a provider that already existed, and the row cascaded from the
registrar's user row, so the person who set sign-in up leaving took the company's sign-in with them.
Reads and removals now go through our own admin-gated routes against the whole table, and the foreign
key is set null.

The client secret for a customer's directory was the one secret here not going through
KEY_ENCRYPTION_KEY. The SSO plugin gives no hook, so the seam is the adapter: oidc_config and
saml_config are ciphertext at rest, plaintext rows written before this still read, and OAuth access
and refresh tokens use Better Auth's own encryptOAuthTokens.

Sign-in left no trace at all. Nothing recorded that somebody who could edit INITIAL_ADMIN_EMAILS had
granted themselves the administrator role, and revoking a person deletes the sessions that were the
only evidence they had ever been here. There are now rows for signing in, for being turned away, and
for the configured floor granting the role, and they never block a sign-in when the trail is down.

Also: a failed registration showed its error on the page behind the dialog, so the dialog sat there
looking as though the button had not worked.

* Show what a Bot is doing, not only what it is looking at

The screen answered half the question. A Bot that spends two minutes in a terminal installing a
package shows a blank browser, and the transcript gives it one grey line, `Ran a command  rg
--version`, with the output nowhere: the model saw it, decided which part mattered, and the person
watching had to take its word for it. That is a poor deal on a machine holding somebody's logins.

Two surfaces, saying the same thing. The transcript line stays one line and opens to show what the
command printed, its exit code, and whether it was cut short or stopped, because a transcript of
twenty commands is unreadable if each one dumps a screenful. Beside the screen there is now an
Activity tab that fills up while the screen sits still: every command, file read, file write and
listing, newest first, with a count on the tab so a Bot working elsewhere is visible without
switching to it.

This session only, in the browser. The record is the audit trail, which is on the server, survives a
reload and is what an investigation reads. This is a window, so it needs no endpoint, no polling and
no second copy of command output in the database.

A saved file shows its path and size and never its contents, for the same reason the write route
declines to echo them: a Bot may be saving something it was told in confidence, and a value repeated
into a pane lives in one more place than it should.

* Bring every doc in line with tonight's changes

The changelog carries the history and the README stays about getting started. Four things moved:
running with no sign-in takes a flag rather than a NODE_ENV guess, the issuer constraint is deferred
to a later release and the reason is now written down where the next person will hit it, registering
an OIDC provider needs its discovery endpoints trusted, and there is a second surface beside the
screen showing what a Bot ran.

* Keep the history in the changelog, not the reference docs

Four passages explained a setting by describing what it used to be. That is the changelog's job.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant