Skip to content

trunk-merge/pr-1875/5777afee-5039-4c4a-830f-4aa3593ae137 - #1980

Closed
trunk-io[bot] wants to merge 17 commits into
mainfrom
trunk-merge/pr-1875/5777afee-5039-4c4a-830f-4aa3593ae137
Closed

trunk-io[bot] wants to merge 17 commits into
mainfrom
trunk-merge/pr-1875/5777afee-5039-4c4a-830f-4aa3593ae137

Conversation

@trunk-io

@trunk-io trunk-io Bot commented Oct 8, 2026

Copy link
Copy Markdown
Trunk Merge Pull Request Banner

This pull request was created and is being managed by Trunk Merge.

This pull request is based on the main branch at SHA 419f3ada001eed2d39feae6a59a1adb5d8ed1763.

See more details here.

When CI completes, this pull request will be closed automatically.

Pull Requests Being Tested

This pull request is testing the changes from pull request 1875.

Dependencies

This pull request depends on the changes from pull requests 1693, 1711, and 1748.

rigel-mintaka and others added 17 commits October 5, 2026 12:31
…enticated (RIG-4529)

auth.ResolveToken folded every store error into ErrTokenNotFound, so a
database blip during Runner enrollment came back Unauthenticated. The
Runner treats that as permanent and exited instead of retrying.

ResolveToken now maps only store.ErrNotFound to ErrTokenNotFound. Any
other non-revoked error, including ctx cancellation, wraps a new
ErrTokenLookupFailed. The Runner door returns a fixed Unavailable for it
and logs the cause server-side. Missing, revoked, wrong-kind and
malformed tokens stay one byte-identical Unauthenticated. The account
door is unchanged. A security second opinion found no new oracle and no
fail-open path.

Spec-impact: none
Refs: RIG-4529, RIG-3691

Co-authored-by: Matt Wilkinson <matt@rigel.build>
Spec-impact: none
Refs: RIG-4529

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…alone (RIG-3937)

DownDetached took a dark socket as proof that the process-backed server
and postgres had stopped, so a recorded group could outlive down. It now
requires the recorded group to be gone or recycled. A container is
confirmed only by its absence. GroupSignaller.Liveness separates gone,
owned, orphaned and recycled groups: an orphaned group keeps our pgid
and is still signalled. Identity is re-checked before every signal, and
the post-SIGKILL zombie shortcut needs a delivered kill.

Co-authored-by: Matt Wilkinson <matt@rigel.build>
kill(0) EPERM means the pgid belongs to another uid, and our children
share our uid. Falling through to the leader read could call that group
orphaned and signal it, or leave a survivor no later down can kill.

Co-authored-by: Matt Wilkinson <matt@rigel.build>
A pgid from a prior boot may now lead a stranger group, so down drops those
entries unsignalled. Container entries keep the confirm-by-absence path.

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…ocket answers (RIG-4570)

kern.boottime moves on a clock step, so darwin now reads kern.bootsessionuuid.
A live socket proves the stack is from this boot, so down refuses instead.

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…RIG-4571)

Container stops took no ctx and Phase A ran a blocking podman stop per
container, so a teardown could run about 65s before the server budget began.
The app then SIGKILLed down after its record was consumed, which skipped
the survivor rewrite and leaked the stack. ContainerController and
Process.Signal now take ctx, Phase A sends a non-blocking stop signal, the
gateway confirm removes an exited container, and runStackDown cancels with
SIGTERM plus a wait delay.

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…out (RIG-4571)

Phase A signalled every tier at once, so postgres and nats stopped while the
server still drained. Teardown now signals and drains the consumer tier
before the infra tier. rm --force skips the stop-timeout, podman calls get a
wait delay, and the gateway confirm only removes a present container.

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…-4571)

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…RIG-4571)

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…IG-4514)

compass-stack up on the microVM backend never passed --microvm-runroot, so
the runner died at preflight ("run-root is not configured") unless the
operator exported COMPASS_MICROVM_RUNROOT by hand.

runnerSpec now passes --microvm-runroot <RuntimeDir> on the microVM backend
when COMPASS_MICROVM_RUNROOT is unset. The flag would beat the inherited
env, so an operator's value is left alone. The RuntimeDir budget Validate
already enforces for the agent socket also covers the worst-case microVM
gateway socket; a test pins that.

Making up fail while the runner is dead is a separate decision (RIG-4519).

Spec-impact: none
Refs: RIG-4514

Co-authored-by: Matt Wilkinson <matt@rigel.build>
…(RIG-4786)

The advisory fetched the parked message under an AckWait timeout. The
server emits that advisory only after AckWait expires, so the budget
measured nothing and a loaded run timed the fetch out, which left the
park to the callback after the test had moved on. Floor the fetch at
5s. Under concurrent package load the two advisory tests failed 10/200
and 14/200 before and 0/400 after.

Refs: RIG-4786

Co-authored-by: Matt Wilkinson <matt@rigel.build>
Asserting only "above AckWait" let a 2x AckWait bound pass; the test now
requires about the 5s floor and fails that mutant.

Refs: RIG-4786

Co-authored-by: Matt Wilkinson <matt@rigel.build>
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Compass engineering docs preview: https://trunk-merge-pr-1875-5777afee.compass-eng-docs.pages.dev

Deployed from trunk-merge/pr-1875/5777afee-5039-4c4a-830f-4aa3593ae137 at 99b6e65.

Changed pages:

@trunk-io trunk-io Bot closed this Oct 8, 2026
@trunk-io
trunk-io Bot deleted the trunk-merge/pr-1875/5777afee-5039-4c4a-830f-4aa3593ae137 branch October 8, 2026 05:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant