Skip to content

The supervisor loops forever when another install holds the pinned ports, and never says so #1122

Description

@mrsabath

Summary

When a second Cortex install holds the pinned loopback ports, the supervisor restarts the child
every 30 seconds forever, logging bind: address already in use each time and never escalating.
The user sees a restart loop that reads as a crash, plus a self-signed-certificate error from their
agent, and nothing in either output says "another install owns this port."

#1104 fixed the install-time case: install.sh now detects a foreign proxy holding the forward
port and refuses to proceed, naming the pid and path. This issue is the runtime case, which that
fix does not reach — the install succeeded, and the conflict began afterwards (the other install was
started later, or the service was restored at login).

Reproduction

Two installs with different $HOME — an ordinary one and a sandbox / second checkout:

$ ps auxww | grep authbridge-proxy | grep -v grep
user   1475  ...  ~/sandbox/<name>/.local/bin/authbridge-proxy --local --supervise
user   1031  ...  ~/.local/bin/authbridge-proxy --config ~/.cortex/config.yaml
user   1030  ...  ~/.local/bin/authbridge-proxy --supervise --config ~/.cortex/config.yaml
user  97657  ...  ~/sandbox/<name>/.local/bin/authbridge-proxy --local --supervise

The global install (1030/1031) owns :47600. Both sandbox supervisors loop:

level=ERROR msg="forward-proxy listen: listen tcp 127.0.0.1:47600: bind: address already in use"
WARN supervisor: proxy exited; restarting ran=37ms err="exit status 1" restart_in=30s

Why it is hard to diagnose

Three things conspire:

  1. The ports are pinned, so coexistence is impossible. The --local preset binds literal ports
    rather than free ones — 47600 forward, 47601 session, 47602 stats, 47603 transparent, 47604 health
    (authbridge/cmd/authbridge-proxy/local.go). Second install in, permanently locked out.

  2. The supervisor cannot distinguish "crashed" from "locked out." supervise.go doubles the
    backoff to superviseMaxDelay = 30 * time.Second and then repeats at that ceiling indefinitely. A
    flapping restart_in=30s is the saturated backoff, not a diagnosis. There is no bind-failure
    classification, no escalation, and no give-up.

  3. The certificate error points at the wrong thing. The request reaches whichever install won the
    port, which MITMs with its CA, while the agent was configured to trust the other install's CA.
    Every generated CA carries the same subject — CommonName: "authbridge-tls-bridge-ca", hardcoded
    at authbridge/authlib/tlsbridge/ca.go:59 — so openssl x509 -subject shows an apparently
    correct certificate and only the fingerprint differs. The user is told to suspect a corporate
    proxy or a private CA. Related: bug: a per-sandbox $HOME gives each sandbox its own CA, and the client only hears "Self-signed certificate detected" #1033 (same class, different trigger), abctl configure claude-code status cannot detect a stale or mismatched TLS-bridge CA #1105 (status cannot
    detect the mismatch).

Proposed behaviour

The supervisor should treat a repeated listen failure as a condition, not a crash:

  • Classify the bind failure. EADDRINUSE on a pinned port is not a transient crash. Distinguish
    it from a child that started and then died.
  • Name the holder. On bind failure, resolve the pid and executable path holding the port and log
    them. install.sh already does exactly this after Fix: Detect a foreign proxy holding the forward port #1104 (port_holder / foreign_proxy_holder) —
    the logic and its platform caveats are settled and could be mirrored, including that ps -o comm=
    cannot carry a path on Linux and that the holder must be probed on IPv4 loopback, IPv6 loopback and
    the wildcard.
  • Escalate rather than loop silently. After N consecutive bind failures, log once at ERROR with
    the holder and the remedy, and either stop or keep retrying with a clearly-stated reason. Looping
    ~2,880 times a day at 30-second intervals with the same unexplained line is the current behaviour.
  • Consider surfacing it in abctl service status, which is the command the docs tell people to
    run. Today it does not report "another install owns the port."

Open question for whoever picks this up: whether unpinning the ports (bind :0 and publish the chosen
port) is a better fix than better diagnostics. That is a larger change, and it would break the fixed
HTTPS_PROXY=http://localhost:47600 contract the docs and abctl configure depend on — so
diagnostics first seems right, but the tradeoff is worth stating.

Docs

The user-facing half is covered by a new troubleshooting section, Two installs on one machine fight
over the ports
(rossoctl/rossoctl) — both symptoms, reading ps, the fingerprint comparison, the
macOS PID-wraparound trap, and that SSL_CERT_FILE is inert for Go programs on macOS. This issue is
the product fix that would make those docs unnecessary.

Context

Reported by @aslom on a v0.7.0-rc.3 laptop install; root cause diagnosed by @huang195. Suggested
there that the installer detect duplicate installs and fail early — #1104 does that at install time,
so this issue carries the runtime half.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions