You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When a second Cortex install holds the pinned loopback ports, the supervisor restarts the child
every 30 seconds forever, logging bind: address already in use each time and never escalating.
The user sees a restart loop that reads as a crash, plus a self-signed-certificate error from their
agent, and nothing in either output says "another install owns this port."
#1104 fixed the install-time case: install.sh now detects a foreign proxy holding the forward
port and refuses to proceed, naming the pid and path. This issue is the runtime case, which that
fix does not reach — the install succeeded, and the conflict began afterwards (the other install was
started later, or the service was restored at login).
Reproduction
Two installs with different $HOME — an ordinary one and a sandbox / second checkout:
$ ps auxww | grep authbridge-proxy | grep -v grep
user 1475 ... ~/sandbox/<name>/.local/bin/authbridge-proxy --local --supervise
user 1031 ... ~/.local/bin/authbridge-proxy --config ~/.cortex/config.yaml
user 1030 ... ~/.local/bin/authbridge-proxy --supervise --config ~/.cortex/config.yaml
user 97657 ... ~/sandbox/<name>/.local/bin/authbridge-proxy --local --supervise
The global install (1030/1031) owns :47600. Both sandbox supervisors loop:
The ports are pinned, so coexistence is impossible. The --local preset binds literal ports
rather than free ones — 47600 forward, 47601 session, 47602 stats, 47603 transparent, 47604 health
(authbridge/cmd/authbridge-proxy/local.go). Second install in, permanently locked out.
The supervisor cannot distinguish "crashed" from "locked out."supervise.go doubles the
backoff to superviseMaxDelay = 30 * time.Second and then repeats at that ceiling indefinitely. A
flapping restart_in=30s is the saturated backoff, not a diagnosis. There is no bind-failure
classification, no escalation, and no give-up.
The supervisor should treat a repeated listen failure as a condition, not a crash:
Classify the bind failure.EADDRINUSE on a pinned port is not a transient crash. Distinguish
it from a child that started and then died.
Name the holder. On bind failure, resolve the pid and executable path holding the port and log
them. install.sh already does exactly this after Fix: Detect a foreign proxy holding the forward port #1104 (port_holder / foreign_proxy_holder) —
the logic and its platform caveats are settled and could be mirrored, including that ps -o comm=
cannot carry a path on Linux and that the holder must be probed on IPv4 loopback, IPv6 loopback and
the wildcard.
Escalate rather than loop silently. After N consecutive bind failures, log once at ERROR with
the holder and the remedy, and either stop or keep retrying with a clearly-stated reason. Looping
~2,880 times a day at 30-second intervals with the same unexplained line is the current behaviour.
Consider surfacing it in abctl service status, which is the command the docs tell people to
run. Today it does not report "another install owns the port."
Open question for whoever picks this up: whether unpinning the ports (bind :0 and publish the chosen
port) is a better fix than better diagnostics. That is a larger change, and it would break the fixed HTTPS_PROXY=http://localhost:47600 contract the docs and abctl configure depend on — so
diagnostics first seems right, but the tradeoff is worth stating.
Docs
The user-facing half is covered by a new troubleshooting section, Two installs on one machine fight
over the ports (rossoctl/rossoctl) — both symptoms, reading ps, the fingerprint comparison, the
macOS PID-wraparound trap, and that SSL_CERT_FILE is inert for Go programs on macOS. This issue is
the product fix that would make those docs unnecessary.
Context
Reported by @aslom on a v0.7.0-rc.3 laptop install; root cause diagnosed by @huang195. Suggested
there that the installer detect duplicate installs and fail early — #1104 does that at install time,
so this issue carries the runtime half.
Summary
When a second Cortex install holds the pinned loopback ports, the supervisor restarts the child
every 30 seconds forever, logging
bind: address already in useeach time and never escalating.The user sees a restart loop that reads as a crash, plus a self-signed-certificate error from their
agent, and nothing in either output says "another install owns this port."
#1104 fixed the install-time case:
install.shnow detects a foreign proxy holding the forwardport and refuses to proceed, naming the pid and path. This issue is the runtime case, which that
fix does not reach — the install succeeded, and the conflict began afterwards (the other install was
started later, or the service was restored at login).
Reproduction
Two installs with different
$HOME— an ordinary one and a sandbox / second checkout:The global install (1030/1031) owns
:47600. Both sandbox supervisors loop:Why it is hard to diagnose
Three things conspire:
The ports are pinned, so coexistence is impossible. The
--localpreset binds literal portsrather than free ones — 47600 forward, 47601 session, 47602 stats, 47603 transparent, 47604 health
(
authbridge/cmd/authbridge-proxy/local.go). Second install in, permanently locked out.The supervisor cannot distinguish "crashed" from "locked out."
supervise.godoubles thebackoff to
superviseMaxDelay = 30 * time.Secondand then repeats at that ceiling indefinitely. Aflapping
restart_in=30sis the saturated backoff, not a diagnosis. There is no bind-failureclassification, no escalation, and no give-up.
The certificate error points at the wrong thing. The request reaches whichever install won the
port, which MITMs with its CA, while the agent was configured to trust the other install's CA.
Every generated CA carries the same subject —
CommonName: "authbridge-tls-bridge-ca", hardcodedat
authbridge/authlib/tlsbridge/ca.go:59— soopenssl x509 -subjectshows an apparentlycorrect certificate and only the fingerprint differs. The user is told to suspect a corporate
proxy or a private CA. Related: bug: a per-sandbox $HOME gives each sandbox its own CA, and the client only hears "Self-signed certificate detected" #1033 (same class, different trigger), abctl configure claude-code status cannot detect a stale or mismatched TLS-bridge CA #1105 (
statuscannotdetect the mismatch).
Proposed behaviour
The supervisor should treat a repeated listen failure as a condition, not a crash:
EADDRINUSEon a pinned port is not a transient crash. Distinguishit from a child that started and then died.
them.
install.shalready does exactly this after Fix: Detect a foreign proxy holding the forward port #1104 (port_holder/foreign_proxy_holder) —the logic and its platform caveats are settled and could be mirrored, including that
ps -o comm=cannot carry a path on Linux and that the holder must be probed on IPv4 loopback, IPv6 loopback and
the wildcard.
the holder and the remedy, and either stop or keep retrying with a clearly-stated reason. Looping
~2,880 times a day at 30-second intervals with the same unexplained line is the current behaviour.
abctl service status, which is the command the docs tell people torun. Today it does not report "another install owns the port."
Open question for whoever picks this up: whether unpinning the ports (bind :0 and publish the chosen
port) is a better fix than better diagnostics. That is a larger change, and it would break the fixed
HTTPS_PROXY=http://localhost:47600contract the docs andabctl configuredepend on — sodiagnostics first seems right, but the tradeoff is worth stating.
Docs
The user-facing half is covered by a new troubleshooting section, Two installs on one machine fight
over the ports (rossoctl/rossoctl) — both symptoms, reading
ps, the fingerprint comparison, themacOS PID-wraparound trap, and that
SSL_CERT_FILEis inert for Go programs on macOS. This issue isthe product fix that would make those docs unnecessary.
Context
Reported by @aslom on a
v0.7.0-rc.3laptop install; root cause diagnosed by @huang195. Suggestedthere that the installer detect duplicate installs and fail early — #1104 does that at install time,
so this issue carries the runtime half.