[WRONG BRANCH] release: 2.76.0 - #6468
Conversation
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
…6360) Carries the snapshot half of #4732: keep the per-entry strings measured by selectSnapshotEntries and join them, so a snapshot write serializes entries once instead of twice. Output is byte-identical; a pinned-bytes regression test covers it. Co-authored-by: WU, CHI-LUNG <chilung-cgu@users.noreply.github.com>
… (#6366) Carries #6262 by @LilMGenius: a role model picker for omo (Codex / LazyCodex) agent roles. It adds a Codex-tab section, GET/PUT /api/codex-agent-roles and ocx agent roles. It is gated on LazyCodex detection (the omo@sisyphuslabs plugin plus the lazycodex-install.json receipt). The edit changes only the root model pin and optionally mirrors it to ~/.omo/omo.jsonc. The maintainer follow-up refuses an invalid role TOML and keeps filesystem paths out of write errors. Co-authored-by: LilMGenius <smsmeee@naver.com>
…arry #6269) (#6367) Carries #6269 by @LilMGenius: Auto-assign for omo (Codex / LazyCodex) agent roles. One sizing call asks a model the user already has to sort each role into a capability tier and an effort level, never naming a model. A deterministic mapper then turns each tier into a concrete model and reasoning effort. POST /api/codex-agent-roles/auto-assign only proposes changes. Applying a proposal goes through the existing PUT. The CLI form is ocx agent roles suggest [--apply]. Maintainer follow-ups: sizing failures now return only a fixed category or HTTP status, the effort writer keeps the role-TOML validation from #6366, and two React Doctor warnings are fixed. Co-authored-by: LilMGenius <smsmeee@naver.com>
…6274) (#6389) Carries #6274 by @LilMGenius. It adds Suggest next to the Subagents page's "Model to call first": the user describes the delegated work, one sizing call applies the role-sizing rubric, and the deterministic mapper proposes the cheapest sufficient offered model and an effort. POST /api/injection-model/suggest is read-only, and "Use this" saves through the existing PUT /api/injection-model. On the CLI, ocx agent injection suggest <work> [--apply] skips settings that already match. The maintainer follow-up sanitizes sizing failures on this path as #6367 did for role auto-assign. Co-authored-by: LilMGenius <smsmeee@naver.com>
Co-authored-by: LilMGenius <smsmeee@naver.com>
…complete row (#6354) * fix(claude): log a stalled or over-cap passthrough stream as a 502 incomplete row tapAnthropicSseForLog ended a body stall or byte-cap overflow with an Anthropic error frame but finalized the row as status 200 with only closeReason body_stall / body_overflow. usage.jsonl persists failure diagnostics only for rows at status >= 400 or with a non-completed terminalStatus, so these cut-short turns reached it as plain 200 rows with no closeReason and no reason: anything reading usage.jsonl counted them as successes. The Responses relay logs the same events (a stall-timeout incomplete) as 502 with terminalStatus "incomplete"; httpStatusForRequestLogTerminal maps only a max_output_tokens incomplete to 200. The tap now does the same: 502, terminalStatus "incomplete", the existing closeReason, and the proxy's own message in upstreamError. A mid-stream reset stays 502 / failed, so the two differ only in terminalStatus. The client stream is unchanged. This settles the status split raised in the #6235 review. * fix(claude): treat a stall or overflow after message_stop as a finished turn Review follow-up. failBody did not check whether the turn's own terminal (message_stop or an upstream error event) had already gone through, unlike the read-error branch. An upstream that sent message_stop and then idled past bodyStallSec, or kept sending past the byte cap, got an error frame appended after message_stop and, with the 502 row, would have counted as a failed turn the client had fully received. It now logs 200 / terminal and closes without a frame. Tests: A1/A2 compare the full finalize meta; the managed lane gains a streaming stall case that checks terminalStatus and the tap's reason, and its fold test pins terminalStatus. The structure doc limits the Responses parity claim to the stall and describes the byte cap as matching the non-stream passthrough's 502. * fix(claude): find passthrough SSE frames delimited by CRLF or CR CodeRabbit follow-up. The tap's inspection split frames on "\n\n" only, so a CRLF- or CR-delimited stream never had its frames seen: usage was not read, and with this PR's terminal check a finished turn that then stalled or overflowed would be logged as a 502 incomplete with an error frame after message_stop. The inspection copy is now normalized to LF (forwarded bytes untouched), holding a trailing CR until the next chunk so a CRLF split across chunks stays one line ending. The tail flush the read-error branch already did is now shared with failBody, so a terminal block still in the buffer (an upstream that went quiet right after it, or a held CR) counts there too; the blank line is restored only when the client has not received one.
…mum (#6355) * fix(anthropic): clamp budget thinking to the model's real output maximum The budget-thinking branch sized `max_tokens` with a flat REASONING_MAX_TOKENS_CEILING of 32000, so a caller asking for 128000 on Opus 4.6 or Sonnet 4.6 silently got 32000 — a quarter of what those models document. The adaptive branch directly above already honours an explicit caller limit above 32k; this one overrode it. The constant dates from Claude 4.x models that really stopped near 32k. Opus 4.6 and Sonnet 4.6 document 128K synchronous output and Haiku 4.5 documents 64K, so the flat value is no longer a safe ceiling for the families that take this path: `usesAdaptiveThinking` sends sonnet >= 5, opus >= 4.7 and fable down the adaptive branch, leaving exactly opus-4-6, sonnet-4-6 and haiku-4-5 here, and all three exceed 32000. It now clamps to `configuredMaxOut`, the per-model value already resolved a few lines above for the omitted-limit default (`modelMaxOutputTokens` falling back to `defaultMaxOutputTokens`), so this reuses that lookup rather than adding a second one. REASONING_MAX_TOKENS_CEILING stays the fallback for a target whose ceiling is unknown, which keeps today's behaviour for an unconfigured provider. Measured before and after, with the provider carrying the seeded maxima: opus-4-6 asked 128000 -> 32000 before, 128000 after sonnet-4-6 asked 128000 -> 32000 before, 128000 after haiku-4-5 asked 128000 -> 32000 before, 64000 after (its real maximum) opus-4-6 asked 10000 -> 24576 both (budget+headroom, unchanged) no maxima asked 128000 -> 32000 both (fallback preserved) Anthropic's invariant `max_tokens > thinking.budget_tokens` holds in every case above; the test asserts it alongside the figures. Follow-up to #6329, where the per-model maxima landed and the reviewer noted this clamp as still outstanding. Verification: `bun run typecheck` clean; 73/73 in tests/adapters/anthropic/anthropic-reasoning.test.ts. The new case fails with "Expected: 128000, Received: 32000" without the source change. tests/adapters/anthropic reports the same 654 pass / 30 fail with and without it — those are pre-existing account-pause/pacing failures on dev, unrelated to max_tokens. Co-Authored-By: Claude Code <noreply@anthropic.com> * fix(anthropic): let the stated ceiling only raise, never lower, the thinking clamp Addresses CodeRabbit's finding on #6355: the first revision treated `configuredMaxOut` as a capability ceiling, but that value is two different things depending on the provider. On the canonical Anthropic entries it is a documented Messages API maximum. On any other provider it is a plain fallback budget for requests that omit `max_output_tokens`. Reading the second as the first meant an operator's deliberately cheap `defaultMaxOutputTokens: 8192` was interpreted as "this model cannot emit more", so an explicit 128000 request was clamped to 8192 with the thinking budget squeezed to 4096 — below the 32000/16384 this path produced before the change. Reproduced before fixing: defaultMaxOutputTokens 8192, explicit 128000, high effort before this commit: max_tokens 8192, budget_tokens 4096 base (pre-PR): max_tokens 32000, budget_tokens 16384 The clamp now takes the LARGER of the configured value and REASONING_MAX_TOKENS_CEILING, so the historical ceiling acts as a floor: a stated maximum can lift it where the vendor documents more, and a low fallback budget can never drag it below what the flat constant already allowed. seeded maxima opus-4-6 explicit 128000 -> 128000 (its documented maximum) seeded maxima haiku-4-5 explicit 128000 -> 64000 (its documented maximum) cheap 8192 opus-4-6 explicit 128000 -> 32000 (floor held, budget 16384) no maxima opus-4-6 explicit 128000 -> 32000 (unchanged) The fallback budget still decides an omitted request, which is what it is for, and that is asserted too so the two roles stay distinguished. Verification: `bun run typecheck` clean; 74/74 in anthropic-reasoning.test.ts, the new case covering the low-budget regression with the explicit assertion that budget_tokens stays 16384. tests/adapters/anthropic unchanged at the same pre-existing 30 failures. Co-Authored-By: Claude Code <noreply@anthropic.com> --------- Co-authored-by: Claude Code <noreply@anthropic.com>
…6371) Resolve Anthropic vision and web-search helper credentials through the same per-model account route as primary requests and carry the admitted config snapshot through sidecar planning and execution. An empty strict route is refused locally, so a helper can no longer use an account the model route excludes.
…ers (#6369) Gemini/Devin tool-schema normalization no longer copies surrounding constraints into every type-array branch: constraints stay once on the original node, type alternatives become type-only branches, an existing anyOf keeps its position with the type constraint appended to allOf, and a null-only result yields a single null branch instead of an empty anyOf. Output now grows linearly. Reviewed at 7b22728; hosted Cross-platform CI 36842797911 green.
…failures (#6372) Apply the existing one-sibling Antigravity authentication rotation guard to the terminal-refresh branch, so a request that already rotated after a structured VALIDATION_REQUIRED 403 cannot rotate again when the sibling's refresh fails terminally. Regression: A 403 -> B terminal refresh failure -> C never sent. Reviewed at 83355c9; hosted Cross-platform CI 36838289022 green.
A Kiro connection-reset retry now uses the rebuilt request URL for the physical retry and for send-budget admission, and derives legacy gateway fallback from the rebuilt region. A rebuilt destination is recorded separately from the one permitted alternate-target transition, so the new region keeps its normal gateway fallback while the total send cap still applies. Reviewed at 2534809; hosted Cross-platform CI 36838356858 green.
Keep request cancellation visible for the whole Kiro request, pass the abort signal into lease acquisition, and release any lease granted across the abort race before it can be installed or used. src/server/responses/core.ts reaches exactly its 210-line file-size cap. Reviewed at badac2a; hosted Cross-platform CI 36838656271 green.
Grok Fast mode can serialize grok-4.7-build-fast while the logical route names grok-4.7. The effective wire model is now checked against the caller allowlist across Responses, compact, Chat Completions and Messages; an explicit provider fastWire destination is authorized by its logical id. Docs (en + 7 locales) distinguish the built-in build-fast lane from an explicit fastWire destination. Reviewed at 581b963; hosted Cross-platform CI 36844019764 green.
Owner-registry publication could follow a symlink planted at the final entry, and a permissive registry directory could expose the publication surface. Reject a symlinked or non-directory registry path, require the owner-only directory hardening helper (also on Windows), and publish entries with the no-follow atomic writer.
…6368) Native-model discovery accepted JSON-escaped lone UTF-16 surrogates in model descriptions or nested metadata, which could be persisted and published into a catalog a strict JSON consumer cannot read. After the existing bounded JSON round trip, reject a discovered row containing an unpaired surrogate in any string value, object key or array element. Valid surrogate pairs remain accepted.
…system template (#6385) The openai-chat leading-system template selector matched only Qwen3.8-27B (optionally Qwen/-prefixed). Internal Eliza serves the same pinned template under dashed checkpoint ids (qwen3-8-27b, -fp8, -lora), which failed upstream with "System message must be at the beginning". Match those ids too. Reviewed at fc61562; hosted Cross-platform CI 36851335539 green.
#6323) fix(command-code): preserve canonical path case in project confinement (#6323) Project-context loading lowercased Windows paths before checking containment, so a case-sensitive sibling directory such as PROJECT could pass as the project directory. Containment now uses an exact, case-preserving canonical prefix check with component boundaries, including at drive and UNC roots. The opened file must also resolve back to the byte-identical canonical path before its contents are published.
fix(gui): keep native login confirmation readable and contained (#6388) The native-login confirmation inherited the shared warning notice's horizontal flex layout, so German copy and a long home path collapsed into narrow columns and the confirm button overflowed the panel. The confirmation now renders as a block, and its action row and button text wrap within the panel. A rendered CSS-cascade regression test covers the real view against the shared stylesheet. Addresses the layout part of #6387.
…, SOCKS-only) (#6205) macOS `proxy: "auto"` now explains why discovery was refused: SOCKS-only and disabled transports, the specific toggle (ExcludeSimpleHostnames, PAC, WPAD), and counts of untranslatable exception entries by shape, without logging the entries themselves. With no HTTP(S) proxy, PAC/WPAD are reported before disabled or SOCKS-only.
…ilure test (#6399) test(codex): match the canonical role path in the role-route write-failure test (#6399) getCodexHome() canonicalizes CODEX_HOME, so the role route reads the real path. On macOS the temp root is under /var, a symlink to /private/var, and on Windows runners tmpdir() is the 8.3 short path. The readFileSync spy matched the spelled path, never fired, and the write succeeded, so the test failed on both platforms and passed only on Linux. The spy now matches the canonical role path, the test asserts that the spy denied a read, and it checks that neither the spelled nor the canonical root leaks into the response.
Relay upstream anthropic-ratelimit-* response headers on native Anthropic passthrough (SSE, JSON, error and count_tokens responses) so Claude Code statusLine rate_limits are populated. Only that header family is forwarded; cookies, auth and request-id headers are not. Closes #6297 Co-authored-by: lhj6102 <67728205+lhj6102@users.noreply.github.com>
#6259) (#6392) A Responses turn that produced only a tool call recorded output tokens with no first-output time. Count non-empty response.function_call_arguments.delta and response.custom_tool_call_input.delta in native Responses SSE, and non-empty tool_call_delta arguments in the adapter bridge, as first output. Empty deltas, tool scaffolding, lifecycle and control frames still do not start the timer, and reporting stays once-only. Carried from #6259. Co-authored-by: xyjk <buchanliang@gmail.com>
…(carry #6379) (#6397) A managed service could select a recorded runtime without proving that the live service-manager definition belonged to the same OpenCodex and Codex homes. Bind census delegation to that definition and refuse missing, ambiguous, or mismatched identity evidence before probing or launching a candidate. Decode the quoted home assignments emitted by the systemd writer. For the WinSW backend, also require trusted sc.exe qc to report the definition's own executable as the registered binary. Carried from #6379. Co-authored-by: luvs01 <27862058+luvs01@users.noreply.github.com>
…(carry #5912) (#6362) Carries #5912 by @andrew05060414. It adds an experimental, use-at-your-own-risk Zed Hosted AI provider: - native-app RSA callback login - an account-scoped short-lived LLM token - a bounded live model roster - Anthropic, Google, Responses and Chat requests wrapped in Zed's /completions envelope The provider is opt-in behind the elevated-risk login warning. Maintainer follow-ups: - requests sign with Zed's user_id (OAuthAccessSnapshot.providerUserId) rather than the store slot id - the stream fails closed on malformed, oversized, partial or unterminated frames - catalog calls get an 8 s deadline - the manual-paste login stops once login settles - upstream failure text is scrubbed of the account token and user id Co-authored-by: Andrew <shiyuanwang@ucsb.edu> Co-authored-by: Andrew <andrewwangmail@gmail.com>
Kiro refusal failover replaced the OAuth account snapshot without rebinding the request continuation owner. Bind the continuation and reasoning replay scope to the replacement account after it is selected and before the request is rebuilt or sent; foreign continuation state is removed and response state belongs to the admitted account. The generic OAuth 429 rotation arm had the same gap and now rebinds the same way. Regressions cover buffered and streamed refusals and the source order of both arms. Reviewed at 02f32ae; hosted Cross-platform CI 36865640257 green.
…pool (#6383) When every Anthropic OAuth account is cooled, the pool refuses locally with a 429 whose Retry-After restates the earliest per-account cooldown. The combo target now cools for the local fallback on that pool-local refusal instead of the restated wait (up to the 24h server-delay cap), so a resumed or newly added account is not ignored. Genuine upstream 429s (API-key providers, identified pool accounts) keep their full Retry-After. Follow-up to #6255. Co-authored-by: lidge-jun <jun@lidgeai.com>
…carry #6233) (#6394) Carries #6233 by @chiperman onto dev. For loopback-admitted, recognized Codex Responses clients, a completed hosted image_generation_call is projected into an assistant final_answer message with a Markdown link to a validated local artifact (SSE, JSON, and JSON-to-SSE). When the client replays history, on both ordinary requests and remote compaction, generated local links in assistant content are replaced by opaque artifact references before dispatch, so local paths never reach an upstream. Artifacts reuse the image validation, byte budgets and retention rules. They are written exclusively with random names, display state is bounded, and URL results are not downloaded. Maintainer commits: compaction-path redaction and normalized artifact-link matching (security review finding), plus a structure-doc line-budget move. Closes #6231 Co-authored-by: chiperman <34889778+chiperman@users.noreply.github.com>
…d state directories (#6398) Refs #6314. A ledger file that fails a safety check now raises SpendLedgerFileRefusedError, a SPEND_LEDGER_OWNER_UNAVAILABLE owner error. The error names the file role (journal, journal-compaction, salt) and the failed condition (not-regular-file, symbolic-link, extra-hard-link, foreign-owner, invalid-salt), and never a path, salt, alias or request content. The guard is unchanged. On macOS, startup now warns once when the state directory resolves inside iCloud Drive, a File Provider folder, or Desktop/Documents while iCloud Desktop & Documents sync appears to be on. A sync service briefly holding a second hard link to the journal turns ordinary requests into intermittent 502s. The warning is advisory and never refuses. A new troubleshooting page explains each condition and the fix.
…(carry #6380) (#6396) Policy fallback could authorize a redirect from a bounded diagnostic trace or reselect a public alias, letting a retry escape the original eligible provider/model set. Keep the first policy evaluation's complete membership as request-owned state, evaluate retries as concrete candidates, and check the final destination before dispatch. Subagent selection, recovery, alias collisions and virtual wire models stay inside that membership. The shadow-intercept target is probed without capturing that scope and is captured only when interception is accepted, so a declined policy target never scopes the request's own route. Carried from #6380. Co-authored-by: luvs01 <27862058+luvs01@users.noreply.github.com>
…probe deleted registrations (#6401) Three probe-scoping defects in the guarded Windows manager proof turned a provable absent verdict into a permanent unknown (approval-changed) or a permanent failed stop: - windowsSchedulerCsvIncludesTask treated a backslash as a name boundary, so a same-named task inside a task folder (\Tools\opencodex-proxy) read as present while every /tn operation resolves a bare name in the root folder only — the bound read/stop then failed on a task nothing could reach. - windowsScheduledTaskState ran Get-ScheduledTask -TaskName without -TaskPath, a folder-agnostic filter that returns foreign task state and arrays. - inspectWindowsSchedulerManager read the registration XML before the state check, so an unreadable registration blocked the absent verdict an inert task does not need; and a task deleted between probes was indistinguishable from an unreadable one. The XML now loads only on the running path, and the unanswerable state/XML reads re-probe presence: proven-gone gets the stray-wrapper inert verdict, persisting or unanswerable stays unknown. observeWindowsGuardedManagerStopped re-probes the same way before failing closed, and proven absence still owes the unreadable former-manager check. Verified live on Windows Server 2022: \DevinProbe\opencodex-proxy exists while schtasks /query /tn, /end and /xml all fail root-scoped, the CSV listing emits "\DevinProbe\opencodex-proxy", and Get-ScheduledTask -TaskName returns the foldered task's state plus arrays for duplicate names. Regression tests pin each failure and the preserved fail-closed cases. Co-authored-by: jun <bitkyc08@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: JUN <jun@lidgeai.com>
) (#6404) * fix(update): release the restart lease before the service refresh (#5760) The dashboard update worker held the ownership-mutation lease across the entire restart, including 'ocx service repair'. The service manager's 'ocx start' child is not a descendant of the worker — Task Scheduler, launchd and systemd spawn it with the stored registration environment — so it carries no delegated token, dies at the lease acquire deadline, and the repair's serving probe fails. Every service-installed dashboard update then fell through to an unmanaged detached start beside a dead supervisor — the same class bin/ocx.mjs fixed for failed npm updates. The restart callback now receives an UpdateRestartLeaseControl: the worker releases the lease immediately before the service refresh and re-runs the recorded-owner veto at the fallthrough boundary, so a claim landing during the unleased refresh stops the direct start instead of killing or replacing it. The job records the veto as succeeded/restarted:false. Regression coverage spawns a real supervised 'ocx start' against the held lease (dies with the lease error), then drives restartAfterUpdate through the service path post-fix (child binds and serves) and through the fallthrough boundary under a mid-refresh claim (no kill, no spawn). Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(update): re-acquire the restart lease before the direct-start fallthrough (#5760) The previous seam released the lease for the service refresh but left the fallthrough's veto re-check, killAllOcxOnPort reclaim and pinned start unleased — a desktop claim landing in that window could be killed. reacquireForDirectStart takes the lease back (bounded; a still-claimed lease fails closed), re-arms the delegation token, then re-runs the veto under the held lease. Regression coverage: claim-in-window veto, lease re-held through reclaim/direct start with a live contender probe, and a real lease-holding child proving the bounded acquire fails closed. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(update): resolve the restart-lease authority the way the module does The suite pinned the lock dir to home/.opencodex, which is only the Windows authority: USERPROFILE tracks the sandbox home there. On POSIX homedir() ignores $HOME and the armed test-home guard drops that legacy entry from serviceStatePaths(), so the lease binds to the OPENCODEX_HOME record instead — existsSync(lockDir) was false at every step and the spawned child's contention was never the path the test probed. Resolve the authority with serviceStatePaths().at(-1) and mirror leasePath()'s realpath fallbacks, so both the lock-dir assertions and the foreign-claimant probes address the path the module actually locks. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(update): dump the held-lease child's output when it outlives the watchdog The POSIX CI shard keeps this child alive past the watchdog; without its stdout the stuck point is invisible. Log the captured output before the watchdog throws. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(update): report the child's effective lock path on POSIX survival The Linux child bound and served under the held lease, so its serviceStatePaths() resolved a different authority — likely the legacy home entry that the armed guard did not drop in the spawned env. Log the candidates (os-homedir lock, env REAL_HOME, resolved paths) so the next CI run shows which path it used. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Fix restart lease test child authority fixture * Align restart child home with test guard * fix(update): fail the job when the restart lease stays claimed after a non-serving refresh --------- Co-authored-by: jun <bitkyc08@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: JUN <jun@lidgeai.com>
…ng guard fact (#6402) * fix(cli): stabilize guarded-stop re-verification and report the failing guard fact Three proven defects let an unchanged approved runtime read as changed: - isOcxCommandLine/isOcxStartCommandLine rejected the bundled sidecar's `ocx-<target triple>[.exe]` binary name, so every desktop-spawned runtime failed identity verification at approval and re-resolution. - The managing-CLI `--version` probe invoked a Windows command shim as `cmd /c <shim> --version`; cmd strips the outer quotes when more than one part is quoted, truncating the command at the first space whenever the shim path itself contains one. The shim now runs through a single verbatim-quoted /c line. - Single-shot probes (the shim `--version` probe and the Windows WMIC/PowerShell command-line read) treated one transient timeout as 'unobservable', flipping the compatibility fingerprint or PID identity of an unchanged process. Both get one bounded retry; a persistent failure still answers fail-closed. The approval-changed summary also collapsed ~10 distinct re-checks into one opaque message, so a transient re-probe was indistinguishable from genuine drift. The summary now carries `detail` naming the single guard fact that failed (token fingerprint vs live PID/port/hostname/CLI vs manager verification), threaded through runGuardedManagerStep as well. Addititive on ocx-stop/1; the desktop StopSummary tolerates the field. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cli): scope probe retry to Windows; drop unproven sidecar-triple identity widening Review correction: real Tauri builds run the bundled CLI as ocx.exe and native baseline ownership succeeded against it, so accepting ocx-<target-triple> names in process identity was broadened on an unproven launch path — reverted with its tests. The --version probe retry is now Windows-only (the platform the flake family was verified on); POSIX probes keep exactly one attempt. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(cli): name the revalidation failure when the manager reads unknown An unprovable manager (`kind: "unknown"` with a reason) is not a proven swap, but the pre-action recheck reported both as 'the service manager changed between approval and stop', losing why revalidation could not answer. Unknown now surfaces as 'the service manager could not be re-verified (reason)'; genuine mismatches keep the changed message. Stop and signal still never run against an unknown target. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: jun <bitkyc08@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: JUN <jun@lidgeai.com>
#6419) * fix(claude): rebuild a CLI picker snapshot whose rows no longer decode A snapshot persisted under an earlier provider setup can list routes the current registry no longer decodes. The provider filtered them out and answered with no opencodex rows until the 5-minute refresh, and Claude Code cached that empty list for an hour. Rebuild the snapshot within the cold-wait bound instead. * docs(structure): note the CLI picker snapshot rebuild * fix(claude): share one cold-wait deadline and keep the snapshot when the registry build fails
… a restart (#6428) * devlog: Claude UX roadmap (intercept on demand, Claude page, app rebuild) * fix(claude): start the intercept pair on demand instead of asking for a restart Saving CLI or Desktop first-party settings refused with "restart needed" when the intercept pair had not bound at startup (Claude routing off, port busy, etc). The lifecycle now has a serialized ensure() that starts the pair in the running process through the same dispatch wrapper, records why a start failed (disabled, client role, ephemeral port, port in use, port mismatch) and is called by the first-party, Desktop picker and Claude routing toggles plus a new local-only POST /api/claude-intercept/start. The GUI shows the reason and a Start interception action; no copy asks for a restart. * fix(claude): gate intercept start on enable paths, reject a changed port, add ocx claude intercept start * fix(claude): report intercept start failures from the native toggle and the picker bind port
…ings tabs (#6430) * feat(gui): top-level Claude page with Account, Code, Desktop and Settings tabs * feat(gui): start interception from the Claude Settings tab * fix(gui): address react-doctor findings on the Claude page * fix(gui): scope the Claude Account tab's reads and sign-in to Anthropic
…#6440) * fix(gui): conditions-based Claude account-pool and OAuth warning copy The Claude account-pool card showed an always-on role="alert" box ("Experimental and not battle-tested ... Keep this off unless you understand the risk") that told operators neither what the pool changes nor that 429 failover keeps running with it off. The card now states the conditions the pool is meant for as static helper text: your own or authorized subscriptions, the genuine Claude Code client, a person supervising the session, Anthropic's current terms, and shared organization quota. A closed "How account selection works" disclosure explains what enabling changes, that a 429 or classified account refusal still fails over when the pool is off (pause an account to exclude it), and the default background activity (no keep-warm requests, no background token refresh or usage polling; usage is read when the dashboard, menu bar app or an ocx command asks). Threshold help calls the threshold a preference, not a usage or billing cap. Save failures remain the only live alert. The OAuth terms dialog gets Anthropic-specific title, body, conditions, API-key alternative, acknowledgement and continue copy, selected by oauthTosCopyKeys. Other providers keep the shared keys; the unchecked acknowledgement, Cancel/Escape/backdrop and single submit are unchanged. All ten locales carry the same claims; the Claude guide and config references drop the old "not battle-tested" wording. The copy describes only behavior that ships today and claims no account-pool mechanism that is still in progress. Framing informed by KarpelesLab/teamclaude docs/compliance.md (MIT); no code ported. * fix(gui): match the Claude pool status line to the selected strategy The enabled status line said new sessions and refusal recovery follow the quota window even under round-robin, where pickUnboundStrategyAccount and pickAlternateAnthropicAccount rotate through the ring without reading usage, the threshold or the window (and the window selector is disabled). The line is now chosen by strategy: - round-robin: sessions and refusal recovery take turns; usage, threshold and window unused - fill-first: the active account is drained to the threshold in the window (or, at 0, kept until cooldown or sign-in), then the next account in order; recovery also takes the next - quota: unchanged above 0; at 0 a healthy active account stays and the lowest-usage account in the window is chosen only when it is unavailable or during refusal recovery The details line now says the selected strategy chooses accounts instead of naming the threshold and window. All ten locales carry the same claims and placeholders. structure/dashboard-and-usage.md said load and save failures are both alerts; only a failed save has role="alert", while a failed load replaces the status line and disables the toggle.
…e serving OAuth account (#6442) * fix(anthropic): bind native metadata to serving OAuth credential * fix(anthropic): preserve observed Claude Code native request identity * fix(anthropic): preserve operator headers and guard all OAuth builds * docs(structure): scope Claude identity continuity to unpooled native switches
Preserve journal reconciliation and local startup refusal while avoiding a false Codex shim autostart error. Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com> Co-authored-by: Andrew Wang <59988150+andrew05060414@users.noreply.github.com>
…the Claude account pool (#6443) * fix(anthropic): classify 429 before account health mutation * feat(anthropic): admit requests with family weekly quota evidence * fix(anthropic): tighten family admission and 429 classification edges
…ks (carry #6439) (#6446) * fix(streaming): decode CR-delimited server-sent events * fix: address review feedback for PR #6439 * fix(streaming): keep CR/LF delimiter search native and cover split CRLF framing Carry follow-up for #6439. Replace the per-character delimiter loop with cached indexOf cursors for the next CR and LF, so the shared decoder keeps native, linear scanning on every provider stream. Add explicit cases for a CRLF split between chunks (including an empty chunk in between), LFCR, many records in one chunk for each framing, a UTF-8 code point split next to CR, and raw versus escaped CR in data. Give the Korean guide section a Korean heading and record the cursor search in the byte-accounting owner doc. Co-authored-by: 정우철 <oocheol@naver.com> --------- Co-authored-by: 정우철 <oocheol@naver.com>
…6431 #6432 #6433 #6434 #6435 #6437 #6438) (#6447) * fix(cli): reject ignored capabilities arguments (carry #6429) Carries #6429 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * fix(cli): reject ambiguous and unsafe integer options (carry #6431) Carries #6431 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * fix(cli): accept UTF-8 BOM in imported config files (carry #6432) Carries #6432 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * fix(cli): validate setup ports before saving config (carry #6433) Carries #6433 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * fix(cli): support JSON output on the default alias command (carry #6434) Carries #6434 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * feat(cli): show request IDs in human log output (carry #6435) Carries #6435 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * fix(cli): return not-found for missing routing profiles (carry #6437) Carries #6437 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * fix(responses): keep usable fallback upstream error messages (carry #6438) Carries #6438 by @oocheol onto current dev as part of the oocheol CLI batch. Co-authored-by: 정우철 <oocheol@naver.com> * fix(cli): ask for the setup port again instead of abandoning ocx init Amends the #6433 carry: an invalid interactive port now reports the error and re-asks, so a typo no longer discards every earlier answer. EOF and SIGINT at the repeated prompt still cancel without saving (exit 1 / 130); the wizard tests cover all three paths. Co-authored-by: 정우철 <oocheol@naver.com> * test(responses): pin recovery consumers of the upstream message fallback Follow-up to the #6438 carry: encrypted-output resend and quota alternate-account detection now see a fallback message hidden behind a blank primary one; a nonblank primary still wins. Co-authored-by: 정우철 <oocheol@naver.com> * docs(cli): separate the carried reference sections Co-authored-by: 정우철 <oocheol@naver.com> * chore: keep the lane plan record out of the carry --------- Co-authored-by: 정우철 <oocheol@naver.com>
…tGPT credits opt-in (#6444) An account at 100% of a usage window is now switched out by default and returns after its reset; spending ChatGPT credits is opt-in through creditCodexAccountIds. The Codex Auth header has one global "Use credits" switch beside the credits display, and each account's switch lives in its card's more menu. Refines #6422 and closes #6334. Co-authored-by: Sungyong Cho <dev@sungyongcho.com>
#6441) * fix(service): wait out the npm Bun placeholder instead of executing it - The Windows service wrapper now applies the shared REAL_BUN_MIN_BYTES rule before launching Bun; during an in-place npm install it logs and retries every 5s until bun's postinstall replaces the placeholder, instead of exiting 216 behind a modal 16-bit dialog * fix(service): keep waiting when Bun vanishes between the exist check and the size read - An empty OCX_BUN_BYTES made the size comparison a cmd syntax error that ended the wrapper; the read now resets the value and treats an unset size as not ready - The wait message names a persistent placeholder's remedy * docs(service): describe the Windows service wrapper builder - buildWindowsServiceScript names what the generated batch wrapper does, answering the CodeRabbit docstring check * docs(structure): record the Windows wrapper's wait on the npm Bun placeholder - docs-and-release and runtime state that the wrapper re-checks a placeholder or vanished Bun every five seconds, that a Bun path missing at the exist check keeps the bun_missing exit, and how a permanently blocked postinstall recovers
…ry and Sionic-first ordering (#6448) * feat(providers): add OpenGateway (Sionic AI) preset with live discovery and Sionic-first ordering * fix(opengateway): read preferFirst from captured gather policy, register test layout, add provider mark - routed-gather read the live registry after capture (codex-gather-authority); the ordering decision now comes from the captured discovery spec, and preferFirst is part of the captured discovery-policy snapshot. - register tests/providers/opengateway-provider.test.ts in the test layout map and its expected fixture. - add OpenGateway's header mark (opengateway.ai/logo.svg, byte for byte), masked, wired in the GUI and desktop shell; the favicon was rejected for a <text> glyph. * fix(opengateway): mark public catalog non-validating, admit only rows served on their wire - apiKeyValidation: "unknown": GET /v1/models answers 200 for any Bearer, so it cannot prove a key (Codex review P2). - discovery admits active chat_completions rows plus the Responses-only rows pinned to Responses (openai/o3-pro); a future Responses-only row stays hidden instead of being published on the Chat default and 404ing (CodeRabbit major). Live: still 70 admitted, Sionic rows first. - docs/structure wording follows the new filter in all locales. * fix(opengateway): keep a custom replacement in its preferred slot; pin derived non-validating key login A custom model that replaces a discovered row of a provider whose discovery declares an order (preferFirst) now takes that row's slot instead of moving to the end of the list. Providers without preferFirst keep the previous order: discovered rows first, custom rows after. Also pins that the derived OpenGateway key-login entry keeps apiKeyValidation "unknown" and that validation never probes the public catalog.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Important Review skippedToo many files! This PR contains 578 files, which is 278 over the limit of 300. To get a review, reduce the PR to 300 files or fewer by splitting it into smaller PRs or changing its base branch. Usage-priced reviews support at most 300 files. ⚙️ Run configurationConfiguration used: Repository: lidge-jun/opencodex/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: ⛔ Files ignored due to path filters (2)
📒 Files selected for processing (578)
You can disable this status message by setting the
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
✅ Deterministic PR hygiene checks passed. |
⏳ DRAFT
What to do
Its title has been prefixed with |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d2f8fd3562
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // A validated client delegates inference to its hub; no local startup is needed. | ||
| return clientState.kind === "connected" ? 0 : 1; |
There was a problem hiding this comment.
Stop reporting successful connected ensure as an error
When ocx ensure runs with a validated connected-client state, this new branch returns success but still executes the preceding console.error("Client mode does not start..."). Direct users and wrappers that treat any stderr output as a failure therefore receive an error-looking diagnostic for a successful no-op. Suppress the message on the connected path or emit it as informational stdout, while retaining stderr for invalid or mismatched states.
Useful? React with 👍 / 👎.
|
|
||
| function zedCacheKey(credentials: ZedCredentials): string { | ||
| // In-memory caches are keyed by an irreversible digest so no map key holds the raw token. | ||
| return createHash("sha256").update(JSON.stringify([credentials.userId, credentials.accessToken])).digest("hex"); |
| // catalog cache bound to the same pair so a multi-account switch cannot reuse a stale | ||
| // roster even when the provider destination is unchanged. | ||
| const authorityIdentity = createHash("sha256") | ||
| .update(JSON.stringify([auth.oauthAccountId, apiKey])).digest("hex"); |
|
Maintainer promotion into
|
Adds 035_release_2_76_0.md (promotions lidge-jun#6467/lidge-jun#6468, Release runs 37041498025/37047146196, both tags published 25 assets) and 080_wave_1002_tail.md (lidge-jun#6448 OpenGateway preset, lidge-jun#6462 dev-open). Window extended to 10-02 16:55 UTC; totals updated to 81 first-parent commits, 68 uncovered merges, 4 release rounds. Co-Authored-By: Epinephrine <luvs01@hanmail.net>
Summary
Promote the verified
devcandidatee0af52c8a2701dccd81fa5e672c92744e59d2e29tomainas stable2.76.0. The tree is identical to the candidate (which already carries 2.76.0 in all four version sources); the branch descends fromorigin/mainthrough anoursmerge.2.76.0 contents since v2.75.0 (60 commits): Claude account pool hardening (#6442 serving-account identity, #6443 429 classification and per-model weekly admission), Claude Code CLI first-party
/modelpicker (#6418, #6419), on-demand intercept start (#6428), top-level Claude dashboard page (#6430) and conditions-based copy (#6440), ChatGPT credits opt-in with 100% auto-switch default (#6444), window-aware main-account hard lock (#6363), app-server quota-gate shim (#6361, #6412), JEV decision methods (#6364), LazyCodex role models (#6366, #6367, #6389), Zed Hosted AI and Antigravity TLS carries (#6362, #6359), Windows service and update hardening (#6391, #6395, #6400–#6404, #6407, #6441), Windows Bun placeholder wait (#6441), SSE CR/CRLF decoding (#6446), oocheol's CLI fixes (#6447),ocx ensureconnected-client exit (#6427), OpenGateway preset (#6448), plus the Kiro, Gemini, xAI, Anthropic and spend-ledger fixes listed in the dev log.Release authorization: the repository owner asked on 2026-10-03 to release 2.76.0 (preview, then stable) with regression verification.
GUI changes carried from dev (screenshots from the source PRs):
Verification
e0af52c8a2: workflow_dispatch Cross-platform CI run 37012309480 success (windows 9/9 rerun once; the sibling-home timeout passed in run 37000419882 with unchanged owning code and passed on rerun) and Service lifecycle run 37012313136 success.devat 2.77.0 (merged as b4616be).git diff e0af52c8a2 HEADis empty: the merge commit's tree equals the candidate's.release.yml.Checklist