chore(release): v1.3.0 — Terminal-Bench 2.1 result + Opus 5 - #765
Merged
Conversation
Bump version to 1.3.0 across pyproject.toml, src/__init__.py, install.sh, ui-tui/src/gatewayClient.ts, and uv.lock. Add the [1.3.0] CHANGELOG section and a news item announcing 80.9% pass@1 (72/89) on Terminal-Bench 2.1 with claude-opus-5 — a single-run (k=1) result that would slot around third on the public leaderboard, ahead of Claude Code on Opus 4.8. Sync README / NEWS archive / ZH README to the top-10 news window. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release v1.3.0
Version bump + Terminal-Bench 2.1 achievement announcement.
Version (all spots → 1.3.0)
pyproject.toml,src/__init__.py,install.sh(INSTALLER_VERSION, was stale at 1.1.0),ui-tui/src/gatewayClient.ts(CLAWCODEX_VERSION, was stale at 1.1.0),uv.lock.Terminal-Bench 2.1 result
Running headless on
claude-opus-5ateffort=xhigh, ClawCodex solved 72 of 89 tasks — 80.9% pass@1 on a single run (k=1). On the public 2.1 leaderboard (k=5 averages) that would slot around third — behind Claude Code / Fable 5 (83.8%) and Codex / GPT-5.5 (83.1%), and ahead of Claude Code on Opus 4.8 (78.9%) and Sonnet 5 (74.6%). Benchmarked onmainat #756 (before this tag). Run artifacts:eval/harbor/jobs/tb21-clawcodex-3.Docs
CHANGELOG.md: new[1.3.0]section (docs(readme): add PyPI installation to translations #719–feat(tui): rework the header box to the reference element allocation #764).README.md/docs/NEWS.md/docs/i18n/README_ZH.md: news item added; README + ZH held at exactly 10 items (DeepSeek prefix-cache item moved to the NEWS archive).Every benchmark claim was independently re-derived from the
result.jsonfiles and cross-checked against the live leaderboard; the write-up is deliberately conservative about the k=1-vs-k=5 comparison, timeouts, and the model/harness confound.🤖 Generated with Claude Code