Skip to content

feat: organize multimodal capabilities by capability tabs - #253

Merged
yujiezhang-ops merged 1 commit into
mainfrom
claude/multimodal-capability-tabs
Oct 10, 2026
Merged

yujiezhang-ops merged 1 commit into
mainfrom
claude/multimodal-capability-tabs

Conversation

@yujiezhang-ops

Copy link
Copy Markdown
Collaborator

Summary

The Multimodal page listed one block per Provider, so answering "who can generate an image for me" meant reading every block. It now treats each keyed Provider as an account and organizes the page by capability:

  • One tab per capability: image generation, video generation, image understanding, video understanding, text-to-speech, speech-to-text. Each tab shows how many accounts can provide it and a check mark when its Skill is installed. The first tab an account can provide opens by default.
  • Within a tab, accounts are ordered as on the Providers page: user-added first (newest first), then built-ins in catalog order. User-added accounts are marked.
  • Each tab holds the installed Skill card, one row per account (model, evidence, install state, probe, install), which accounts have no model for it, which do not support it, and a manual model ID entry that can target any account.
  • A failed model listing is reported once above the tabs; the ffmpeg hint once per tab instead of per account.
  • Providers without a key stay hidden, as before. BootAgent's protocol converters are left out, as on the Providers page.

Frontend only: account order comes from the existing status.providers via byProviderCreatedAt. No backend or binding change.

Tests

  • pnpm run test: 70 files / 590 tests pass, including new tests for tab order, default tab, arrow-key navigation, account order and marking, converter exclusion, the installed marker, and per-tab notes
  • tsc --noEmit and pnpm run build pass
  • Checked in the browser against the e2e backend with seeded accounts: all six tabs, ordering, empty states and notes; no horizontal overflow at 560px
  • Ran the desktop build against a real ~/.bootagent with built-in and user-added Providers

🤖 Generated with Claude Code

The Multimodal page listed one block per Provider, so finding who can
generate an image meant reading every block. It now has one tab per
capability, image and video generation first, and opens the first one an
account can provide. Each tab lists the keyed Providers (accounts) that
can provide it, the ones the user added first, newest first, then
built-ins in catalog order, as on the Providers page.

Within a tab: the installed Skill, one row per account, which accounts
have no model for it, and a manual model ID entry that can target any
account. A failed model listing is reported once above the tabs, and the
ffmpeg hint once per tab. BootAgent's protocol converters are left out,
as on the Providers page.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@yujiezhang-ops
yujiezhang-ops requested a review from a team October 10, 2026 06:09
@yujiezhang-ops
yujiezhang-ops merged commit e3d1826 into main Oct 10, 2026
4 checks passed
@yujiezhang-ops
yujiezhang-ops deleted the claude/multimodal-capability-tabs branch October 10, 2026 07:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant