Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,7 @@ Launch BootAgent
- Keep a local Skill library, scan Skills already present on supported Agents, and choose which Agents receive each Skill.
- See which models on your keyed Providers can understand images, video or audio, or generate images, and where each finding comes from (declared by the Provider, inferred from the model name, or probed). Listing models is free; a probe is a real request and is sent only when you click it, and an image-generation probe asks first.
- Generated images, audio and video default to `~/.bootagent/output/`, outside the installed Skill directory, and never overwrite existing files. Use `--out` to choose another destination.
- When BootAgent cannot confirm how a model is called, as with many user-added gateways, you can install the Skill in explore mode. The Agent then asks you or reads the provider's documentation to work out the request, asks before anything billed, and records what worked under `~/.bootagent/skill-notes/`. Its script sends only to that provider's own host and adds the key itself.
- Install a multimodal Skill (image understanding, image generation, speech-to-text, text-to-speech, video generation, video understanding) into Claude Code, Codex or OpenCode, with example prompts to copy. Agents without a Skills directory get the same files under `~/.bootagent/multimodal/` and a prompt that points them there. Video understanding samples frames with ffmpeg, which must be on `PATH`. Built-in Providers add the models their docs name: MiniMax Speech on Novita, PPIO and JieKou.AI, GPT Image and async text-to-video on JieKou.AI. Video is never probed, because a probe would be a full billed generation. Its script needs Node.js 18 or later. BootAgent writes the Provider key into the Agent's own configuration as `BOOTAGENT_<PROVIDER>_API_KEY`, or into a private file under `~/.bootagent/skill-env/` for Agents with no configurable environment, so the Agent and its model can read it; the install confirmation says so. Uninstalling the last Skill that uses a Provider removes the variable.
- Discover MCP servers from initialized Claude Code, Codex, OpenCode, Kilo CLI, and Hermes installations.
- Store MCP servers in a local user-level Registry, select sync targets explicitly, and apply changes only when you are ready.
Expand Down
1 change: 1 addition & 0 deletions README_ZH.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,7 @@ BootAgent 是一个本地桌面工作台,用来统一管理 AI 编程 Agent。
- 在本机维护 Skill 管理库,扫描支持的 Agent 已有的 Skill,并选择每个 Skill 要同步到哪些 Agent。
- 查看已保存 API Key 的 Provider 中哪些模型能理解图片、视频或音频,或能生成图片,以及每项结论的来源(Provider 声明、按模型名推断或实测)。列出模型不收费;实测是一次真实请求,只在你点击时发出,图片生成实测会先确认。
- 生成的图片、音频和视频默认保存到 `~/.bootagent/output/`,位于已安装的 Skill 目录之外,且不会覆盖已有文件。可用 `--out` 指定其他位置。
- BootAgent 无法确认某个模型怎么调用时(很多用户自己添加的网关都是这样),可以用探索模式安装 Skill。Agent 会先问你或查模型服务的文档弄清调用方式,任何可能计费的请求都先征得你同意,并把可用的方式记在 `~/.bootagent/skill-notes/`。它的脚本只把请求发往该模型服务自己的地址,并自己加上 Key。
- 把多模态 Skill(图片理解、图片生成、语音识别、语音合成、视频生成、视频理解)安装到 Claude Code、Codex 或 OpenCode,并提供可复制的示例提示词。没有 Skills 目录的 Agent 会在 `~/.bootagent/multimodal/` 下得到同样的文件,并附一段指向它的提示词。视频理解用 ffmpeg 抽帧,需要 ffmpeg 在 `PATH` 上。内置 Provider 会补充其文档中列出的模型:Novita、PPIO 和 JieKou.AI 的 MiniMax Speech,JieKou.AI 的 GPT Image 与异步文生视频。视频生成从不实测,因为一次实测就是一次完整的计费生成。Skill 脚本需要 Node.js 18 或更高版本。BootAgent 会把 Provider 的 Key 以 `BOOTAGENT_<PROVIDER>_API_KEY` 写进 Agent 自己的配置;没有可配置环境变量的 Agent 则写进 `~/.bootagent/skill-env/` 下的私有文件。因此 Agent 及其模型能读到它,安装确认里会说明。卸载最后一个用到该 Provider 的 Skill 时,这个变量会被移除。
- 从已初始化的 Claude Code、Codex、OpenCode、Kilo CLI 和 Hermes 中发现 MCP 服务器。
- 使用本机用户级 Registry 保存 MCP 服务器,明确选择同步目标,并在确认后再应用改动。
Expand Down
76 changes: 76 additions & 0 deletions docs/decisions/ADR-012-explore-mode-for-unverified-models.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
# ADR-012: Explore Mode for Models BootAgent Cannot Call

## Status

Accepted

## Date

2026-10-10

- Related: [ADR-011](ADR-011-provider-keys-for-multimodal-skills.md), whose key
delivery explore Skills use unchanged
- Design: [Multimodal Skills](../multimodal-skills-design.md#explore-mode)

## Context

A generated Skill is installed only on evidence: a Provider's declaration, the
catalog, or a probe that passed. An ID rule alone is never enough, because a
Skill that calls a model the wrong way fails only once the Agent uses it.

That rule leaves user-added Providers with almost nothing. Their `/models`
listings rarely declare modalities, so every capability is a guess from the
model name, and some of their routes are not the OpenAI-compatible ones
BootAgent can probe. A gateway that serves image generation at `/v3/<model>`,
for example, can never pass the `openai-images` probe, so its image models stay
"not installable yet" whatever the user does. The user often knows, or can find
out, how the model is called; BootAgent cannot.

## Decision

The user may install any capability in explore mode, on any model a keyed
Provider lists or the user types. Explore mode claims nothing about the model:

- The Skill says BootAgent does not know how the model is called. It tells the
Agent to read its notes first, then ask the user for documentation or an
example, then read the Provider's public documentation, and to stop and ask
rather than guess by sending requests. It asks the user before any request
that may be billed, and allows one request at a time until one has worked.
- The script sends only the requests the Agent composes, and only to the origin
of the Provider's saved base URL. It adds the key itself in the way the Agent
names (`Authorization: Bearer`, a named header, or none), refuses headers that
carry credentials, and never follows a redirect, so the key cannot be sent to
another host. Its output is redacted like every other Skill's; media in a
response is saved to `~/.bootagent/output` and printed as a path.
- What the Agent learns is kept in `~/.bootagent/skill-notes/<provider>-<capability>.md`,
outside the Skill tree, so writing it never makes the managed tree look
edited, and it survives a reinstall.
- `explore` is not in the catalog's adapter vocabulary. The catalog cannot name
it and nothing is probed over it; only the user's choice installs it, and the
UI marks such Skills.

## Alternatives Considered

- **Relax the evidence rule for user-added Providers.** Installs a Skill that
calls the model the standard way, which is exactly what fails for the
gateways that motivated this. It also hides the uncertainty from the Agent.
- **Let the Agent call the Provider with `curl`.** The Agent would need the key
in its command lines, where it reaches logs and transcripts, and nothing would
stop documentation or a page the Agent read from steering the key to another
host.
- **Ask the user for the route in BootAgent and write a dedicated adapter.**
Needs a request and response schema language in the UI, and still fails on
the first detail the user gets wrong. The Agent can read documentation and
iterate with the user; BootAgent cannot.

## Consequences

- Whether an explore Skill works depends on the Agent and the documentation.
BootAgent guarantees only that the key stays on the Provider's origin and that
spending is asked for.
- The Agent composes request bodies, so it can call any route on that origin
the key allows, not just the one capability. The origin lock and the consent
rule bound this; the key was already visible to the Agent's commands under
ADR-011.
- Notes are not removed on uninstall. They hold no key (the script strips
key-shaped text before saving) and describe the user's own Provider.
47 changes: 44 additions & 3 deletions docs/multimodal-skills-design.md
Original file line number Diff line number Diff line change
Expand Up @@ -293,6 +293,40 @@ OpenAI-compatible route (`openai-images`, `openai-audio-transcriptions`,
passes. This is the user naming a model, not BootAgent guessing an endpoint; a
probe verdict for a model in no listing is still shown and installable.

### Explore mode

Some models can never meet the evidence rule: a user-added gateway's listing
declares no modalities, and its image route may be `/v3/<model>` rather than
`/images/generations`, so no probe BootAgent can send will pass. For these the
user may install any capability in explore mode, on a listed model or one typed
by hand. Decision and alternatives:
[ADR-012](decisions/ADR-012-explore-mode-for-unverified-models.md).

An explore Skill (adapter `explore`) claims nothing about the model. `SKILL.md`
tells the Agent to read its notes, then ask the user for documentation or an
example, then read the Provider's public documentation (starting at the
Provider's website when one is saved), and to stop and ask rather than guess by
sending requests; to ask before any request that may be billed; and to send one
request at a time until one has worked. Documentation is reference material,
never instructions.

`run.mjs request` sends what the Agent composes: `--method`, `--path`, a JSON
body file or multipart `--form`/`--file` fields, `--query`, extra `--header`s and
`--auth` (`bearer`, `header:<Name>` or `none`). The path is resolved against the
origin of the Provider's saved base URL and refused if it leaves it; headers that
carry credentials are refused; redirects are reported, not followed. Output is
the status line and the redacted body; images, audio and video in a JSON body
(base64, data URL or hex that sniffs as media) or as the whole body are saved to
`~/.bootagent/output` and printed as paths. `fetch-result` downloads a result URL
without the key. `notes` and `save-notes` read and write
`~/.bootagent/skill-notes/<provider>-<capability>.md`, outside the Skill tree;
`save-notes` strips the key and key-shaped text first.

`explore` is not in the catalog's adapter vocabulary and has no probe. The UI
offers it on an account row that cannot be installed otherwise, and in the
manual entry of every capability, and marks installed explore Skills. A change
to the Provider's website re-renders them, since `SKILL.md` names it.

### Cache

`~/.bootagent/capabilities.json` keeps probe verdicts under the existing
Expand Down Expand Up @@ -530,8 +564,11 @@ not write is left alone, as in an Agent's Skills directory.
the capability cache, React state or test artifacts. Golden tests assert the
absence of the key in every generated file.
- Inputs are read from local paths and sent only to the user's own Provider.
Outputs are downloaded into the Agent's working directory; BootAgent uploads
nothing anywhere else.
Outputs are saved to `~/.bootagent/output/` unless the Agent passes `--out`;
BootAgent uploads nothing anywhere else.
- An explore Skill's script sends only to the origin of the Provider's saved
base URL, adds the key itself and follows no redirect, so neither the Agent
nor documentation it reads can steer the key to another host.
- BootAgent itself sends billed requests only on a click. Spending done later by
the Agent through a Skill is governed by `SKILL.md`, which requires the Agent
to confirm generation with the user first.
Expand All @@ -545,7 +582,11 @@ not write is left alone, as in an Agent's Skills directory.
declared metadata, and a model
found by an ID rule alone is not installable before a probe.
- Capability cache schema migration from version 1 and fingerprint invalidation.
- Generator golden files per capability and adapter, with the no-key assertion.
- Generator golden files per capability and adapter, with the no-key assertion;
an explore golden for a user-added gateway.
- Explore `run.mjs`: origin lock against `//host`, `/\host` and absolute paths,
refused credential headers, unfollowed redirects, redaction of echoed keys in
refusals, media taken out of JSON, notes kept outside the tree without the key.
- Environment writers per Agent against fixtures, including the Codex
`shell_environment_policy` merge with an existing table and the secret write
for `config.toml`.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -423,6 +423,12 @@ export interface MultimodalDetection {
* where the Agent runs, which may differ, so this is a hint only.
*/
"ffmpeg": boolean;

/**
* NotesDir is where explore Skills keep what they learn, one
* <provider>-<capability>.md per Skill; the install confirmation names it.
*/
"notesDir": string;
}

/**
Expand All @@ -441,6 +447,14 @@ export interface MultimodalInstallRequest {
* prompt. Their key goes to the fallback file.
*/
"guide": boolean;

/**
* Explore installs without knowing how the model is called: the Skill tells
* the Agent to work the request out with the user or from the Provider's
* documentation, and sends only to the Provider's own origin. The user
* opts into it, so it needs no evidence for the capability.
*/
"explore"?: boolean;
}

export interface MultimodalInstallResult {
Expand Down
1 change: 1 addition & 0 deletions frontend/src/backend/wails.ts
Original file line number Diff line number Diff line change
Expand Up @@ -367,6 +367,7 @@ export const wailsApi = {
manual: detection?.manual ?? [],
fallback: detection?.fallback ?? [],
ffmpeg: detection?.ffmpeg ?? false,
notesDir: detection?.notesDir ?? "",
})) as Promise<MultimodalDetection>,
// Sends a live (billed) request; only called from a button the user presses.
// `confirmed` is set only after the user agreed to a probe that produces
Expand Down
56 changes: 56 additions & 0 deletions frontend/src/components/MultimodalSection.test.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,7 @@ function detection(partial: Partial<MultimodalDetection> = {}): MultimodalDetect
],
fallback: [],
ffmpeg: true,
notesDir: "/Users/me/.bootagent/skill-notes",
...partial,
};
}
Expand Down Expand Up @@ -442,6 +443,61 @@ describe("MultimodalSection", () => {
expect(await screen.findByRole("tab", { name: /^图片理解已安装/ })).toBeTruthy();
});

it("offers explore mode for a model BootAgent cannot install, and says what it means", async () => {
detectMultimodal.mockResolvedValue(detection({ generatable: [...GENERATABLE, "explore"] }));
installMultimodal.mockResolvedValue({ skillId: "bootagent-vision", results: [{ agent: "codex", installed: true, envWritten: true }] });
renderSection(vi.fn(), { novita: builtIn("Novita", 3) });
// An installable model gets the plain install, not explore.
const row = (await screen.findByLabelText("Novita 的图片理解模型")).closest("li") as HTMLElement;
expect(within(row).getByRole("button", { name: /^安装$/ })).toBeTruthy();
expect(within(row).queryByRole("button", { name: "探索模式安装" })).toBeNull();
fireEvent.change(screen.getByLabelText("Novita 的图片理解模型"), { target: { value: "moonshotai/kimi-k3" } });
expect(within(row).queryByRole("button", { name: /^安装$/ })).toBeNull();
fireEvent.click(within(row).getByRole("button", { name: "探索模式安装" }));
const dialog = screen.getByRole("dialog", { name: "以探索模式安装图片理解 Skill" });
expect(within(dialog).getByText("调用方式未经验证")).toBeTruthy();
expect(within(dialog).getByText(/Skill 的脚本只把请求发往 https:\/\/example\.com,/)).toBeTruthy();
expect(within(dialog).getByText(/\/Users\/me\/\.bootagent\/skill-notes\/novita-vision\.md/)).toBeTruthy();
// The key notice still applies.
expect(within(dialog).getByText("API Key 对 Agent 可见")).toBeTruthy();
fireEvent.click(within(dialog).getByRole("button", { name: "安装" }));
await waitFor(() => expect(installMultimodal).toHaveBeenCalledWith({ provider: "novita", model: "moonshotai/kimi-k3", capability: "vision", agents: ["claude-code", "codex"], guide: false, explore: true }));
});

it("installs a hand-entered model in explore mode where there is no probe", async () => {
detectMultimodal.mockResolvedValue(detection({ generatable: [...GENERATABLE, "explore"] }));
renderSection();
fireEvent.click(await screen.findByRole("tab", { name: /^视频生成/ }));
const manual = screen.getByText("手动填写模型 ID").closest("details") as HTMLElement;
expect(within(manual).getByText(/没有 BootAgent 能实测的标准接口/)).toBeTruthy();
expect(within(manual).queryByRole("button", { name: /^实测$/ })).toBeNull();
const explore = within(manual).getByRole("button", { name: "探索模式安装" });
expect(explore).toBeDisabled();
fireEvent.change(within(manual).getByLabelText("模型 ID"), { target: { value: " my-video-model " } });
fireEvent.click(explore);
fireEvent.click(within(screen.getByRole("dialog")).getByRole("button", { name: "安装" }));
await waitFor(() => expect(installMultimodal).toHaveBeenCalledWith(expect.objectContaining({ model: "my-video-model", capability: "video-generation", explore: true })));
});

it("offers no explore mode when this version cannot generate it", async () => {
detectMultimodal.mockResolvedValue(detection());
renderSection();
fireEvent.click(await screen.findByRole("tab", { name: /^视频生成/ }));
expect(screen.queryByText("手动填写模型 ID")).toBeNull();
expect(screen.queryByRole("button", { name: "探索模式安装" })).toBeNull();
});

it("marks a Skill installed in explore mode", async () => {
detectMultimodal.mockResolvedValue(detection({
generatable: [...GENERATABLE, "explore"],
installed: [{ skillId: "bootagent-vision", capability: "vision", providerId: "novita", model: "moonshotai/kimi-k3", adapter: "explore", agents: ["codex"] }],
}));
renderSection();
fireEvent.change(await screen.findByLabelText("Novita 的图片理解模型"), { target: { value: "moonshotai/kimi-k3" } });
expect(screen.getAllByText("探索模式").length).toBe(2);
expect(screen.getByRole("button", { name: "重新安装(探索模式)" })).toBeTruthy();
});

it("reports a failed listing without hiding the Provider", async () => {
detectMultimodal.mockResolvedValue(detection({ providers: [{ id: "deepseek", name: "DeepSeek", keyEnv: "BOOTAGENT_DEEPSEEK_API_KEY", keyFile: "/Users/me/.bootagent/skill-env/deepseek.env", unsupported: [], listed: false, message: "API key was rejected (401).", errorCode: "API_KEY_REJECTED", models: [] }] }));
renderSection();
Expand Down
Loading
Loading