From 6ab3b3fb20c4cf12ebabd23f7357d6628b8607ea Mon Sep 17 00:00:00 2001 From: David McKay Date: Fri, 2 Oct 2026 12:10:50 -0700 Subject: [PATCH 1/2] Bring the docs and the changelog up to date with the code before release Every claim in the README, .env.example, docs/ and the chart README was checked against the code on main and corrected where it had drifted. The README points to running Intelligence locally on the free Developer plan and to having CopilotKit build OpenBot with you. The Unreleased changelog gains entries for Automatic Learning, Parallel Search, constant-time token checks and the lodash-es pin, an upgrade note naming migrations 0042 to 0051, and the entries that had been written into the already released 0.0.15 and 0.0.14 sections after their tags. --- .env.example | 32 +++++--- CHANGELOG.md | 167 ++++++++++++++++++++++----------------- README.md | 68 +++++++++++++--- charts/openbot/README.md | 39 +++++++-- docs/README.md | 1 + docs/architecture.md | 47 ++++++----- docs/configuration.md | 57 +++++++++---- docs/coworkers.md | 27 ++++--- docs/development.md | 2 +- docs/plugins/composio.md | 9 ++- docs/releasing.md | 9 ++- docs/routines.md | 9 ++- docs/windows-signing.md | 7 +- 13 files changed, 307 insertions(+), 167 deletions(-) diff --git a/.env.example b/.env.example index 190dce0c2..b03d6cab2 100644 --- a/.env.example +++ b/.env.example @@ -108,21 +108,25 @@ OPENBOT_SINGLE_USER=true # OPENBOT_APP_URL=http://localhost:3010 TRUSTED_ORIGINS=http://localhost:3010 -# CopilotKit Intelligence. Required: the server refuses to start without all four because -# Intelligence owns durable threads and memory and a deployment without it forgets every -# conversation. There is no degraded mode. -# -# Do not run Intelligence yourself. These two point at the managed service and should be left -# unchanged; the CLI below provisions a free licence for it. Running Intelligence on your own -# infrastructure is an Enterprise Intelligence Platform feature deployed by Helm chart, and is not -# self-serve: https://docs.showcase.copilotkit.ai/premium/self-hosting +# CopilotKit Intelligence. Required: the server refuses to start without the URL, the gateway URL +# and the key below, because Intelligence owns durable threads and memory and a deployment without it +# forgets every conversation. There is no degraded mode. +# +# These two point at the managed service; leave them unless you run Intelligence yourself. To try +# it on one Mac, the CLI's local evaluation writes both for you (see "Intelligence on your own +# machine" in README.md): https://docs.copilotkit.ai/intelligence/self-hosting-local +# Running it in production on your own infrastructure is an Enterprise Intelligence Platform +# feature deployed by Helm chart: https://docs.showcase.copilotkit.ai/premium/self-hosting INTELLIGENCE_API_URL=https://api.intelligence.copilotkit.ai INTELLIGENCE_GATEWAY_WS_URL=wss://realtime.intelligence.copilotkit.ai # -# Get both of the next two from the CopilotKit CLI: +# Get the key from the CopilotKit CLI: # # npx --yes copilotkit@latest login # browser sign-in -# npx --yes copilotkit@latest project select # prints the cpk-... runtime key -> INTELLIGENCE_API_KEY +# npx --yes copilotkit@latest project select # writes the cpk-... key to .env as CPK_INTELLIGENCE_API_KEY +# bun scripts/setup-learning.ts # copies it here, as INTELLIGENCE_API_KEY +# +# The server reads only INTELLIGENCE_API_KEY, so a key left as CPK_INTELLIGENCE_API_KEY is not seen. # # The runtime key is also under "API Keys" in your project at https://intelligence.copilotkit.ai INTELLIGENCE_API_KEY= @@ -284,8 +288,9 @@ COMPUTER_TOKEN= # volumes, so one Bot cannot read another's files or use another's logins, and every action records # which Bot took it. A rule can still restrict a single Bot with `bot.id`. # -# Attributes: tool.name, bot.id, actor.id, page.url, page.host, element.ref/role/name/type, -# key, file.path, file.name, file.extension. +# Attributes: tool.name, intent, bot.id, actor.id, page.url, page.host, +# element.ref/role/name/type, key, command, file.path, file.name, file.extension, +# mcp.server, mcp.tool, mcp.effect, initiator.kind, initiator.id. # # Name every route to the same effect. A form submits from a keypress in any of its fields, so a rule # that only blocks a Submit button does not block Enter from another field. The example below refuses @@ -405,7 +410,8 @@ AGENT_TOOL_TOKEN= # # Left empty, the internal endpoint refuses every call and no routine ever fires — the correct state # for a deployment that has not stood up a worker. Set for one that has: openssl rand -base64 32. -# Do not accept a default in production. +# `scripts/start.sh` falls back to a fixed development value locally. Do not accept a default in +# production. WORKER_SHARED_SECRET= # Composio, the broker that holds people's accounts for a few hundred apps, so a Bot can act in Gmail diff --git a/CHANGELOG.md b/CHANGELOG.md index 18ae593db..888608113 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,18 @@ Newest first. `Unreleased` is what is on `main` and not yet tagged. ## Unreleased +**Before upgrading.** Four things change for an existing deployment: +- Automatic Learning is on unless an administrator saved it off. It does nothing until a Learning + container is assigned; see below. +- A Bot's computer refuses the network until the server pushes its policy. A computer run without + an API server can set `EGRESS_POLICY_REQUIRED=0` for the old behaviour. +- The upgrade runs migrations `0042_user_preferences`, `0043_plugin_logos`, `0044_voice_sessions`, + `0045_channel_activity_source`, `0046_automatic_learning`, `0047_agent_pinning`, + `0048_coworker_parity`, `0049_coworker_parity_lanes`, `0050_review_fixes` and + `0051_routine_enabled_at`. +- An existing Windows clone checks text files out with LF only after + `git rm -r --cached . && git reset --hard` on a clean tree. + ### The Helm chart configures Slack, Teams, text messages, push, SCIM, inbound email and OpenTelemetry These settings had no chart values and could only be passed through `config.extraEnv`. They now have @@ -43,7 +55,6 @@ not draw. Playground components now publish and withdraw both rows in one transa empty description or HTML is refused instead of leaving a half-published component, and repeating an unchanged publish no longer advances the revision. - ### A proxy password containing `%` no longer stops every shell command A proxy password with a `%` that does not start an escape, such as `p%zz`, made decoding it throw. @@ -99,6 +110,13 @@ the two scripts came back empty. Where those tools are missing on Windows, the scripts now ask PowerShell, which ships with Windows. Everywhere the tools exist they are used exactly as before. +### The supervisor and the Python Bots compare their tokens in constant time + +The supervisor compared its bearer token with a plain string comparison, and the eleven Python Bots +compared the shared agent token as text, which answered 500 rather than 401 on a header carrying a +non-ASCII character. Both now compare bytes in constant time, as the server and the computer already +did. A wrong or missing token is refused exactly as before. + ### A clone on Windows builds an image that starts On Windows, where Git converts line endings by default, a clone checked every text file out with @@ -115,7 +133,9 @@ A routine that fails ten times in a row is switched off, and someone has to swit first-failure message never appeared. - **Now:** failures are counted from when the routine was last switched on, recorded in a new `routines.enabled_at` column. The migration sets it to the time of the upgrade, so any failure - streak already under way starts again from zero at that point. + streak already under way starts again from zero at that point. Adds migration + `0051_routine_enabled_at`. + ### Bots work as coworkers A Bot can now carry on without anyone watching it. It runs standing **Responsibilities** fed by @@ -147,6 +167,21 @@ refused in every mode, including `allow_all`, and the browser's WebRTC traffic n filter instead of around it. A computer run without an API server can set `EGRESS_POLICY_REQUIRED=0` to keep the old behaviour. +### Parallel Search is in the plugin catalogue + +Two catalogue entries reach Parallel's public-web search and extraction at +`https://search.parallel.ai/mcp`: **Parallel Search**, anonymous with provider-managed limits, and +**Parallel Search (API key)**, which sends a deployment credential as a bearer token. Nothing is +granted automatically. When a Bot holds both `web_search` and `web_fetch`, the built-in Bot guidance +describes them for public-web research, and the fintech example's Research Desk ships a +`research-public-web` skill that declares them. See +[Public-web research with Parallel](docs/parallel-research.md). + +### `lodash-es` is pinned to the patched 4.18.0 + +A root `overrides` entry pins the transitive `lodash-es` to 4.18.0, the patched release, wherever a +dependency pulls it in. + ### Provider and Bot lookups ignore inherited object properties Unknown names such as `constructor` and `__proto__` no longer return an inherited @@ -155,7 +190,9 @@ missing Bots raise the existing startup error. Configured providers and Bots are ### A malformed `%` in a stream URL no longer returns a 500 -A request to `/api/computers//stream` whose id held a broken percent-escape, such as `%zz`, made the server throw and answer 500. It is now treated as not matching the stream route and goes through normal routing. Valid ids behave as before. +A request to `/api/computers//stream` whose id held a broken percent-escape, such as `%zz`, +made the server throw and answer 500. It is now treated as not matching the stream route and goes +through normal routing. Valid ids behave as before. ### `start.sh` names the port to change on macOS @@ -340,6 +377,56 @@ the choice is a plain OpenAI key, and the OpenAI SDK only defaults an absent URL as the address. The Bot now falls back to `https://api.openai.com/v1` for an empty value, as its Anthropic branch already did for `ANTHROPIC_BASE_URL`. An OpenAI-compatible endpoint is unchanged. +### The live screen keeps reconnecting after it has recovered + +A dropped live screen retries five times, waiting half a second, then one, two, four and eight, and +then asks for Retry. The count of retries never went back to zero after a retry worked, so a screen +left open through five short drops over an afternoon gave up on the sixth, although each had +recovered within a second. The count now starts over once a reconnected screen shows a frame again, +so only five failures in a row end in Retry. + +### A failed save of a Bot's browser control leaves no copy behind + +The computer keeps who holds a Bot's browser, and its handoff requests, in one file per Bot under +the profiles volume, written to a temporary file first and renamed over it. When the write or the +rename failed, the temporary file stayed, a readable copy of that state beside the real one, and +every later failure added another. It is now removed whether or not the save succeeds, as the +learning setup and the model sign-in file already do. + +### Every Bot's provider defaults live in one spec file + +`shared/model-providers.json` now holds the provider facts and the default provider and model of +the thirteen Bots that read it: the three TypeScript Bots through `shared/model-providers.ts` and the +ten Python Bots through `shared/model_providers.py`. The Claude Agent SDK Bot does not read it. +`BOT_PROVIDER` and `BOT_MODEL` still win over both, and each Bot keeps its existing default, +including Mastra's `gpt-4o-mini`. Moving a Bot to a different model is one row in one file. + +Both loaders check the file against their own list of providers and refuse in the same words, so a +wrong row now stops all thirteen at startup, naming the key, where the Python Bots used to start +clean and meet it at their first model call. The Mastra Bot also refuses a `BOT_PROVIDER` it does +not recognize (such as `google`) instead of quietly answering through OpenAI with a different model. + +Compose used to substitute `gpt-5.5` for `agent-langgraph` whenever `BOT_MODEL` was unset, whatever +`BOT_PROVIDER` named. It now passes the unset value through, so the Bot's own row answers: an OpenAI +deployment keeps `gpt-5.5`, and a Google or Anthropic one stops being handed a model its vendor has +never heard of. The picked harness in Compose now also receives `GOOGLE_API_KEY` and +`GOOGLE_GENERATIVE_AI_BASE_URL`, so a harness picked on `BOT_PROVIDER=google` has its key. + +### Automatic Learning, on by default + +Every shipped Bot can now contribute completed conversations to a CopilotKit Intelligence Learning +container and receive the skills published from it. **Admin → Automatic Learning** chooses a default +container, overrides or excludes individual Bots, and pauses Learning. Chat, channels, routines and +handoffs all follow the same settings. + +Learning is on unless an administrator has saved it off, and a saved off stays off. It collects and +delivers nothing until a container exists in the Intelligence project and is assigned: +`bun scripts/setup-learning.ts` creates or reuses one called `openbot` and writes it to `.env`, or +set `CPK_INTELLIGENCE_LEARNING_CONTAINER_ID` (Helm: `config.learning.containerId`), or enter it on +the Admin page. OpenBot starts and chats without one. Skills are reviewed and published in +Intelligence; nothing is approved automatically. See +[Automatic Learning](docs/automatic-learning.md). Adds migration `0046_automatic_learning`. + ### Dictate messages and talk to a coworker in a live voice call Deployments can configure transcription separately from their Bots' models, with a waveform composer @@ -351,8 +438,10 @@ affects the person's microphone. See [configuration](docs/configuration.md#live- Voice summaries use the configured chat provider, including Anthropic keys and Claude or ChatGPT plan sign-in. Retrying a failed summary refreshes its sidebar preview without replacing newer -activity. The macOS app includes the microphone permission description and audio-input entitlement -needed for dictation and voice calls. +activity. A call the voice service turns down says why, such as "The voice service is busy. Please +retry shortly.", and a call that cannot be saved says "Could not save this voice chat." and stays on +its card to retry. Caps on a call's context, captions and answers cut between characters, never +inside an emoji. ### The Pydantic AI Bot answers on a plain OpenAI key or an Anthropic key @@ -379,14 +468,6 @@ expandable group instead of filling the conversation with screenshots. ## 0.0.15 -### The live screen keeps reconnecting after it has recovered - -A dropped live screen retries five times, waiting half a second, then one, two, four and eight, and -then asks for Retry. The count of retries never went back to zero after a retry worked, so a screen -left open through five short drops over an afternoon gave up on the sixth, although each had -recovered within a second. The count now starts over once a reconnected screen shows a frame again, -so only five failures in a row end in Retry. - ### A model provider's own sign-in can stand in for an API key `OPENBOT_MODEL_OAUTH_FILE` names a credential file holding a Google or xAI OAuth grant. Set it and @@ -409,13 +490,6 @@ still uses the compatibility endpoint. **A provider 403 no longer reads as an expired sign-in.** It usually means a missing project or resource permission, which signing in again cannot fix, so only a 401 now raises "sign in again". -### A voice call no longer cuts an emoji in half - -A voice call caps what it carries: the chat context it joins with, the live captions, and the answer -a delegated request comes back with. Each cap cut on UTF-16 code units, and an emoji is two of them, -so a cap landing inside one left half of it: a box at the edge of a caption, and a broken character -in what the voice model was given. The caps now cut between characters, as the chat's own do. - ### Desktop setup shows progress, chooses its own local ports, and can sign in to a provider Downloads report transferred bytes, every running step reports elapsed time, and running, completed @@ -439,14 +513,6 @@ removes that deployment's database volume and nothing else. Startup failures keep enough of the log to name the cause, with every secret value redacted, and carry a support link a whitelabel build can point elsewhere. -### A failed save of a Bot's browser control leaves no copy behind - -The computer keeps who holds a Bot's browser, and its handoff requests, in one file per Bot under -the profiles volume, written to a temporary file first and renamed over it. When the write or the -rename failed, the temporary file stayed, a readable copy of that state beside the real one, and -every later failure added another. It is now removed whether or not the save succeeds, as the -learning setup and the model sign-in file already do. - ### A Bot's image pull finds the Docker credential helper beside Docker A Docker install whose credential helper sits next to the `docker` binary rather than on the desktop @@ -455,28 +521,12 @@ resolved `docker`, and the directory holding what it points at when it is a syml to the PATH the engine is invoked with. Appended, so an existing helper still wins, and the inherited PATH is now kept rather than replaced, which it was not before. -### A voice call that cannot start says why - -When the voice service turned a call down, the provider wrote a message for the caller, such as -"The voice service is busy. Please retry shortly." when it answered 429, and the call route replaced -every one with "The voice service could not start a call. Please retry." The route now passes those -fixed messages on, as the dictation route already does. Any other failure still reads the generic -line, so nothing from an upstream response reaches the browser. - ### OpenBot starts only on the Bun it pins An installed or cached Bun that is not the pinned version is no longer accepted, on install and on every start, and OpenBot acquires its own copy instead. The version already on the machine is left exactly as it is and simply not used. -### A voice chat that could not be saved says so in a sentence - -Saving a finished voice call read the server's answer as JSON without a fallback. When something in -front of OpenBot answered instead, such as a proxy's 502 page, the call's card gave the JSON -parser's error as the reason (in Chrome, "Unexpected token '<' ... is not valid JSON"); an answer -without a saved session failed on a property read the same way. Both now read "Could not save this voice chat.", the -message the card already uses, and the call stays on the card to retry. - ### Organization sign-in survives a callback that arrives in pieces The loopback listener that receives an organization or provider sign-in read the callback once and @@ -484,15 +534,6 @@ gave up if the whole request had not arrived, and on Windows the accepted socket listener's non-blocking mode, so a timeout did not apply. A good sign-in could be answered "Sign-in did not match". Both paths now read until the request line is complete, with a real timeout. -### The Bots agree on one set of provider defaults - -The three TypeScript Bots now read a single shared list of provider facts instead of keeping their -own copies, which is what makes a default changeable in one place rather than in three. The Mastra -Bot also refuses a `BOT_PROVIDER` it does not recognize (such as `google`) instead of quietly -answering through OpenAI with a different model. The picked harness in Compose now receives -`GOOGLE_API_KEY` and `GOOGLE_GENERATIVE_AI_BASE_URL` as well, so a harness picked on -`BOT_PROVIDER=google` has the key it needs. - ### Compose file lists separate correctly on Windows The separator between Compose files fell back to `:` everywhere, which is right on macOS and Linux @@ -501,26 +542,6 @@ reachable on every platform now that a port overlay is passed, where before it w ## 0.0.14 -### One spec file, in every language - -`shared/model-providers.json` now holds the provider facts and every Bot's default provider and -model. The TypeScript Bots read it through `shared/model-providers.ts` and the ten Python Bots -through `shared/model_providers.py`, with `BOT_PROVIDER` and `BOT_MODEL` still winning over both -as they always have. Each Bot keeps its existing default, including Mastra's `gpt-4o-mini`. -Moving a Bot to a different model, or giving a Bot written in any other language its first one, -is editing one row in one file instead of one line per language. - -Both loaders check the file against their own list of providers, in both directions, and refuse in -the same words. A wrong row in the file used to stop the three TypeScript Bots while the ten -Python Bots started clean and met it at their first model call instead; all thirteen stop at -startup now, naming the key that is wrong. Adding a provider is one row in the file and one entry -to `PROVIDER_IDS` in each loader. - -Compose used to substitute `gpt-5.5` for `agent-langgraph` whenever `BOT_MODEL` was unset, whatever -`BOT_PROVIDER` named; it now passes the unset value through, so the Bot's row — or the moved -provider's default row — is what answers. An OpenAI deployment keeps the same `gpt-5.5` either way; -a Google or Anthropic one stops being handed a model its vendor has never heard of. - ### A tool cannot be granted for an app this deployment has not added Granting a Bot a connector's tool checked only that the person asking was an administrator, so a diff --git a/README.md b/README.md index 9ec581efa..c22e57007 100644 --- a/README.md +++ b/README.md @@ -35,9 +35,12 @@ your own machine. > **Runs on your machine.** Everything below is written for a laptop. `.env.example` carries `OPENBOT_SINGLE_USER=true`, which admits every request as one administrator, so a fresh clone reaches the product without registering an OAuth client first. [Sign-in](#sign-in) turns that off, and is required before anybody else can reach the deployment. -> **Do not want to build it yourself?** We will. Our engineers will stand OpenBot up inside your -> infrastructure, customize it into something that looks like your own product, and hand it back to you to -> keep changing. [**Start the conversation**](https://copilotkit.ai/talk-to-an-engineer?ref=openbot_readme). +> **We can build this for you, or with you, at your company.** Our engineers will stand OpenBot up +> inside your infrastructure, customize it into something that looks like your own product, and hand it +> back to you to keep changing. [**Start the conversation**](https://copilotkit.ai/talk-to-an-engineer?ref=openbot_readme). +> +> In the meantime, run all of it on your laptop, Intelligence included: CopilotKit's free Developer plan +> covers [running Intelligence locally](#intelligence-on-your-own-machine) in Docker on a Mac. ## What it is @@ -92,9 +95,11 @@ A Bot is any endpoint speaking [AG-UI](https://github.com/ag-ui-protocol/ag-ui), `openbot` Learning container, then writes both settings to `.env`. It preserves custom container assignments and refuses to replace an existing OpenBot key. For fresh self-hosted setup, set `INTELLIGENCE_API_URL` to that deployment's - HTTPS API origin and run the helper without the two `npx` commands. It opens - an isolated Chrome or Edge window for your usual sign-in and project choice; - install either browser first. Existing key-only setups can assign a container + HTTPS API origin (or `http://localhost` for a local one) and run the helper + without the two `npx` commands. It opens an isolated Chrome or Edge window for + your usual sign-in and project choice; install either browser first. It writes + the API URL, the key and the container, not `INTELLIGENCE_GATEWAY_WS_URL`, so + set that one to the same deployment yourself. Existing key-only setups can assign a container through [Admin → Automatic Learning](docs/automatic-learning.md). Managed Intelligence needs no separate licence token. @@ -116,9 +121,31 @@ A Bot is any endpoint speaking [AG-UI](https://github.com/ag-ui-protocol/ag-ui), 5. Open . -`scripts/start.sh` starts Docker services, applies migrations, starts the API server on port 3001, starts the app on port 3010, and checks that the services answer their own health routes before printing next steps. +`scripts/start.sh` starts Docker services, applies migrations, starts the API server on port 3001, starts the routine worker, starts the app on port 3010, and checks that the services answer their own health routes before printing next steps. -`scripts/stop.sh` takes the same things down, including each Bot's computer, which compose does not own. Nothing is deleted: the database, the Bots' files and their browser profiles are volumes. +`scripts/stop.sh` takes the same things down, including each Bot's computer, which compose does not own; `--keep-computers` leaves the computers running. Nothing is deleted: the database, the Bots' files and their browser profiles are volumes. + +### Intelligence on your own machine + +To try Intelligence on one Mac instead of the managed service, follow +[Evaluate Intelligence locally](https://docs.copilotkit.ai/intelligence/self-hosting-local). The free +[Developer plan](https://www.copilotkit.ai/pricing) qualifies. The local licence lasts 30 days and +`npx copilotkit@latest local renew` renews it as often as you need, while your account is in good +standing. It is a preview for macOS with Docker Desktop, +not a production installation. Run its +`npx copilotkit@latest local connect` from the OpenBot root, where `.env` is. + +That command writes `INTELLIGENCE_API_URL`, `INTELLIGENCE_GATEWAY_WS_URL` and +`CPK_INTELLIGENCE_API_KEY` to `.env`. OpenBot reads the first two as they are, but takes its project +key only from `INTELLIGENCE_API_KEY` and refuses to start without it, so copy the `cpk-...` value +across before `bash scripts/start.sh`: + +```sh +INTELLIGENCE_API_KEY=cpk-... # the value of CPK_INTELLIGENCE_API_KEY +``` + +To give the Bots a Learning container on that stack, see +[Automatic Learning](docs/automatic-learning.md). ## Deploy it @@ -157,10 +184,19 @@ Leave `EMBEDDED_POSTGRES` off and set `DATABASE_URL` to point at a database you | `/` | Start and browse channels. | | `/agents` | Create, edit, duplicate, pin, hide, delete, and launch coworkers. | | `/channel/:id` | Converse with one coworker, watch its screen, and see what it ran. | +| `/group/new` | Start one conversation with two or more Bots. | | `/bot` | Direct chat with a Bot; `?agent=` selects one. | +| `/bots` | What each of your Bots is doing, what it needs from you, and whether it is paused. | +| `/team-bots` | Bots your teammates published, and the ones you share. | +| `/responsibilities` | Give a Bot a lasting goal, follow its progress, and decide when it should work. | +| `/reachability` | Continue a conversation in Slack, Microsoft Teams, by text message, or on your phone. | +| `/memory` | Review what your Bots remember, and choose which connected apps can contribute facts. | +| `/approvals` | Choose what your Bots may do, and review actions waiting for you. | | `/skills` | Create and enable personal skills. | | `/routines` | See the routines that are standing, and stop one. | | `/settings` | User preferences. | +| `/settings/connected-accounts` | Connect your apps so your Bots can work with them. | +| `/settings/passwords` | Logins you saved when signing a Bot in to a website. | | `/admin/credentials` | Store write-only encrypted credentials. | | `/admin/computers` | View, stop, and reset Bot computers. | | `/admin/boundaries` | Configure browser/file/MCP action policy. | @@ -171,6 +207,7 @@ Leave `EMBEDDED_POSTGRES` off and set `DATABASE_URL` to point at a database you | `/admin/learning` | Assign Learning containers to Bots, or pause Automatic Learning. | | `/admin/people` | List, promote, demote, and remove people who have signed in. | | `/admin/identity-providers` | Register a company SAML or OIDC provider, routed by email domain. | +| `/admin/enterprise` | What members may use, how Bot computers reach the network, and what is recorded about it. | | `/admin/audit` | Review permitted, refused, and failed actions. | ## Features @@ -184,7 +221,7 @@ Leave `EMBEDDED_POSTGRES` off and set `DATABASE_URL` to point at a database you - **Secrets never enter the transcript**: the trail records that a secret was requested and how long it was, not what it said. - **Bring your own agent**: any AG-UI endpoint is a Bot, on a framework or hand-written. Endpoints are validated with the same target checks used for browser navigation, and an auth header is stored write-only. - **Components instead of prose**: compiled React components live in `app/src/components/gallery/`, sandboxed ones are authored in `/admin/playground` and published with no deployment. Every call asks the server whether the component exists, is published, and is not withheld from that Bot. Data functions are granted per component. -- **Governed MCP**: Google Drive and Notion ship in the catalogue, and Composio brokers a few hundred more apps behind one account, each reached as the person asking. The catalogue carries only vendors this deployment stands behind, so adding one is a review of that vendor. Custom servers must pass URL checks; unknown tools and custom-server tools are treated as writes, and a catalogue tool the server advertises but does not name as a write classifies as a read. A Bot is told which connectors exist here and which it holds, so it says it has not been granted one rather than browsing to the vendor's website. +- **Governed MCP**: Google Drive, Notion and [Parallel Search](docs/parallel-research.md) ship in the catalogue, and Composio brokers a few hundred more apps behind one account, each reached as the person asking. The catalogue carries only vendors this deployment stands behind, so adding one is a review of that vendor. Custom servers must pass URL checks; unknown tools and custom-server tools are treated as writes, and a catalogue tool the server advertises but does not name as a write classifies as a read. A Bot is told which connectors exist here and which it holds, so it says it has not been granted one rather than browsing to the vendor's website. - **Skills are instructions, not capabilities**: personal skills attach only to Bots their author owns, deployment skills are admin-owned, and both are invoked with `/` in the composer. A Bot granted the shipped `skill-creator` skill can write one with you in the conversation, and saves it only when you press the button on the card. - **Sign in with what your company already has**: Google, Microsoft or Okta from the environment, or a company's own SAML or OpenID Connect provider registered while the deployment runs and routed by email domain. Any one turns sign-in on; several may be configured at once. - **Decide who gets in**: `/admin/people` lists everybody who has signed in, promotes and demotes them, and removes access, which ends the session they are using and stops the next sign-in. Every change is on the audit trail. @@ -192,7 +229,7 @@ Leave `EMBEDDED_POSTGRES` off and set `DATABASE_URL` to point at a database you - **Credentials encrypted at rest**: stored through `/admin/credentials`, never returned by an API, and redacted from audit events. - **Loopback by default**: computers bind to `127.0.0.1` and require a per-container token, so nothing reaches a logged-in browser by knowing its port. The supervisor binds there too, because it holds the Docker socket and its token is a shared secret rather than a network boundary. - **Durable threads and memory**: conversations survive restarts through CopilotKit Intelligence, and each deployment stamps the threads it owns. -- **Routines**: ask a Bot to do something on a schedule and it does, running as you, in the channel you asked in. A 15-minute floor and a cap of 20 enabled routines keep a sentence from scheduling more than a person meant, and ten failures in a row switch a routine off rather than burn model spend forever. Needs a worker process; see [docs/routines.md](docs/routines.md). +- **Routines**: ask a Bot to do something on a schedule and it does, running as you, in the channel you asked in. A 15-minute floor and a cap of 20 enabled routines keep a sentence from scheduling more than a person meant, and ten failures in a row switch a routine off rather than burn model spend forever. Needs a worker process, which `scripts/start.sh` starts locally; see [docs/routines.md](docs/routines.md). ## Bring your own agent @@ -270,6 +307,7 @@ Full reference: [docs/configuration.md](docs/configuration.md). | `agent-bot` | 4200 | Proof-of-concept AG-UI Bot. | | `agent-langgraph` | 4201 | LangGraph AG-UI Bot. | | `supervisor` | 4500 host / 4300 container | Creates and manages one computer per Bot. | +| `worker` | none | Fires due routines by handing each run to the server. | | PostgreSQL with pgvector | 5432 | Product data, policy, audit, credentials, grants, channels, and component metadata. | | CopilotKit Intelligence | external | Durable threads and memory. | @@ -321,8 +359,9 @@ A company's own SAML or OpenID Connect provider is registered while the deployme Admin → Identity providers, and routed by email domain. An OIDC registration needs every host in the provider's discovery document listed in `TRUSTED_ORIGINS`, not only the issuer. -- `INITIAL_ADMIN_EMAILS` is required, because nothing else grants the administrator role and no - screen can promote somebody afterwards. It is re-read on every sign-in, so editing it takes effect +- `INITIAL_ADMIN_EMAILS` is required, because it is what grants the first administrator. After that + `/admin/people` can promote others, but an address named here stays an administrator and cannot be + demoted or removed from that screen. It is re-read on every sign-in, so editing it takes effect the next time that person signs in. - `MICROSOFT_OAUTH_TENANT_ID` defaults to `common`, which admits personal Microsoft accounts as well as work ones. On a multi-tenant app registration Entra may send no `email` claim at all, so @@ -330,7 +369,7 @@ provider's discovery document listed in `TRUSTED_ORIGINS`, not only the issuer. sign-in is refused and the reason is logged: add `email` as an optional claim, or use your directory GUID here. - A half-configured provider is refused at start-up rather than at somebody's first attempt to sign - in: a client id with no secret, a secret shorter than 32 characters, or an Okta issuer with no + in: a client id with no secret, a `BETTER_AUTH_SECRET` shorter than 32 characters, or an Okta issuer with no credentials behind it. - **SAML and OIDC** are registered while the deployment runs rather than configured here. Sign in as an administrator and go to Admin → Identity providers with the metadata your identity team gave @@ -393,7 +432,10 @@ Use `bash scripts/start.sh` for the whole stack and `bash scripts/stop.sh` to ta - [docs/configuration.md](docs/configuration.md) - [docs/development.md](docs/development.md) - [docs/coworkers.md](docs/coworkers.md) +- [docs/routines.md](docs/routines.md) +- [docs/automatic-learning.md](docs/automatic-learning.md) - [docs/deployment.md](docs/deployment.md) +- [charts/openbot/README.md](charts/openbot/README.md): the Helm chart - [docs/releasing.md](docs/releasing.md) - [docs/parallel-research.md](docs/parallel-research.md): public-web search and extraction through Parallel diff --git a/charts/openbot/README.md b/charts/openbot/README.md index 1002ba719..a8b1ae41e 100644 --- a/charts/openbot/README.md +++ b/charts/openbot/README.md @@ -297,9 +297,16 @@ installing on somebody's bare-metal cluster. ## Refused at install, not in a crash loop The chart fails the install, naming the value to change, when: there is no database or two of them; -nobody would be an administrator; `singleUser` is combined with a public URL; both an Ingress and an -HTTPRoute are enabled; both `externalSecrets` and an existing Secret are named; a Bot endpoint is -named with no token to call it with; a browser is asked for inside more than one API replica; or +nobody could sign in or nobody would be an administrator; an identity provider is configured without +`secrets.betterAuthSecret` or `config.publicUrl`; `singleUser` is combined with a public URL or a +public LoadBalancer; the Intelligence URLs or `secrets.intelligenceApiKey` are missing; +`secrets.keyEncryptionKey` is not a base64 32-byte value, or is the public example key; there is no +`secrets.computerToken`; `computers.mode` is `external` with no `computers.url`, or `sandbox` on a +cluster with no Sandbox CRD; both an Ingress and an HTTPRoute are enabled; both `externalSecrets` and +an existing Secret are named; a Bot endpoint is named with no token to call it with; a browser is +asked for inside more than one API replica; `networkPolicy.enabled` is set with an external database +and no `networkPolicy.extraEgress`, or with `computers.mode: sandbox` and no +`networkPolicy.kubernetesApiCidr`; a computer proxy is named that the network policy blocks; or `routines.enabled` is set with no `secrets.workerSharedSecret` — and, on `externalSecrets`, no `worker-shared-secret` key named for it to read instead. One combination gets no refusal at all: `secrets.existingSecret` with `routines.enabled`, because the Secret this chart would otherwise @@ -323,9 +330,29 @@ secrets: The token travels on every call and is required whenever a url is set. With an existing Secret or an external store, the key is `managed-agent-token`. -Left empty, this deployment has the Bots its tenant package declares as built-in and no others. A -package entry pointing at an endpoint that resolves to nothing is dropped rather than registered as a -coworker nobody can talk to. +Left empty, this deployment has the Bots its tenant package declares as built-in, and a coworker +somebody creates in the app without an endpoint runs here on its role description. A package entry +pointing at an endpoint that resolves to nothing is dropped rather than registered as a coworker +nobody can talk to. + +## Automatic Learning + +Learning is on by default, and it needs a container in the Intelligence project before it collects or +delivers anything. `config.learning.containerId` sets the initial default container +(`CPK_INTELLIGENCE_LEARNING_CONTAINER_ID`) and `config.learning.revision` the skills revision +(`CPK_INTELLIGENCE_SKILLS_REVISION`). Both are only starting values: settings saved in Admin → +Automatic Learning take precedence, including switching it off. See +[docs/automatic-learning.md](../../docs/automatic-learning.md). + +## Settings with no value of their own + +Several server settings have no dedicated chart value and are passed through `config.extraEnv` or +`config.extraEnvFrom`: OpenTag delivery to Slack and Teams (`OPENTAG_URL`, `OPENTAG_SHARED_SECRET`), +text messages (`TWILIO_*`), mobile push (`EXPO_ACCESS_TOKEN`, `EXPO_PROJECT_ID`), inbound email for +Responsibilities (`OPENBOT_INBOUND_EMAIL_DOMAIN`, `OPENBOT_INBOUND_EMAIL_SNS_TOPIC_ARNS`), SCIM +(`SCIM_*`), OpenTelemetry export (`OPENBOT_OTEL_EXPORT`, `OTEL_*`) and `DELIVERY_PUBLIC_URL`. +`.env.example` describes each. Put the secret ones in a Secret and name it in `config.extraEnvFrom` +rather than writing them into a values file. ## Slack, Teams, text messages, push, SCIM, inbound email and OpenTelemetry diff --git a/docs/README.md b/docs/README.md index d39560335..6f76028ce 100644 --- a/docs/README.md +++ b/docs/README.md @@ -8,6 +8,7 @@ Start with the root [README](../README.md), then use these references: - [Coworkers](coworkers.md): durable Bot profiles, channels, visibility, deletion, and external AG-UI registration. - [Routines](routines.md): standing instructions a Bot runs on a schedule, the worker that fires them, and who they run as. - [Automatic Learning](automatic-learning.md): which Bots contribute conversations to a Learning container, and receive its published skills. +- [Parallel research](parallel-research.md): public-web search and extraction through the Parallel Search connector. - Plugins, one connector per page — what an administrator registers, what each person consents to, and what the failures mean: - [Composio](plugins/composio.md): the broker, and so the one page here that is a catalogue of apps rather than a single connector. - [Google Drive](plugins/google-drive.md) diff --git a/docs/architecture.md b/docs/architecture.md index fc6111da8..f98b90e6f 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -23,7 +23,7 @@ Regenerate it with `bun run diagram` after changing anything it shows. | PostgreSQL with pgvector | 5432 | Product data, audit rows, credentials, policy, grants, channels, and components. | | CopilotKit Intelligence | external | Durable threads, memory, and realtime gateway. | -`scripts/start.sh` starts PostgreSQL, `agent-computer`, `agent-bot`, `agent-langgraph`, and the supervisor through Docker Compose, then starts `server` and `app` on the host. +`scripts/start.sh` starts PostgreSQL, `agent-computer`, `agent-langgraph`, `agent-bot` (skipped when `BOT_PROVIDER=anthropic`, since that sample only takes OpenAI keys) and the supervisor (unless `OPENBOT_ONE_COMPUTER_EACH=false`) through Docker Compose, and runs `migrate`. It then starts `server`, the routines worker (`bun worker/src/index.ts`, the local stand-in for the routines CronJob) and `app` on the host. The compose file also defines optional SPIRE services. `start.sh` does not start them. @@ -80,14 +80,16 @@ Compose puts it on a different network from PostgreSQL. A Bot has a shell, and a With `COMPUTER_SUPERVISOR_URL`, each Bot gets its own computer container, workspace volume, and browser profile. Without it, all Bots share `AGENT_COMPUTER_URL`. +Each computer also filters its own outbound connections. The browser and shell commands go through a filter proxy that applies the Bot's network policy (`allow_all`, `defaults_plus_allowlist`, `allowlist_only` or `deny_all`, which is what a Bot gets when its owner's Cloud network access is switched off). The server pushes the policy before the computer acts and whenever it wakes, and until one arrives nothing may leave. `EGRESS_POLICY_REQUIRED=0` is for a computer run with no API server. Cloud metadata addresses are refused in every mode. A process that ignores the proxy variables and opens a raw socket is not stopped by this; that is a Kubernetes NetworkPolicy's job. + A command on the computer inherits PATH, locale and terminal names, and the proxy variables, not the rest of the process environment. Userinfo is stripped from a proxy URL. `COMPUTER_SHELL_ENV` names anything else a deployment wants passed. The supervisor exposes only ensure, stop, reset, and list operations. It holds the Docker socket, so do not expose it outside the deployment network: Docker Compose binds it to `127.0.0.1:4500`, and a deployment running the server inside the compose network reaches it as `supervisor:4300` and needs no published port at all. Set `COMPUTER_RUNTIME=runsc` to run computers under gVisor on hosts that support it. ## What started a run -Every audit row records on whose authority an action was taken. A routine asserts its owner, and a -hop between Bots asserts the person who began the conversation, so that column alone cannot say +Every audit row records on whose authority an action was taken. A routine or a Responsibility asserts +its owner, and a hop between Bots asserts the person who began the conversation, so that column alone cannot say whether anybody was there when it happened. An interactive run has somebody watching who will notice a wrong tool call; an unattended one does not, which is the case worth being able to find. @@ -98,9 +100,12 @@ Each row therefore also names what caused it: | `person` | none | Somebody was in the room. The default. | | `deployment` | none | The deployment itself, at start-up or refusing a caller it could not identify. | | `routine` | the routine's id | A schedule fired it, as its owner, with nobody there. | +| `responsibility` | the Responsibility's id | One of a Bot's standing Responsibilities was triggered, as its owner, with nobody there. | +| `memory` | the memory source's id | Memory read a connected app to import facts, as the source's owner. | | `handoff` | the Bot that handed on | Another Bot asked for this, on the person's behalf. | -The Audit screen filters on it, and **Nobody watching** is `routine` and `handoff` together, which is +The Audit screen filters on it, and **Nobody watching** is `routine`, `handoff`, `responsibility` and +`memory` together, which is the question of what ran on somebody's authority while they were away. `deployment` is deliberately outside that filter: a boundary held at start-up is not work done on anybody's behalf. @@ -117,9 +122,9 @@ person, and a stream that stalls all say the same thing without each being told cannot relabel its own run, because the assertion is signed by the deployment, and a kind this deployment does not write is read as a person rather than kept. -Computer actions are the one family that carries no initiator, and correctly so: the computer tools -are browser actions, executed by the person's own session, so a headless run has no way to drive the -computer at all today. A row there is a person's because a person's browser wrote it. +Computer actions carry the initiator too. A run with no browser attached (a routine, a Responsibility, +a group turn) is given the computer tools server-side, with the Bot and the person bound before the +model supplies any arguments, and the gateway writes that run's initiator on every decision. ## Human control and secrets @@ -129,7 +134,7 @@ Handovers are audited as control events: - `computer.control_taken` - `computer.control_released` -While a person controls the browser, Bot actions are refused rather than queued. +While a person controls the browser, Bot actions are refused rather than queued. A run with no browser attached pauses instead, waiting on the handover. Secret entry is separate from chat content. The audit trail records that a secret was requested or supplied and the character count, not the secret value. @@ -149,6 +154,8 @@ A coworker is a durable Bot profile: A channel is a conversation with one coworker and a CopilotKit Intelligence thread mapping. Starting a new channel creates a new thread. +A group conversation holds several Bots. An Intelligence thread admits exactly one agent, so each Bot gets a thread of its own in the group (`group_bot_threads`), and what the people in it read is a shared transcript (`group_messages`) with every reply attributed to its speaker. A Bot that names a peer as `@Name` hands the conversation to it under the same grant, caps and audit rows as any other hop. + Who may reach one is decided by membership: every channel route resolves the caller in `channel_memberships` and refuses without a row. `channels.allowed_groups` is declared in the tenant package and stored, and is not part of that decision — `users.groups` is never populated by @@ -259,20 +266,18 @@ Reaching a second Bot spends a model call, may wake a computer and can fan out; already in the conversation costs nothing and cannot be aimed anywhere they cannot see. A deployment able to switch off the safe exit and keep the expensive one would be backwards. -Both tools are for Bots that run here. A Bot at its own endpoint runs its own loop and is handed -descriptions of the tools it may call back for, and the callback path executes MCP refs only, so -neither `message_bot` nor `ask_person` can reach it. +A Bot at its own endpoint runs its own loop. A call it makes back to `message_bot` or `ask_person` +through the signed callback route is dispatched the same way as one from a Bot running here. Being +handed work is not the same as being able to hand it on, so the target of a grant may live at +its own endpoint. -It is the Bot **doing the asking** that has to run here. Being handed work is not the same as being -able to hand it on, so the target of a grant may perfectly well live at its own endpoint. A grant -whose *grantee* is remote is refused rather than stored, so an administrator finds out at the point -of granting rather than from a Bot that never hands anything on. +One limit remains on the asking side. The grant check a hop goes through (`botsReachableFrom`) counts +only grants held by built-in Bots, so a Bot at its own endpoint can be granted another Bot and can ask +its person, but every hop it attempts is refused as not granted. -That is a real limit rather than a detail, and it is worth being plain about which Bots it leaves -out: **a Bot created through the UI is a remote one**, because creating a coworker here means -pointing it at an AG-UI endpoint. Only Bots a tenant package declares as built-in run in this -process. So on a deployment with no package, nothing can be granted `message_bot` at all, and the -screens say nothing about why. +A coworker created through the UI with an endpoint is a remote one. Created without one, it runs at +the managed Bot's endpoint when the deployment has one, and otherwise runs here as a built-in Bot on +its role description. Who "a person" is, is a seam. This template answers the person in the conversation, which is the only answer a template can give honestly; a company has an on-call rota or a duty desk, and that is a @@ -287,7 +292,7 @@ MCP servers and skills share the plugin grant table, but they have different own - MCP tools are admin-governed because they can reach external systems with stored credentials. - Skills are reusable instructions. A person can create personal skills and attach them only to Bots they own. Administrators create deployment skills. -The curated MCP catalogue contains Google Drive and Notion. Custom MCP servers must pass URL checks; unknown tools and custom-server tools are treated as writes unless positively classified as reads. +The curated MCP catalogue contains Parallel Search (anonymous, or with this deployment's API key; see [parallel-research.md](parallel-research.md)), Google Drive, Notion, and the built-in Routines server. Custom MCP servers must pass URL checks; unknown tools and custom-server tools are treated as writes unless positively classified as reads. A catalogue entry says whose credential a Bot reaches it with, which is a different question from whether it is reachable at all. A deployment-wide token answers the same for everybody; Google Drive and Notion are both `user-oauth`, so a Bot reaches them as the person asking and sees only what that person can see. An administrator enabling the connector and a person connecting their own account are two decisions, and neither can be made for the other. See [Google Drive](plugins/google-drive.md) and [Notion](plugins/notion.md). diff --git a/docs/configuration.md b/docs/configuration.md index de2486a9e..7561713de 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -42,7 +42,7 @@ at `agent-langgraph` on a laptop. | Variable | Default | Meaning | | -------------------- | ---------------------------------- | ------------------------------------------------------------------- | -| `PORT` | `3001` | API server port. | +| `PORT` | `3001` | API server port. `SERVER_PORT` names the same port; set either, or both to the same value, or the server refuses to start. | | `NODE_ENV` | unset | `production` refuses the example `KEY_ENCRYPTION_KEY`. It does not decide whether sign-in is required; see `OPENBOT_SINGLE_USER`. | | `TENANT_PACKAGE_DIR` | `../examples/fintech` | Tenant package directory, resolved from `server/`. | | `DEPLOYMENT_ID` | the tenant package's id | Names this deployment inside a shared Intelligence project. | @@ -63,7 +63,7 @@ at `agent-langgraph` on a laptop. | `AUDIT_RETENTION_DAYS` | unset | Whole number of days to keep audit rows; older ones are removed. Unset keeps the trail forever. | | `WORKER_SHARED_SECRET` | unset; `start.sh` uses a fixed local default | The secret the routines worker presents to fire a due routine. Without it the server refuses every handoff, whether or not a worker exists to send one. | | `OPENBOT_GENERATIVE_UI` | unset (capability on) | Set `false` or `0` to stop Bots from answering with generated interfaces. | -| `OPENBOT_ACCESSIBILITY_DISABLED` | `true` or `1` stops naming OpenBot on the analytics the runtime already sends. | +| `OPENBOT_ACCESSIBILITY_DISABLED` | unset | `true` or `1` stops naming OpenBot on the analytics the runtime already sends. | | `COMPOSIO_API_KEY` | unset | One key for the whole deployment, for the broker that holds people's accounts for a few hundred apps. Unset, there is nothing to connect, nothing to grant and no Composio tool for a Bot to call; what remains is one row that goes nowhere, under **More apps** on the admin Plugins page, naming this variable. See [Composio](plugins/composio.md). | **`OPENBOT_GENERATIVE_UI`** enables generated interfaces by default: streamed HTML/CSS/JavaScript @@ -132,7 +132,8 @@ in-cluster Service address. `shared/model-providers.json` is one file, and every language in the box reads it: the TypeScript Bots through `shared/model-providers.ts`, the Python Bots through `shared/model_providers.py`, and -any other implementation straight as JSON. It has two sections — the facts per provider, and the +any other implementation straight as JSON. The one Bot that does not is `agent-claude-sdk`, which has +no `bots` row. It has two sections: the facts per provider, and the provider and model each Bot runs: ```json @@ -233,7 +234,7 @@ the chat model provider; neither `OPENAI_BASE_URL` nor `OPENAI_API_KEY` is inher For OpenAI, set the base URL to `https://api.openai.com/v1` and select an available transcription model such as `gpt-transcribe`. A compatible local endpoint can instead use -`http://localhost:8000/v1` and its own model name. Compatibility with chat completions alone does +`http://127.0.0.1:8000/v1` and its own model name. Compatibility with chat completions alone does not imply transcription support. See the [OpenAI transcription guide](https://developers.openai.com/api/docs/guides/speech-to-text). The composer shows the microphone when the service is configured. Browser microphone access needs @@ -330,8 +331,11 @@ policies apply. See [OpenAI WebRTC](https://developers.openai.com/api/docs/guide | `OPENBOT_ORGANIZATION_AUTH_URL` | An OpenBot deployment that verifies employee identity and roles. Decides sign-in ahead of `OPENBOT_SINGLE_USER`. | **Some variables belong to the desktop app, not to you.** A desktop installation writes these into -its own deployment's `.env` and owns their values: `OPENBOT_MODEL_OAUTH_FILE`, `CHATGPT_AUTH_FILE` -and `CLAUDE_CODE_OAUTH_TOKEN`. When a model provider is connected by OAuth rather than by key, the +its own deployment's `.env` and owns their values: `OPENBOT_MODEL_OAUTH_FILE`, `CHATGPT_AUTH_FILE`, +`CLAUDE_CODE_OAUTH_TOKEN`, and the `PICKED_HARNESS_*` names that describe the Bot picked during setup +(`PICKED_HARNESS_IMAGE`, `PICKED_HARNESS_URL`, `PICKED_HARNESS_PORT` and the rest). The server reads +`PICKED_HARNESS_IMAGE` and `PICKED_HARNESS_URL` to hand that Bot `MANAGED_AGENT_TOKEN`, and refuses to +start when they are set without it. When a model provider is connected by OAuth rather than by key, the desktop also points `OPENAI_BASE_URL` at OpenBot's own loopback route and sets `OPENAI_API_KEY` to a local proxy credential rather than a provider key, so those two do not mean what the table above says in that mode. A server you configure yourself is unaffected by all of this. @@ -340,7 +344,7 @@ in that mode. A server you configure yourself is unaffected by all of this. nothing to sign anybody in and does not say that was deliberate refuses to start, naming what to configure, because a public URL where every visitor is an administrator fails silently. `NODE_ENV` does not enter into it. `.env.example` ships the line switched on, so a clone runs with no -configuration at all. +configuration at all. `OPENBOT_DEV_NO_AUTH=true`, the flag's former name, is still honoured. **But not on a public address.** The flag says you meant an open deployment; it does not say who can reach it. If `OPENBOT_PUBLIC_URL`, `OPENBOT_APP_URL` or any `TRUSTED_ORIGINS` entry is an address @@ -509,6 +513,18 @@ SNS asks. Both are needed. With either missing, the email route is not mounted, and an email trigger has no address: its page and the Bot both say inbound email is not configured on this deployment. +## Automatic Learning + +Learning is on by default; with no container assigned, nothing is collected or delivered and the +deployment still starts and chats. Both variables are optional defaults that an administrator's saved +settings under **Admin → Automatic Learning** override, including a saved off. See +[automatic-learning.md](automatic-learning.md). + +| Variable | Meaning | +| ---------------------------------------- | --------------------------------------------------------------------------------------------- | +| `CPK_INTELLIGENCE_LEARNING_CONTAINER_ID` | Default Learning container for this deployment's Bots. 1 to 64 lowercase letters, digits and single hyphens, or the server refuses to start. | +| `CPK_INTELLIGENCE_SKILLS_REVISION` | An exact published skills revision to pin. Read only beside a container id. Unset follows the latest. | + ## OpenTelemetry export Every audit row is also sent as an OTLP log record over HTTP, so a SIEM or an OpenTelemetry @@ -548,7 +564,7 @@ is still the record. `COMPUTER_SANDBOX` is not the cluster sandbox provider. A Kubernetes deployment can instead run each computer as a sandboxed pod, selected by `COMPUTER_SANDBOX_NAMESPACE` with `COMPUTER_SANDBOX_IDLE_AFTER` and `COMPUTER_SANDBOX_TEMPLATE_FILE` beside it; those are set by the Helm chart, not by Compose, and -are covered in [charts/openbot/README.md](../../charts/openbot/README.md). The similarly named +are covered in [charts/openbot/README.md](../charts/openbot/README.md). The similarly named `COMPUTER_SANDBOX` above only toggles Chromium's own process sandbox on a Docker computer. Changing `COMPUTER_BROWSER_MODE` affects new supervised computers. A computer that already exists is @@ -842,15 +858,24 @@ agents: title: Risk & Compliance role_description: Investigate policies and controls. type: remote-ag-ui - endpoint: ${MANAGED_AGENT_AG_UI_URL} + endpoint: ${MANAGED_AGENT_AG_UI_URL:-} ``` Each agent requires `id`, `name`, `title`, `role_description`, and `type`. -| Type | Required field | -| -------------- | --------------- | -| `built-in` | `system_prompt` | -| `remote-ag-ui` | `endpoint` | +| Type | Required field | +| --------------- | --------------- | +| `built-in` | `system_prompt` | +| `remote-ag-ui` | `endpoint` | +| `remote-mastra` | `endpoint`, and optionally `remote_agent_id` to pick one agent on that Mastra server | + +A remote agent whose `endpoint` resolves to an empty string is left out rather than refused, and it +is dropped from every channel's `permitted_agents` too. That is how the example package carries a row +for a Bot that only exists once something is configured. + +An agent may list `skills:`, slugs of skills this package ships in `skills.yaml`. A slug the package +does not ship, or the same slug twice, stops the server. Two agents with the same `id` in +`agents.yaml` stop it as well. The two types are told different amounts, which is easy to miss. A `built-in` agent gets its `system_prompt`; a `remote-ag-ui` agent has none, and its `role_description` is the only instruction @@ -909,7 +934,7 @@ channels: allowed_groups: [risk, compliance] ``` -Each channel requires `id`, `name`, `description`, `permitted_agents`, and `allowed_groups`. Every `permitted_agents` entry must match an agent id. +Each channel requires `id`, `name`, `description`, `permitted_agents`, and `allowed_groups`. Every `permitted_agents` entry must match an agent id. A channel `id` declared twice, or an agent listed twice in one channel, stops the server. `allowed_groups` is validated and stored, and nothing reads it. It decides nothing today, and a deployment that writes one must not treat it as an access control. Both halves of that control are @@ -932,7 +957,7 @@ model: default_model: gpt-5.6-terra ``` -`provider` must be `openai`. `credential_secret_ref` is a reference to a stored credential, not a credential value. `default_model` is passed through as written, so an OpenAI-compatible endpoint reached through `OPENAI_BASE_URL` takes the name that endpoint publishes. +`provider` must be `openai` or `anthropic`. `credential_secret_ref` is a reference to a stored credential, not a credential value. `default_model` is passed through as written, so an OpenAI-compatible endpoint reached through `OPENAI_BASE_URL` takes the name that endpoint publishes. ### `knowledge.yaml` @@ -968,7 +993,7 @@ Refs are `serverId/toolName`, the same form a grant is written in. A package may One slug is load-bearing. A Bot granted `skill-creator` is offered the four tools that let a conversation end in a saved skill, so a package shipping that skill should also grant it to a Bot in `agents.yaml` — shipping it and granting it to nobody boots a deployment where writing a skill in the composer quietly does nothing. It declares no `tools`, and should not: those four are the app's own rather than a connector's, so they are not `serverId/toolName` refs. See [architecture.md](architecture.md#writing-a-skill-in-a-conversation). -Slugs are lowercase letters, digits and hyphens. If a package ships a slug somebody in the deployment already wrote a skill under, theirs keeps the name, the package loses that skill, and startup continues. +Slugs are 2 to 40 lowercase letters, digits and hyphens, starting and ending with a letter or digit, the same rule the skills API and the app's form apply. A slug declared twice in `skills.yaml` stops the server. If a package ships a slug somebody in the deployment already wrote a skill under, theirs keeps the name, the package loses that skill, and startup continues. Omit the file entirely for a package with no skills. diff --git a/docs/coworkers.md b/docs/coworkers.md index 5f4558c5c..d73d5df29 100644 --- a/docs/coworkers.md +++ b/docs/coworkers.md @@ -4,13 +4,14 @@ A coworker is a Bot with a durable profile and standing role. The role is sent w ## Data model -| Piece | Table | Purpose | -| -------------------- | ------------------------------- | --------------------------------------------------------------------- | -| Runtime agent | `agents` | AG-UI endpoint and optional key reference. | -| Profile | `agent_profiles` | Name, title, role, avatar seed, owner, visibility, and soft deletion. | -| Personal roster | `agent_preferences` | Per-user hidden state. | -| Channel | `channels` | Conversation membership and coworker binding. | -| Intelligence mapping | `intelligence_channel_mappings` | Channel-to-thread mapping. | +| Piece | Table | Purpose | +| -------------------- | ------------------------------- | ---------------------------------------------------------------------------------------- | +| Runtime agent | `agents` | Type (built-in, remote AG-UI or remote Mastra), endpoint or prompt, optional key reference. | +| Profile | `agent_profiles` | Name, title, role, avatar seed, owner, visibility, and soft deletion. | +| Personal roster | `agent_preferences` | Per-user hidden and pinned state. | +| Team publication | `team_bot_publications` | A Bot its owner has published to the whole team or to named people and groups. | +| Channel | `channels` | Conversation membership and coworker binding. | +| Intelligence mapping | `intelligence_channel_mappings` | Channel-to-thread mapping. | Package-provided agents are public and ownerless. User-created coworkers are owned by the creator. @@ -28,8 +29,9 @@ This standing role applies in every channel. Treat channel messages as task-spec A further provenance block is appended by the deployment rather than the package: it tells the coworker to say where each answer came from, to mark plainly anything it answers from its own -knowledge rather than from a source, and never to present the latter as the former. Being deployment-wide, -it cannot be forgotten from the next coworker somebody adds. +knowledge rather than from a source, and never to present the latter as the former. A second block +tells it how to read the untrusted-content envelope its tool results carry. Being deployment-wide, +both cannot be forgotten from the next coworker somebody adds. The message is ordinary AG-UI system content, so it works with any AG-UI-compatible backend. Editing the role affects the next run. @@ -42,6 +44,11 @@ The message is ordinary AG-UI system content, so it works with any AG-UI-compati Filtering happens in server/database queries. Package-provided agents cannot be edited or deleted through the product. +An owner can also publish a Bot as a Team Bot, to everybody signed in or to named people and groups +(`/api/team-bots`, the **Team Bots** screen). Publishing is a separate, revocable record; the Bot's own +visibility is unchanged by it. A published Bot stays invisible to teammates until it has a role +description and a name other than a placeholder such as "New Bot". + ## Channels Starting a channel creates a new conversation and Intelligence thread. Two channels with the same coworker stay separate. @@ -52,7 +59,7 @@ Each channel routes through a channel-local proxy agent id, pinned to that chann Deleting is soft. The coworker stops running, but existing channels remain readable for their members and restore as tombstones. -Hiding is personal roster state. It removes the coworker from one user's list without disabling the coworker for anyone else. +Hiding and pinning are personal roster state. Hiding removes the coworker from one user's list without disabling the coworker for anyone else; pinning moves it into the **Pinned** section at the top of `/agents` for that user only. ## Default endpoint diff --git a/docs/development.md b/docs/development.md index ade36ded4..f6da20edd 100644 --- a/docs/development.md +++ b/docs/development.md @@ -54,7 +54,7 @@ Use `bun run dev` only when you want the app and API server without starting the | `supervisor` | 4500 host / 4300 container | | PostgreSQL | 5432 | -`start.sh` leaves existing matching services alone and reports when a port is held by another process. +`start.sh` hands every selected Docker service to `docker compose up -d` on every run, so one whose configuration changed is recreated and an unchanged one is left as it is. It leaves an API server or app that already answers as OpenBot alone, restarts the API server when it refuses the worker's secret or has no routines route, and reports when a port is held by another process. **Nothing here sweeps staged attachments.** A file dropped into the composer is stored before the message is sent, and the only thing that reclaims the ones never sent is diff --git a/docs/plugins/composio.md b/docs/plugins/composio.md index c50bf69e5..7ee1bba98 100644 --- a/docs/plugins/composio.md +++ b/docs/plugins/composio.md @@ -494,10 +494,11 @@ offers, and should be chosen rather than discovered. ## Not built yet -**No approval step before a destructive action.** The destructive marker is recorded and now -visible, but it gates nothing: a Bot granted a destructive action performs it without anybody being -asked. That is the same position every other connector is in — but it is now a position reachable -through the UI rather than only through a database insert, which is a real change in exposure. +**The destructive marker gates nothing by itself.** It is recorded and visible, but a Bot granted a +destructive action is not asked about it because of the marker. What can stop such a call is the +same as for every other connector's write: the person's **Ask before making changes** switch, an +approval rule on the **Approvals** page (a person's own, or a team rule), or a built-in safety +requirement. With none of those matching, the action runs. **No way to give a Bot a whole large app to search.** A Bot carries the actions somebody switched on for it, one at a time. There is no search-and-run path for an app too large to tick through, which diff --git a/docs/releasing.md b/docs/releasing.md index a7a0f79fd..5433cd6b7 100644 --- a/docs/releasing.md +++ b/docs/releasing.md @@ -116,9 +116,13 @@ A job added to `ci.yml` is covered by it without anybody updating a list. | check | what it would catch | | --- | --- | | `format, lint, types` | the ordinary things, across every workspace including `agent-computer` and the supervisor | +| `types (agent-computer)`, `types (supervisor)` | a type error in either deployable package, each checked on its own install | +| `native types` | a type error in the `mobile` app, which installs from its own npm lockfile | +| `computer (real browser)` | the agent-computer behaviour only a real Chromium shows, such as a password reaching the snapshot a model reads, a session cookie lost on restart, or WebRTC leaving around the egress filter | | `tests` | a decision made wrongly, in isolation | -| `chart` | a Helm values file that renders a server which cannot start, across the EKS, GKE, AKS and self-hosted targets | +| `chart` | a Helm values file that renders a server which cannot start, across the self-hosted, EKS, EKS sandbox, GKE and AKS targets | | `python harness regressions` | a provider-boundary regression in the Python Bot harnesses | +| `startup (macos-latest)`, `startup (windows-latest)` | the app's serve-or-build cache deciding wrongly on macOS or Windows | | `build` | the app not compiling | | `migrations` | a schema change with no migration, or a snapshot that has drifted | | `image` | an image that builds but does not boot, or a supervised service that respawns | @@ -133,7 +137,8 @@ These checks run again, against the release commit, when the release PR is merge publish rather than the proposal, which is why the release PR arriving without its own checks does not matter: a pull request opened by a workflow does not trigger them. -**No secrets are required.** Every workflow here uses only the built-in `GITHUB_TOKEN`. +**No secrets are required to release.** CI and both release workflows use only the built-in +`GITHUB_TOKEN`. The desktop build workflows also read `OPENBOT_GOOGLE_MODEL_OAUTH_CLIENT_SECRET`. ## The one thing CI cannot do diff --git a/docs/routines.md b/docs/routines.md index 000176178..d903eb8f5 100644 --- a/docs/routines.md +++ b/docs/routines.md @@ -37,7 +37,8 @@ deleting one is not required. A routine that fails posts exactly one message about it — the first failure after a success, not every failure. Ten consecutive failures switch the routine off and post a second, final message -saying so; nothing further fires until a person turns it back on. +saying so; nothing further fires until a person turns it back on. Turning it back on starts the +count again: only failures since the routine was last switched on count toward the ten. This is deliberately not a retry policy. A retry policy answers "did this one attempt make it through a dispatch that failed for a moment" — a busy queue, a server that hiccuped — and that question is @@ -107,9 +108,9 @@ separate, because as far as the channel is concerned, that is exactly what it is ## Scope This ships the core: creating, listing, changing and deleting routines from chat; the schedule, the -cap and the fatigue rule; the worker that fires them. Four follow-ups are tracked in -[#193](https://github.com/CopilotKit/OpenBot/issues/193) and deliberately not in this pass; the first -of them has since been closed. Audit rows now say what started the run they came out of, so a +cap and the fatigue rule; the worker that fires them. Four follow-ups were listed in +[#193](https://github.com/CopilotKit/OpenBot/issues/193), which is now closed; the first of them has +been built. Audit rows now say what started the run they came out of, so a routine's action is told apart from the same person's own by reading the row rather than by correlating timestamps against `routine_runs`. See [Architecture](architecture.md#what-started-a-run). Still open: there is no admin view of diff --git a/docs/windows-signing.md b/docs/windows-signing.md index 83a0b7cb8..795a54995 100644 --- a/docs/windows-signing.md +++ b/docs/windows-signing.md @@ -7,8 +7,7 @@ with the binaries. See [desktop build versions](releasing.md#desktop-build-versi The [Desktop Windows signing workflow](../.github/workflows/desktop-signing.yml) builds OpenBot and its NSIS installer with the existing DigiCert certificate in Azure Key Vault. It retains verified binaries and signature evidence as Actions -artifacts for 14 days. It does not create or publish a release. Desktop version -`0.0.0` remains a validation build. +artifacts for 14 days. It does not create or publish a release. Ordinary Desktop CI and fork PR builds remain unsigned. The `tauri.windows-signing.conf.json` overlay is passed explicitly to Tauri only by @@ -17,13 +16,13 @@ automatically merges that filename into every Windows build. ## Request a signed validation build -Before this workflow is merged to `main`, add the `windows-signing` label to a +To sign a pull request's build, add the `windows-signing` label to a same-repository PR and approve its `windows-signing` environment deployment. The workflow checks out the exact PR head SHA from the labeling event. New pushes run the regression job; remove and reapply the label to sign the new SHA. Fork PRs cannot enter this signing job. There is no `pull_request_target` trigger. -After merge, use **Actions → Desktop Windows signing → Run workflow**, select the +To sign any ref, use **Actions → Desktop Windows signing → Run workflow**, select the ref to validate, and set `signing-mode` to `keyvault`. The default `none` runs only credential-free regressions. Environment reviewers should check the exact source SHA and workflow changes before approving access to the publisher's key. From eac7920d1d18c4c606b6a9d0b6918503f3fbe96e Mon Sep 17 00:00:00 2001 From: David McKay Date: Fri, 2 Oct 2026 13:13:50 -0700 Subject: [PATCH 2/2] Say that a remote Bot hands work on, and drop the extraEnv note the chart replaced architecture doc no longer describes every such hop as refused. #717 gives the delivery, SCIM, inbound email and OTel settings their own chart values, which supersedes the extraEnv guidance. --- charts/openbot/README.md | 10 ---------- docs/architecture.md | 7 ++++--- 2 files changed, 4 insertions(+), 13 deletions(-) diff --git a/charts/openbot/README.md b/charts/openbot/README.md index a8b1ae41e..236b5786a 100644 --- a/charts/openbot/README.md +++ b/charts/openbot/README.md @@ -344,16 +344,6 @@ delivers anything. `config.learning.containerId` sets the initial default contai Automatic Learning take precedence, including switching it off. See [docs/automatic-learning.md](../../docs/automatic-learning.md). -## Settings with no value of their own - -Several server settings have no dedicated chart value and are passed through `config.extraEnv` or -`config.extraEnvFrom`: OpenTag delivery to Slack and Teams (`OPENTAG_URL`, `OPENTAG_SHARED_SECRET`), -text messages (`TWILIO_*`), mobile push (`EXPO_ACCESS_TOKEN`, `EXPO_PROJECT_ID`), inbound email for -Responsibilities (`OPENBOT_INBOUND_EMAIL_DOMAIN`, `OPENBOT_INBOUND_EMAIL_SNS_TOPIC_ARNS`), SCIM -(`SCIM_*`), OpenTelemetry export (`OPENBOT_OTEL_EXPORT`, `OTEL_*`) and `DELIVERY_PUBLIC_URL`. -`.env.example` describes each. Put the secret ones in a Secret and name it in `config.extraEnvFrom` -rather than writing them into a values file. - ## Slack, Teams, text messages, push, SCIM, inbound email and OpenTelemetry Each is off until its values are set, and each is the same setting [docs/configuration.md](../../docs/configuration.md) diff --git a/docs/architecture.md b/docs/architecture.md index f98b90e6f..6eccfad74 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -271,9 +271,10 @@ through the signed callback route is dispatched the same way as one from a Bot r handed work is not the same as being able to hand it on, so the target of a grant may live at its own endpoint. -One limit remains on the asking side. The grant check a hop goes through (`botsReachableFrom`) counts -only grants held by built-in Bots, so a Bot at its own endpoint can be granted another Bot and can ask -its person, but every hop it attempts is refused as not granted. +The asking side works the same way. The grant check a hop goes through (`botsReachableFrom`) counts a +grant whatever the asking Bot's type, so a Bot at its own endpoint hands work on through the signed +callback, the same grant check, caps and audit rows as a built-in one, to a built-in Bot or to another +Bot at its own endpoint. A coworker created through the UI with an endpoint is a remote one. Created without one, it runs at the managed Bot's endpoint when the deployment has one, and otherwise runs here as a built-in Bot on