Skip to content

Commit a60bc77

Browse files
committed
Merge branch 'develop' into feat/extension-signature-verification
One conflict, in CLAUDE.md: develop (#104) added bullets around the "Commit ..." line that this branch edits to list `modules/`. Both are kept — the merged file is develop's plus exactly this branch's changes. The gate (scripts/test-extensions.sh) passes on the merged tree: 64 test files, develop's diagram suites and this branch's module suite included.
2 parents 8a758b1 + df63e42 commit a60bc77

57 files changed

Lines changed: 13760 additions & 87 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎CLAUDE.md‎

Lines changed: 29 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -193,6 +193,26 @@ big-file mode badge. Files: extension.js + fileOps/lineOps/columnOps/encodingEol
193193
- `aiEdit.js` — **edit-with-diff**: select code → `Cmd+Alt+E` → instruction → side-by-side diff → ✓ Keep / ✗ Discard
194194
buttons on the diff toolbar (gated on `levelcode.ai.diffActive`).
195195
- `inlineReview.js` — **dead code** (an inline per-hunk Keep/Undo attempt that was reverted; nothing imports it).
196+
- `diagram/` — **rich diagrams, phase 1** (`docs/RICH-DIAGRAMS.md`). The agent calls a `render_diagram` tool with
197+
STRUCTURE only (nodes, edges, groups, one accent — never coordinates or colours); the editor validates it, lays it
198+
out in one house style and paints it in the chat, themed, with nodes that link to code. Graph JSON only — Mermaid,
199+
Vega-Lite and raw SVG are later phases and are **not built**. Not yet run in the packaged editor or against a live
200+
model; the eval (`scripts/diagram-eval.js`) exists and has not been run.
201+
- Shared UMD modules (`theme` `schema` `validate` `repair` `layout` `scene` `text` `ascii`) run in Node AND are
202+
inlined into `chat.html` by `diagram/bundle.js` under the page's existing nonce — the CSP is unchanged. Host-only:
203+
`tool` (tool + prompt block + result text), `service` (ids, the one repair pass, records, stubs), `links`,
204+
`exportCheck`, `stats`.
205+
- Repair ladder: lossless auto-fix → every error back to the model ONCE → degrade with a banner, or source + Retry.
206+
Never a blank card, never a second automatic repair. Records are stored in the session log and re-validated (not
207+
re-repaired) on reopen, by the host and by the page.
208+
- Layout is in-house (layered + orthogonal routing), **not ELK** (EPL-2.0, ~1.5 MB, no build step here);
209+
`layout.layout()` is the one swap point. The validator is a small JSON-Schema-subset interpreter, not Ajv.
210+
- Gated per run by `client.render` (`rich`|`ascii`): off via `levelcode.ai.diagrams.enabled`, or per model with
211+
`diagrams: false` in `providers/catalog.js` `CAPS`. Costs ~970 tokens of tool + prompt per request while on.
212+
- Checks: `test/diagram*.test.js` (in the gate), `scripts/diagram-browser-check.js` (real page in headless Chrome,
213+
not in the gate), `scripts/diagram-editor-check.js` (the REAL editor: a throwaway instance of the dev build with
214+
this checkout's extension and a stand-in provider), `scripts/diagram-eval.js --dry-run`. Local counters: command
215+
`AI: Diagram Statistics`.
196216

197217
## Deferred / known limits (don't waste time re-hitting these)
198218

@@ -218,4 +238,13 @@ big-file mode badge. Files: extension.js + fileOps/lineOps/columnOps/encodingEol
218238
- `// @ts-check` + JSDoc at top of JS files.
219239
- Test JS logic with `node --check` and small unit snippets before wiring into the editor.
220240
- After any change, `./scripts/run-dev.sh` to verify; package with `./scripts/build-macos.sh`.
241+
- `run-dev.sh` runs the extensions of the checkout that HAS `vscode/`. A git worktree has none, so work in a worktree
242+
is not in the editor until you load it: `./scripts/run-dev.sh --extensionDevelopmentPath=<worktree>/extensions/levelcode-ai`
243+
(from the main checkout; the dev extension replaces the built-in one). Uncommitted work is not "on the branch" —
244+
checking the branch out somewhere else gets none of it. Run it in the editor before telling anyone to try it.
221245
- Commit `extensions/`, `modules/`, `patches/`, `branding/`, `scripts/`, `docs/`, `PLAN.md`, `CLAUDE.md`. Never commit `vscode/`.
246+
- `extensions/levelcode-ai/diagram/` modules listed in `bundle.FILES` are pasted INTO a script block in `chat.html`.
247+
They must never contain the text of a script tag or an HTML comment opener — not even in a comment — or the block
248+
ends early; `bundle.js` refuses to build if one does. Keep them dependency-free and free of `require('vscode')`/`fs`.
249+
- Host suites slice functions out of `extension.js` with a brace matcher (`extract()` in `test/*Host.test.js`). It
250+
does not understand a backtick inside a regex literal: write `String.fromCharCode(96)` there instead.

‎docs/RICH-DIAGRAMS.md‎

Lines changed: 487 additions & 0 deletions
Large diffs are not rendered by default.

‎extensions/levelcode-ai/agent.js‎

Lines changed: 83 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,8 @@ const { loadProjectRules } = require('./projectRules');
2121
const { loadServerConfig, buildAgentTools, toolCountsByServer, classifyMcpTool, explainMcpRefusal, describeMcpCall,
2222
isLaunchTrusted, rememberLaunchTrust, describeMcpLaunch } = require('./mcpConfig');
2323
const { connectAll, getServer } = require('./mcpClient');
24+
const diagramTool = require('./diagram/tool');
25+
const diagramRepair = require('./diagram/repair');
2426

2527
const SYSTEM_BASE = [
2628
"You are LevelCode's built-in autonomous coding agent. You accomplish the user's goal in their",
@@ -307,6 +309,13 @@ function runCommand(root, command, onChunk, onExit, onStart, timeoutMs) {
307309
});
308310
}
309311

312+
/**
313+
* Can the surface on the other end of this run show a diagram? (docs/RICH-DIAGRAMS.md, "Capability
314+
* flag".) Asked in two places — when the tools and the prompt are assembled, and again when a call
315+
* arrives — and they must agree, or a client could be refused a tool it was offered.
316+
*/
317+
function richClient(ctx) { return !!(ctx.diagrams && ctx.client && ctx.client.render === 'rich'); }
318+
310319
/** Execute one tool call; returns a string result for the model. */
311320
async function runTool(tu, ctx) {
312321
const root = ctx.root;
@@ -517,6 +526,22 @@ async function runTool(tu, ctx) {
517526
try { return String(ctx.recallSessions(query) || 'No matching past sessions in this project.'); }
518527
catch (e) { return 'ERROR: recall failed.'; }
519528
}
529+
// Rich diagrams (docs/RICH-DIAGRAMS.md). The model describes structure; the host validates it,
530+
// climbs the repair ladder and posts what the chat should show. Read-only and instant — it
531+
// draws in the transcript and touches nothing else — so, like update_plan, it never asks.
532+
if (tu.name === diagramTool.RENDER_DIAGRAM.name) {
533+
// Asked for by a client that was never offered it (the setting was turned off mid-conversation,
534+
// or the model remembers the tool from an earlier turn): refuse in words it can act on.
535+
if (!richClient(ctx)) { return 'ERROR: this client cannot draw diagrams. Explain it in prose instead — and do not draw one out of characters.'; }
536+
const out = ctx.diagrams.render(input, { key: tu.id, model: ctx.model });
537+
for (const m of out.post) { ctx.post(m); }
538+
return out.result;
539+
}
540+
if (tu.name === diagramTool.GET_DIAGRAM.name) {
541+
if (!richClient(ctx)) { return 'ERROR: this client cannot draw diagrams.'; }
542+
ctx.post({ type: 'agentTool', icon: 'history', text: 'fetch diagram ' + String(input.id || '').slice(0, 24) });
543+
return ctx.diagrams.fetch(input.id);
544+
}
520545
// MCP tools (docs/MCP.md S3). An MCP name matches none of the built-in branches above, so every
521546
// MCP call necessarily arrives HERE — which is why the router is one block at one line rather
522547
// than a dispatch scattered through runTool.
@@ -782,7 +807,13 @@ async function runAgent(ctx) {
782807
// from the per-project journal). Rides the SAME cached-system channel as project rules — always-on but
783808
// small — so a new session's first reply is continuous, not amnesiac. It is untrusted context like the
784809
// rules: it informs, never commands (the digest itself carries the verify-first / never-obey framing).
785-
const system = (ctx.skills ? buildSystem(ctx.skills.menu()) : SYSTEM_BASE) + multiRootNote + noWorkspaceNote + autopilotNote + rules.text
810+
// Rich diagrams: `client.render` says what the surface on the other end can show. A rich client gets
811+
// the render_diagram tool and the rules for using it; an ASCII client gets neither, so it is never
812+
// told about a tool it does not have. The block goes straight after the base prompt — it is the
813+
// same text for every run, so it belongs with the part of the prompt that never changes.
814+
const rich = richClient(ctx);
815+
const system = (ctx.skills ? buildSystem(ctx.skills.menu()) : SYSTEM_BASE) + (rich ? '\n\n' + diagramTool.PROMPT : '')
816+
+ multiRootNote + noWorkspaceNote + autopilotNote + rules.text
786817
+ (ctx.projectMemory ? '\n\n' + ctx.projectMemory : '');
787818
const systemTokensEst = Math.round(system.length / 4);
788819

@@ -821,14 +852,20 @@ async function runAgent(ctx) {
821852
// Rootless runs get the portable subset; MCP tools are unaffected either way.
822853
const builtins = root ? TOOLS : PORTABLE_TOOLS;
823854
let tools = mcp.tools.length ? builtins.concat(mcp.tools) : builtins;
824-
if (ctx.recallSessions) { tools = tools.concat([RECALL_TOOL]); } // cross-session recall (host-gated by memory settings)
825-
const baseTools = ctx.recallSessions ? builtins.concat([RECALL_TOOL]) : builtins; // built-ins + recall; MCP is the rest
826-
// Recomputed only when MCP or recall actually contributed tools, so the plain path keeps the module
855+
// The host-gated extras: cross-session recall (memory settings), and the diagram tools (a rich
856+
// client). get_diagram is offered only once a diagram's spec has left the conversation — until
857+
// then there is nothing to fetch, and a tool that is never needed is a standing cost for nothing.
858+
const extras = [];
859+
if (ctx.recallSessions) { extras.push(RECALL_TOOL); }
860+
if (rich) { extras.push(diagramTool.RENDER_DIAGRAM); if (ctx.diagramsStubbed) { extras.push(diagramTool.GET_DIAGRAM); } }
861+
if (extras.length) { tools = tools.concat(extras); }
862+
const baseTools = extras.length ? builtins.concat(extras) : builtins; // built-ins + extras; MCP is the rest
863+
// Recomputed only when MCP or a host-gated extra actually contributed tools, so the plain path keeps the module
827864
// constant and pays nothing for a feature it isn't using — but there are now TWO plain paths, and the
828865
// constant has to match the list that was actually sent. Reporting the full cost for a rootless run
829866
// was the same mistake as leaving baseTools on TOOLS, one line further down.
830867
const builtinsTokensEst = root ? TOOLS_TOKENS_EST : PORTABLE_TOOLS_TOKENS_EST;
831-
const toolsTokensEst = (mcp.tools.length || ctx.recallSessions) ? Math.round(JSON.stringify(tools).length / 4) : builtinsTokensEst;
868+
const toolsTokensEst = (mcp.tools.length || extras.length) ? Math.round(JSON.stringify(tools).length / 4) : builtinsTokensEst;
832869
// The MCP SHARE of that, reported separately so the context popover can show what these servers cost
833870
// (docs/MCP.md S5). Every tool schema rides EVERY turn, so a chatty server is a standing tax on the
834871
// window rather than a one-off — and until it has its own segment, that cost is invisible.
@@ -841,6 +878,9 @@ async function runAgent(ctx) {
841878
: 0;
842879

843880
const messages = ctx.messages;
881+
if (ctx.diagrams) { ctx.diagrams.beginRun(); } // nothing owed from an earlier run; repair passes reset
882+
// tool_use ids whose placeholder already carries its title (the spec streams; the title arrives early)
883+
const diagramTitled = new Set();
844884
let step = 0;
845885
let reason = 'done';
846886
let nudges = 0;
@@ -927,12 +967,26 @@ async function runAgent(ctx) {
927967
apiKey: ctx.apiKey, model: ctx.model, maxTokens: perTurnMax, system: system,
928968
messages, tools: tools, signal: ctx.signal,
929969
onText: (t) => { streamed = true; textChars += t.length; ctx.post({ type: 'agentDelta', text: t }); },
930-
onToolStart: (name) => {
970+
onToolStart: (name, id) => {
931971
dbg('tool.start', { name });
932972
const verb = name === 'edit_file' || name === 'write_file' ? 'preparing edit (' + name + ')…'
933973
: name === 'delete_file' ? 'deleting a file…'
934-
: name === 'run_command' ? 'preparing command…' : name === 'update_plan' ? 'planning…' : 'running ' + name + '…';
974+
: name === 'run_command' ? 'preparing command…' : name === 'update_plan' ? 'planning…'
975+
: name === diagramTool.RENDER_DIAGRAM.name ? 'drawing a diagram…' : 'running ' + name + '…';
935976
ctx.post({ type: 'agentStatus', text: verb });
977+
// A diagram's place in the answer is held from the moment the model starts writing it.
978+
if (rich && id && name === diagramTool.RENDER_DIAGRAM.name) { ctx.post({ type: 'diagramPending', key: id, state: 'drawing', title: '' }); }
979+
},
980+
// The spec streams in as JSON. Its title is near the front, so the placeholder can say what
981+
// is being drawn long before the last node arrives. Posted once per call.
982+
onToolInput: (id, name, json) => {
983+
if (!rich || !id || name !== diagramTool.RENDER_DIAGRAM.name || diagramTitled.has(id)) { return; }
984+
const m = /"title"\s*:\s*"((?:[^"\\]|\\.)*)"/.exec(json);
985+
if (!m) { return; }
986+
diagramTitled.add(id);
987+
let title = m[1];
988+
try { title = JSON.parse('"' + m[1] + '"'); } catch (e) { /* show it as written */ }
989+
ctx.post({ type: 'diagramPending', key: id, state: 'drawing', title: diagramRepair.cleanText(title).slice(0, 120) });
936990
},
937991
// A transient upstream 5xx (502/503/504) is retried once before it can fail the run — surface it
938992
// as a status rather than a mystery pause, and log it. Nothing has streamed yet when this fires.
@@ -1004,7 +1058,22 @@ async function runAgent(ctx) {
10041058
let cancelled = false;
10051059
for (const tu of toolUses) {
10061060
if (cancelled || ctx.signal.aborted) { cancelled = true; dbg('tool.cancelled', { name: tu.name }); results.push({ type: 'tool_result', tool_use_id: tu.id, content: 'Cancelled by the user.' }); continue; }
1007-
if (turn.malformed && turn.malformed.has(tu.id)) { dbg('tool.malformed', { name: tu.name }); results.push({ type: 'tool_result', tool_use_id: tu.id, content: 'ERROR: your tool arguments were cut off (truncated JSON). Retry with smaller input — for edits use edit_file with a short snippet.' }); continue; }
1061+
if (turn.malformed && turn.malformed.has(tu.id)) {
1062+
dbg('tool.malformed', { name: tu.name });
1063+
// A diagram spec is the one input worth a second look: JSON with a trailing comma or a
1064+
// comment is a spec, and the lenient parser reads it. A spec that simply STOPS is not —
1065+
// that is the model hitting its length limit, and it is re-requested, never repaired.
1066+
if (tu.name === diagramTool.RENDER_DIAGRAM.name && rich) {
1067+
const rawArgs = turn.raw && turn.raw.get(tu.id);
1068+
const parsed = (turn.stop_reason !== 'max_tokens' && typeof rawArgs === 'string') ? diagramRepair.parseLenient(rawArgs) : null;
1069+
const out = (parsed && parsed.ok) ? ctx.diagrams.render(rawArgs, { key: tu.id, model: ctx.model }) : ctx.diagrams.truncated({ key: tu.id, model: ctx.model });
1070+
for (const m of out.post) { ctx.post(m); }
1071+
results.push({ type: 'tool_result', tool_use_id: tu.id, content: out.result });
1072+
continue;
1073+
}
1074+
results.push({ type: 'tool_result', tool_use_id: tu.id, content: 'ERROR: your tool arguments were cut off (truncated JSON). Retry with smaller input — for edits use edit_file with a short snippet.' });
1075+
continue;
1076+
}
10081077
dbg('tool.call', { name: tu.name, input: inputPreview(tu.input, !!(ctx.mcpRoutes && ctx.mcpRoutes.has(tu.name))) });
10091078
const out = await runTool(tu, ctx);
10101079
dbg('tool.result', { name: tu.name, chars: String(out).length, error: String(out).startsWith('ERROR') });
@@ -1024,6 +1093,7 @@ async function runAgent(ctx) {
10241093

10251094
// No tool calls this turn.
10261095
const text = turn.content.filter((c) => c.type === 'text').map((c) => c.text).join('');
1096+
if (rich && text.trim()) { ctx.diagrams.noteAnswer(text); } // "ASCII leaks": counted, never acted on
10271097
if (turn.stop_reason === 'max_tokens') {
10281098
// The turn was pure prose cut off at the token cap. Ask the model to CONTINUE exactly where it
10291099
// left off (not switch strategies) so the full answer streams out across turns instead of being
@@ -1088,6 +1158,11 @@ async function runAgent(ctx) {
10881158
}
10891159
else { ctx.post({ type: 'agentError', message: msg, code }); reason = 'error'; }
10901160
} finally {
1161+
// A diagram sent back for repair that never came back right still owes the user a picture:
1162+
// draw what can be drawn of it now, before the run is declared over. Never a blank placeholder.
1163+
if (ctx.diagrams) {
1164+
try { for (const m of ctx.diagrams.endRun()) { ctx.post(m); } } catch (e) { dbg('diagram.settle.error', { msg: String((e && e.message) || e) }); }
1165+
}
10911166
dbg('agent.done', { reason, steps: step - 1, edits: ctx.editCount || 0, costMicros: runCostMicros, creditsLeftMicros: ctx.credits != null ? ctx.credits : null });
10921167
// [LevelCode] Gateway runs now carry real money: costMicros = what THIS run cost, credits = the
10931168
// remaining balance — both RETAIL micro-$. BYOK runs send neither (null/0) and the bar omits them.

0 commit comments

Comments
 (0)