From 4faf638253512c6e22ec0228fadc5766ed36537f Mon Sep 17 00:00:00 2001 From: rkoster Date: Thu, 17 Sep 2026 09:55:04 +0200 Subject: [PATCH 1/2] docs: add Agent Substrate research note --- ...026-09-17-agent-substrate-research-note.md | 46 +++++++ ...17-agent-substrate-research-note-design.md | 40 ++++++ research/agent-substrate.md | 122 ++++++++++++++++++ 3 files changed, 208 insertions(+) create mode 100644 docs/superpowers/plans/2026-09-17-agent-substrate-research-note.md create mode 100644 docs/superpowers/specs/2026-09-17-agent-substrate-research-note-design.md create mode 100644 research/agent-substrate.md diff --git a/docs/superpowers/plans/2026-09-17-agent-substrate-research-note.md b/docs/superpowers/plans/2026-09-17-agent-substrate-research-note.md new file mode 100644 index 0000000..47875a0 --- /dev/null +++ b/docs/superpowers/plans/2026-09-17-agent-substrate-research-note.md @@ -0,0 +1,46 @@ +# Agent Substrate Research Note Implementation Plan + +> **For agentic workers:** Execute this plan inline with validation checkpoints. + +**Goal:** Add and publish a concise architecture-first research note on Agent Substrate and its Cloud Foundry relevance. + +**Architecture:** Describe Substrate as an actor lifecycle and sandbox multiplexing layer: a control plane maps suspended actors onto ready workers, snapshots state, resumes actors on demand, and routes traffic. Separate documented demos from aspirational architecture, then map the model to CF process supervision, routing, isolation, and density. + +**Tech Stack:** Markdown, YAML frontmatter, Devbox, Git, GitHub CLI. + +--- + +### Task 1: Write the research note + +**Files:** +- Create: `research/agent-substrate.md` + +- [ ] Add frontmatter with title `Agent Substrate: Multiplexed Sandboxed Actors`, author `Ruben Koster (@rkoster)`, date `2026-09-17`, tags `[sandboxing, workload-isolation, orchestration, ecosystem-survey]`, `cf_areas: [diego, capi]`, `status: draft`, provisional ratings, and sources for the repository, README, architecture document, command tour, and counter demo. +- [ ] Explain the actor/worker model, where many mostly-idle actors are mapped onto fewer ready workers. +- [ ] Describe lifecycle operations: create/destroy, suspend/resume, worker assignment, routing, and full-state snapshots for process memory and filesystem state. +- [ ] Describe the component split: `ateapi` gRPC control plane, `atelet` node supervisor, `atecontroller` WorkerPool reconciler, `atenet` Envoy routing, and gVisor/microVM execution helpers. +- [ ] Distinguish the working counter demonstration from the architecture document's aspirational elements and note that the project is not an officially supported Google product. +- [ ] Assess Cloud Foundry relevance for Diego density, CAPI lifecycle APIs, route-to-resume behavior, sandbox isolation, snapshot storage, and observability. +- [ ] Add open questions about CF-native suspend/resume, trusted snapshot formats, network identity, tenant isolation, failure recovery, and whether actor multiplexing belongs in Diego or an adjacent substrate. + +### Task 2: Validate and inspect + +**Files:** +- Test: `.github/scripts/validate_notes.py` + +- [ ] Run `devbox run validate` and expect all research notes and ideas to be valid. +- [ ] Run `devbox run test` and expect success. +- [ ] Run `git diff --check` and inspect `git status --short`; leave unrelated environment artifacts unstaged. + +### Task 3: Commit and publish + +**Files:** +- Include: `research/agent-substrate.md` +- Include: `docs/superpowers/specs/2026-09-17-agent-substrate-research-note-design.md` +- Include: `docs/superpowers/plans/2026-09-17-agent-substrate-research-note.md` + +- [ ] Stage only the three intended files, using `git add -f` for ignored planning artifacts. +- [ ] Commit with `docs: add Agent Substrate research note`. +- [ ] Push `research/agent-substrate` to origin. +- [ ] Open a PR titled `docs: add Agent Substrate research note` targeting `main`, with the repository checklist completed. +- [ ] Verify the PR URL, branch, state, and CI status with `gh pr view`. diff --git a/docs/superpowers/specs/2026-09-17-agent-substrate-research-note-design.md b/docs/superpowers/specs/2026-09-17-agent-substrate-research-note-design.md new file mode 100644 index 0000000..1cbac0c --- /dev/null +++ b/docs/superpowers/specs/2026-09-17-agent-substrate-research-note-design.md @@ -0,0 +1,40 @@ +# Agent Substrate Research Note Design + +## Goal + +Capture a first-pass, sourced research note on Agent Substrate's architecture and potential +relevance to Cloud Foundry, with later refinement intentionally left open. + +## Scope + +The note will cover Agent Substrate's actor/worker model, lifecycle control, suspend/resume +and snapshotting, sandbox backends, network routing, Kubernetes integration, and the boundary +between implemented demonstrations and aspirational architecture. It will briefly assess +Cloud Foundry implications for Diego, CAPI, routing, workload isolation, density, and state +management. + +## Structure + +Create `research/agent-substrate.md` using the repository template and required sections: + +1. Summary +2. Key findings +3. CF relevance +4. Open questions + +Use concise provisional ratings and label this as a research snapshot. Do not present Agent +Substrate as an officially supported Google product or claim that aspirational architecture is +already implemented. + +## Sources and evidence + +Use the Agent Substrate repository README, architecture document, command/component tour, +observability or API documentation when available, and the counter demo. Claims about Cloud +Foundry will be analysis or open questions, not claims of existing integration. + +## Validation + +Run the repository's configured Devbox validation and test scripts, inspect whitespace and +the staged diff, then commit the note and this approved design spec on `research/agent-substrate`. +Push the branch and open a new PR targeting `main` without staging unrelated environment +artifacts. diff --git a/research/agent-substrate.md b/research/agent-substrate.md new file mode 100644 index 0000000..73e7889 --- /dev/null +++ b/research/agent-substrate.md @@ -0,0 +1,122 @@ +--- +title: "Agent Substrate: Multiplexed Sandboxed Actors" +author: Ruben Koster (@rkoster) +date: 2026-09-17 +tags: [sandboxing, workload-isolation, orchestration, ecosystem-survey] +cf_areas: [diego, capi] +status: draft +ratings: + platform-impact: + value: 86 + note: "The actor/worker model targets the density and isolation problems that become important when every agent needs its own stateful sandbox." + maturity: + value: 55 + note: "The repository has a working counter demonstration and substantial components, but its architecture document explicitly marks much of the design as aspirational." + novelty: + value: 84 + note: "Substrate combines sandbox checkpoint/restore, pre-started workers, actor teleportation, and traffic routing into an agent-oriented execution substrate." + actionability: + value: 70 + note: "The architecture provides concrete questions for CF around Diego density, snapshots, isolation, and route-triggered activation, although no CF integration is provided." +sources: + - https://github.com/agent-substrate/substrate + - https://raw.githubusercontent.com/agent-substrate/substrate/main/README.md + - https://raw.githubusercontent.com/agent-substrate/substrate/main/docs/architecture.md + - https://github.com/agent-substrate/substrate#tour + - https://raw.githubusercontent.com/agent-substrate/substrate/main/demos/counter/README.md +--- + +## Summary + +Agent Substrate is a Kubernetes-backed execution system for running large numbers of +stateful, mostly-idle actors in isolated sandboxes. It separates an actor's lifecycle from +the worker that currently hosts it: actors can be suspended, snapshotted, moved to a ready +worker, and resumed when traffic arrives. The project is infrastructure rather than an agent +SDK, and its architecture documentation explicitly identifies substantial parts of the design +as aspirational. + +## Key findings + +- **The central abstraction is actor versus worker.** An actor is an instance of an + agent-like workload, not necessarily an AI agent. Substrate maps many actors onto a smaller + pool of ready workers, relying on the fact that these workloads spend much of their time + waiting for input or events. +- **Lifecycle is independent of worker placement.** The control plane manages actor creation + and destruction, suspension and resumption, worker assignment, and traffic routing. A + suspended actor can be resumed on any suitable worker instead of remaining tied to the + process or node where it last ran. +- **Pre-started workers target both density and latency.** Rather than waiting for the + Kubernetes scheduler on every activation, Substrate keeps worker Pods ready and assigns an + actor to one when an event arrives. This allows oversubscription while aiming for sub-second + resume operations. +- **State can include volatile process memory.** The counter demo shows a `Full`-scope snapshot + preserving process memory together with filesystem state across suspend and resume. A + `Data`-scope policy preserves durable-volume data but restarts in-memory state, making the + snapshot policy an explicit application and platform trade-off. +- **The implementation is split into focused components.** `ateapi` exposes gRPC lifecycle + operations for actors and workers; `atelet` supervises worker Pods and coordinates snapshots + and state transfers; `atecontroller` reconciles WorkerPool custom resources; and `atenet` + provides Envoy routing and proxy sidecars. Separate helpers integrate gVisor checkpoint and + restore or run actors inside cloud-hypervisor microVMs. +- **Kubernetes is the infrastructure substrate.** Substrate uses Kubernetes for provisioning, + worker lifecycle management, Pod autoscaling, and controller integration, while adding + agent-specific scheduling and lifecycle control above those primitives. The counter demo + uses a WorkerPool CRD and an Agent Substrate actor template managed through an atespace. +- **Security is part of the execution model.** The project targets untrusted workloads and + supports sandbox technologies such as gVisor and microVMs. The README describes zero-trust + kernel and network isolation as design goals, but the exact security boundary and operational + guarantees require further investigation. +- **The evidence has different maturity levels.** The repository includes a runnable counter + demo that preserves state across suspend/resume. In contrast, the architecture document + states that much of its described architecture is aspirational, so performance and scale + claims should be treated as project targets or demonstrations until backed by reproducible + benchmark evidence. +- **This is not an agent programming framework.** Substrate supplies lifecycle, placement, + isolation, state preservation, and routing infrastructure. Agent logic, model access, tools, + and application-level durable workflows remain the responsibility of the workload or other + platform components. +- **The project has an explicit support caveat.** The README states that Agent Substrate is + not an officially supported Google product. That governance and support status matters when + evaluating it as a platform dependency. + +## CF relevance + +Agent Substrate is a useful comparison point for Cloud Foundry because it attacks a problem +that ordinary application scaling does not fully solve: many isolated workloads are idle most +of the time, but resuming them quickly requires preserving state and avoiding a fresh +scheduler placement on every request. Diego already owns process placement and health +management, while CAPI owns application lifecycle and metadata. A Substrate-like layer could +sit beside those systems or extend them with actor identity, suspended state, worker pools, and +route-triggered activation. + +The mapping is not direct. CF applications are normally long-running processes with routing +to currently running instances; Substrate assumes actors can disappear from workers while +retaining execution state. Implementing that model in CF would require trusted checkpoint and +restore support, durable snapshot storage, stable network identity, route lookup during resume, +and clear semantics for in-flight requests. It would also need an isolation boundary comparable +to the selected gVisor or microVM backend rather than relying only on ordinary application +process isolation. + +The most relevant research question is whether actor multiplexing belongs inside Diego or in an +adjacent substrate that Diego manages as a specialized workload. Either option would need +Loggregator-compatible lifecycle and routing events, resource accounting per suspended actor, +snapshot failure visibility, and tenant-aware network policy. The counter demo's distinction +between process-memory snapshots and durable-volume snapshots is especially relevant to CF's +stateful-agent discussions: preserving memory can improve resume fidelity, but increases the +security, storage, compatibility, and failure-recovery burden. + +## Open questions + +- Could Diego support suspended actor identities and fast resume, or should CF expose a + separate actor substrate alongside ordinary applications? +- What checkpoint format, storage backend, encryption, and key lifecycle would be trusted for + process memory containing credentials or sensitive agent context? +- How should CF route a request to a suspended actor, and what happens to requests that arrive + while snapshot restore or worker assignment is failing? +- Which isolation guarantees are required for untrusted agent code, and can they be provided by + the existing Diego execution model or only by gVisor/microVM-style workers? +- How should suspended actors consume quota, memory, storage, network identity, and billing + resources while they are not attached to a worker? +- What lifecycle, snapshot, resume, and routing events should Loggregator expose to operators? +- Which parts of Agent Substrate are validated by the current demo and benchmarks, and which + remain aspirational before it is suitable as a production platform dependency? From a2446eda8ae1007c803ae77ccfba0c68d95cc7aa Mon Sep 17 00:00:00 2001 From: rkoster Date: Sat, 19 Sep 2026 16:29:13 +0200 Subject: [PATCH 2/2] docs: update generated research map --- generated/research-map.html | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/generated/research-map.html b/generated/research-map.html index bbe5f21..7f8ecf0 100644 --- a/generated/research-map.html +++ b/generated/research-map.html @@ -23,9 +23,9 @@

Focus use cases

Attested Workload Authority and Mediated Tool AccessExchange platform-attested workload identity for scoped authority while credentials and outbound tool access remain mediated by the platform.Strategic decision: Decide whether CF should become the portable trust and policy layer between agent workloads and the tools they invoke.
Gap, experiments, and evidence
Current CF gap
CF issues workload identity certificates but does not exchange them for scoped tool authority, keep third-party credentials out of workloads, mediate off-platform access, or record delegation-aware audit events.
Candidate POC
Exchange a Diego instance identity certificate for a short-lived scoped token, invoke one allowed tool through a credential proxy and egress mediator, deny another, and emit attributable audit events.
Candidate RFC scope
Define workload token exchange, authority and delegation claims, credential brokering, outbound mediation and policy enforcement, audit events, revocation, and integration boundaries for UAA, routing, and service brokers.
-
Gap, experiments, and evidence
Current CF gap
CF can stage apps and run ephemeral tasks but cannot cheaply compose a reusable environment with per-session workspace state, select stronger isolation, constrain session networking, or resume the session lifecycle.
Candidate POC
Start two isolated sessions from one content-addressed staged environment, attach separate mutable workspaces, apply per-session egress policy, stop one session, and resume it on fresh compute.
Candidate RFC scope
Define environment and workspace references, session identity and lifecycle, isolation classes, network policy, workspace persistence and cleanup, scheduling, quotas, and compatibility with existing CF staging and task APIs.

ResearchIdea

Platform Impact x Maturity

Emerging < Maturity > EstablishedLocal concern < Platform Impact > Platform-wide concern
Unplaced notes (0)
  • All notes are placed.
+
Gap, experiments, and evidence
Current CF gap
CF can stage apps and run ephemeral tasks but cannot cheaply compose a reusable environment with per-session workspace state, select stronger isolation, constrain session networking, or resume the session lifecycle.
Candidate POC
Start two isolated sessions from one content-addressed staged environment, attach separate mutable workspaces, apply per-session egress policy, stop one session, and resume it on fresh compute.
Candidate RFC scope
Define environment and workspace references, session identity and lifecycle, isolation classes, network policy, workspace persistence and cleanup, scheduling, quotas, and compatibility with existing CF staging and task APIs.

ResearchIdea

Platform Impact x Maturity

Emerging < Maturity > EstablishedLocal concern < Platform Impact > Platform-wide concern
Unplaced notes (0)
  • All notes are placed.
-