Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .apm/prompts/odd-config.prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,7 +73,8 @@ With no arguments, in this order:
the line "or `/odd-instrument-stack create a stack <name>` for a
backend not listed" - and, on a remote stack, "or `target
<environment>` to read the values persisted for another
environment". Anything the user
environment", naming the environments persisted for it (the
`<environment>-<stack>` keys of `stack_config`). Anything the user
picks goes to the `backend-configuration` skill's `## Switch`, which
owns the switch end to end: CLI presence preflight with a guided
install offer, the contract check for a custom stack, the persisted
Expand Down
7 changes: 3 additions & 4 deletions .apm/prompts/odd-verify.prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,10 +44,9 @@ this order:
baseline's pair, never silently retargets the current one, and
never rewrites the configuration: the divergence is stated, not
persisted; the entry the replay reads is the report's pair's, which
step 3's `## Check` resolves itself when the pair is not the
configured one (`effective` answers for the configured pair only),
and says so. That `stack` may be a custom name - a value on no row of
`builtin-stacks.md`: step 3's `## Check` resolves it from
the `baseline` command's `entry:` line names (`effective` answers
for the configured pair only). That `stack` may be a custom name -
a value on no row of `builtin-stacks.md`: step 3's `## Check` resolves it from
`.odd/observability-stacks/<name>/guide.md` in this clone and, when
that directory is absent, stops with the `observability-stack`
reference's error - carry that stop: the report's contract cannot
Expand Down
29 changes: 19 additions & 10 deletions .apm/skills/backend-configuration/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,11 +65,13 @@ connection.
python3 <this skill's directory>/scripts/preflight.py [--benchmark <dir>] [--containers <name>]
```

Which CLIs are installed and at which version, what is running, the
repository's branch and cleanliness, what each `.odd/` store holds, and
a named benchmark's target service and base URLs — every question a
Which CLIs are installed and at which version (the configured stack's
among them), what is running (or that the Docker daemon is
unreachable), the repository's branch and cleanliness, what each
`.odd/` store holds, and a named benchmark's target service and base
URLs — every question a
preflight asks the machine rather than the user. It takes no judgment,
so it takes no turns: one call, about a third of a second, instead of a
so it takes no turns: one call instead of a
shell command per question. Exit is always 0 — an absent CLI or a
missing directory is an answer the steps below act on, not a failure.

Expand Down Expand Up @@ -138,9 +140,13 @@ is persisted for the configured environment, the plain `<stack>`
otherwise — whole, never a merge of the two) and its `stack_config` is
the values; read them there, never composed from `stack_config` by
hand. For a replay's other pair (`odd-verify`, step 1) `effective`
does not apply: resolve the entry yourself from `stack_config` —
`<environment>-<stack>` when that key is present, else `<stack>` — and
say so in the display, next to the divergence. Show that
does not apply: the `baseline` command's `entry:` line names it — say
so in the display, next to the divergence.
An environment configured whose key is the plain `<stack>` (or none) is
a **named degradation**, said as such: "no `<environment>-<stack>`
entry - the runs read the plain entry (none: the CLI's active context),
whose environment is unverified", with the offer `persist ... for
<stack> in <environment>`. Show that
configuration to the user **as-is, no confirmation needed** — it is
informative: which instance, tenant, or site the queries are about to
hit is exactly what a user wants to see before a run, and what catches a
Expand Down Expand Up @@ -312,8 +318,9 @@ is a write of the `environment` field and nothing else:
`unknown` or `local`), or `{"environment": null}` to clear it. It selects which
of the configured stack's entries the missions read (`## Check` step
2's effective entry) and changes no `stack_config` value; the local
stack refuses it. A switch that names both ("switch to cloudwatch in
prod") writes both fields in the one call. Like every path it ends at
stack refuses it, so on local write none and tell the user so. A
switch that names both ("switch to cloudwatch in prod") writes both
fields in the one call. Like every path it ends at
verification (step 5).

The switch alone touches nothing else: it does not boot, reset, or stop
Expand Down Expand Up @@ -349,7 +356,9 @@ excepted, it takes no prefix), under the prefixed key:
`{"stack_config": {"<environment>-<stack>": {...}}}`. The key carries
the environment, the fields stay the stack's, and the entry is whole:
what the missions read for that pair is that entry alone, never the
plain one underneath it. The
plain one underneath it. Creating one, say so, and ask for the
reference's other `What to persist` fields or offer to copy the plain
entry's values into the same write. The
payload is merged into that entry and every other entry
is left untouched, so a one-value correction is a one-value call. Values
are flat scalars (string, number, boolean) and nothing else: identifiers,
Expand Down
71 changes: 62 additions & 9 deletions .apm/skills/backend-configuration/scripts/preflight.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,9 @@

It deliberately does NOT resolve the stack or read the backend's
configuration: those come from the MCP tools and the backend's own
reference, which is the part of the preflight that is not mechanical.
reference, which is the part of the preflight that is not mechanical. It
reads the configured stack's name only, to probe that stack's CLI too (its
row of the observability-cli-guides skill's builtin-stacks.md).

preflight.py
preflight.py --benchmark .odd/benchmarks/<name> --containers llmbench
Expand All @@ -37,6 +39,13 @@
"k6": ["k6", "version"],
"docker": ["docker", "--version"],
}
CONFIG = Path.home() / ".oddyssey" / "config.json"
BUILTIN_STACKS = (
Path(__file__).resolve().parents[2]
/ "observability-cli-guides"
/ "references"
/ "builtin-stacks.md"
)
ODD_STORES = (
".odd/observe-run-reports",
".odd/otel-instrumentation-reports",
Expand All @@ -60,23 +69,48 @@ def run(args: list[str], cwd: Path | None = None, timeout: int = 20) -> str:
return p.stdout.strip()


def stack_cli() -> str | None:
"""The configured stack's CLI - the first backticked name of its row's CLI
column in builtin-stacks.md - or None: no configuration, a custom stack,
or a CLI already probed."""
try:
stack = json.loads(CONFIG.read_text()).get("stack")
table = BUILTIN_STACKS.read_text()
except (OSError, ValueError, AttributeError):
return None
for line in table.splitlines():
cells = [c.strip() for c in line.split("|")]
if len(cells) > 4 and cells[1] == f"`{stack}`":
m = re.search(r"`([\w.-]+)`", cells[3])
if m and m.group(1) not in CLIS:
return m.group(1)
return None


def cli(name: str) -> dict:
"""Present or absent, and the version string when present.

A CLI that prints a JSON object (`gcx version` does) is reduced to its
`version` field: a raw `{...}` in the block reads as a corrupted tool
`version` field, else its first string value (`az version`): a raw `{...}` in the block reads as a corrupted tool
result to a run, which then re-runs the script and reads its source.
"""
if shutil.which(name) is None:
return {"present": False}
out = (run(CLIS[name]) or "").strip()
# a stack's CLI: `version` first (`az --version` checks for updates over
# the network, two seconds), `--version` when it refuses (`aws`)
out = (run(CLIS.get(name, [name, "version"])) or "").strip()
if not out and name not in CLIS:
out = (run([name, "--version"]) or "").strip()
if out.startswith("{"):
try:
parsed = json.loads(out)
except ValueError:
parsed = None
if isinstance(parsed, dict) and parsed.get("version"):
return {"present": True, "version": str(parsed["version"])}
if isinstance(parsed, dict): # `az version`: {"azure-cli": "2.89.1", ...}
first = next((v for v in parsed.values() if isinstance(v, str)), "")
return {"present": True, "version": first}
first = out.splitlines()
return {"present": True, "version": first[0].strip() if first else ""}

Expand All @@ -102,10 +136,15 @@ def machine_line(report: dict) -> str:
else f"{name} ABSENT"
for name, value in report["clis"].items()
)
running = ", ".join(c["name"] for c in report["containers"]) or "nothing"
if report["containers"] is None:
running = "docker daemon unreachable"
else:
running = "running " + (
", ".join(c["name"] for c in report["containers"]) or "nothing"
)
r = report["repo"]
state = "clean" if r["clean"] else f"dirty ({r['dirty_paths']} paths)"
parts = [clis, f"running {running}", f"repo {r['branch']} {r['head']} {state}"]
parts = [clis, running, f"repo {r['branch']} {r['head']} {state}"]
b = report.get("benchmark")
if b:
if not b["exists"]:
Expand All @@ -123,14 +162,24 @@ def report_json(report: dict) -> dict:
return {**report, "machine": machine_line(report)}


def containers(name_filter: str | None) -> list[dict]:
def containers(name_filter: str | None) -> list[dict] | None:
"""The running containers; None when `docker ps` fails - the daemon is
unreachable, which an empty list would misstate as nothing running."""
if shutil.which("docker") is None:
return []
args = ["docker", "ps", "--format", "{{.Names}}\t{{.Status}}\t{{.Image}}"]
if name_filter:
args += ["--filter", f"name={name_filter}"]
try:
p = subprocess.run(
args, capture_output=True, text=True, timeout=20, check=False
)
except (subprocess.TimeoutExpired, OSError):
return None
if p.returncode != 0:
return None
rows = []
for line in (run(args) or "").splitlines():
for line in p.stdout.splitlines():
parts = line.split("\t")
if len(parts) == 3:
rows.append({"name": parts[0], "status": parts[1], "image": parts[2]})
Expand Down Expand Up @@ -237,7 +286,9 @@ def render(report: dict) -> str:
lines.append(f" {path}")
if r.get("dirty_truncated"):
lines.append(f" ... and {r['dirty_truncated']} more")
if report["containers"]:
if report["containers"] is None:
lines.append(" containers daemon unreachable")
elif report["containers"]:
lines.append(" containers")
for c in report["containers"]:
lines.append(f" {c['name']} {c['status']}")
Expand Down Expand Up @@ -276,7 +327,9 @@ def main() -> int:

root = Path(args.root).resolve()
with ThreadPoolExecutor(max_workers=8) as pool:
f_clis = {name: pool.submit(cli, name) for name in CLIS}
extra = stack_cli()
names = [*CLIS, *([extra] if extra else [])]
f_clis = {name: pool.submit(cli, name) for name in names}
f_containers = pool.submit(containers, args.containers)
f_repo = pool.submit(repo, root)
f_stores = pool.submit(stores, root)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -537,9 +537,13 @@ A mission records profiles as a telemetry gap and moves on.
### Display

```bash
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-context.py check --app <app_insights_app> --workspace <workspace>
python3 <Skills>/observability-cli-guides/scripts/azure-monitor-context.py check --subscription <subscription> --resource-group <resource_group> --workspace <workspace> --app <app_insights_app>
```

Whole surface of `check`: the four flags above, each the resolved
entry's value, omitted when not persisted (the output names it as
skipped), and `--json`.

Two sources, and every line says which one it came from - the CLI
identity and the persisted targeting values are different facts and a
mismatch between them is what this display exists to catch.
Expand Down Expand Up @@ -567,9 +571,9 @@ otherwise - shown next to its field:
takes), not its resource name.

A field the user did not persist reads "not persisted - the mission will
ask"; an empty resolved entry (`{}`) is all four
unset, a valid state. `app_insights_app` unset is the one exception to
that neutral wording - a **named degradation**:
ask"; an entry with neither `workspace` nor `app_insights_app` is the
proof's `identity only` below. `app_insights_app` unset is the one
exception to that neutral wording - a **named degradation**:

> no Application Insights configured - `requests`/`dependencies`/
> `customMetrics`/`traces`/`exceptions` and the Profiler are unavailable,
Expand All @@ -592,10 +596,16 @@ alone is not a connected verdict when `app_insights_app` is persisted:
(never `-g` beside it, never `--subscription`: the data plane needs
neither), and with `--workspace` the same against the workspace; about
a second each. Skipped - not failed - when `--app` is not given.
- `--subscription` is proved by `az account show --subscription`, and
`--resource-group` by `az group exists` (`false` is `not-found`); under
a second each, verified 2026-09-26.

The exit code is the verdict, and the output carries the diagnosis:

- **0** - connected (both parts). Verified 2026-09-11.
- **3** with `NOT connected - identity only` - neither `--workspace` nor
`--app` given: nothing the queries read is proven; route to the
switch for the entry's values. Verified 2026-09-26.
- **3** - the persisted value does not resolve: an unknown appId
(`ApplicationNotFoundError`, az's exit 3) or a value that is not an
appId GUID (`The Application Insight is not found. Please check the app
Expand Down Expand Up @@ -633,6 +643,7 @@ stated above, and the mission proceeds logs-only having said so.
- "persist workspace <guid> for azure-monitor"
- "persist app insights <name-or-guid> for azure-monitor"
- "clear the workspace for azure-monitor"
- "persist workspace <guid> for azure-monitor in prod", "target prod"

## What to persist

Expand Down
24 changes: 12 additions & 12 deletions .apm/skills/observability-cli-guides/references/cloudwatch.md
Original file line number Diff line number Diff line change
Expand Up @@ -610,12 +610,12 @@ report's stack-friction section.
Two sources, labelled per line - the CLI's effective credentials and
the persisted targeting values.

**If the resolved entry persists `profile`, run every
command below (display and connection proof alike) with `--profile
<profile>`** - a bare call answers for whatever profile happens to
resolve without a flag, which on an SSO setup with no `default` is
routinely none at all, reporting a degradation on an account that is
configured and working.
**`profile` and `region` are required: run every command below
(display and connection proof alike) with `--profile <profile>`** - an
entry missing either is incomplete, the proof's exit 3 - and a bare
call answers for whatever profile happens to resolve without a flag,
which on an SSO setup with no `default` is routinely none at all, or
another account.

From the `aws` CLI:

Expand Down Expand Up @@ -655,9 +655,8 @@ otherwise - shown next to its key:
- `xray` - the X-Ray group name the service graph reads (`--xray-group`),
when persisted; the default group otherwise.

Every field the user did not persist is listed as "not persisted - the
mission will ask", and a present-but-empty resolved entry
(`{}`) means exactly that for all of them: a valid state, not an error.
Every other field the user did not persist is listed as "not persisted
- the mission will ask": a valid state, not an error.
Call out a persisted `region` that differs from the CLI's effective one
- the query targets the persisted value.

Expand All @@ -674,7 +673,8 @@ Never echo: an access key, a session token, the SSO cache.
python3 <Skills>/observability-cli-guides/scripts/cloudwatch-context.py check --profile <profile> --region <region> --log-group <log_group> --metrics-log-group <metrics_log_group>
```

Whole surface: `--profile`, `--region` (required), `--log-group`,
Whole surface: `--profile`, `--region` (required: either missing, exit
3 - the entry is incomplete, nothing is run), `--log-group`,
`--metrics-log-group` (each optional: given, proved to resolve;
omitted, skipped and said so), `--json`. Two parts: **identity** - `aws
sts get-caller-identity --profile <profile>` (needs no permission: a
Expand Down Expand Up @@ -726,8 +726,8 @@ information:
under (`--profile <name>` / `AWS_PROFILE`). SSO setups routinely have
**no `default` profile at all** - without this, `aws sts
get-caller-identity` fails with `NoCredentials` even though the CLI is
configured and working under its named profile. Skip the field only
when a `default` profile truly resolves on its own.
configured and working under its named profile. The proof requires
it: persist `default` when that profile truly resolves on its own.
- `log_group` - the CloudWatch Logs group the missions read for
**application logs**. When the services follow a convention rather
than one fixed group, store the **naming pattern** instead
Expand Down
Loading