diff --git a/.apm/prompts/odd-config.prompt.md b/.apm/prompts/odd-config.prompt.md index 5696a2b..cfa6dbb 100644 --- a/.apm/prompts/odd-config.prompt.md +++ b/.apm/prompts/odd-config.prompt.md @@ -73,7 +73,8 @@ With no arguments, in this order: the line "or `/odd-instrument-stack create a stack ` for a backend not listed" - and, on a remote stack, "or `target ` to read the values persisted for another - environment". Anything the user + environment", naming the environments persisted for it (the + `-` keys of `stack_config`). Anything the user picks goes to the `backend-configuration` skill's `## Switch`, which owns the switch end to end: CLI presence preflight with a guided install offer, the contract check for a custom stack, the persisted diff --git a/.apm/prompts/odd-verify.prompt.md b/.apm/prompts/odd-verify.prompt.md index a0198e3..40a3ca1 100644 --- a/.apm/prompts/odd-verify.prompt.md +++ b/.apm/prompts/odd-verify.prompt.md @@ -44,10 +44,9 @@ this order: baseline's pair, never silently retargets the current one, and never rewrites the configuration: the divergence is stated, not persisted; the entry the replay reads is the report's pair's, which - step 3's `## Check` resolves itself when the pair is not the - configured one (`effective` answers for the configured pair only), - and says so. That `stack` may be a custom name - a value on no row of - `builtin-stacks.md`: step 3's `## Check` resolves it from + the `baseline` command's `entry:` line names (`effective` answers + for the configured pair only). That `stack` may be a custom name - + a value on no row of `builtin-stacks.md`: step 3's `## Check` resolves it from `.odd/observability-stacks//guide.md` in this clone and, when that directory is absent, stops with the `observability-stack` reference's error - carry that stop: the report's contract cannot diff --git a/.apm/skills/backend-configuration/SKILL.md b/.apm/skills/backend-configuration/SKILL.md index f89eeb7..89bac06 100644 --- a/.apm/skills/backend-configuration/SKILL.md +++ b/.apm/skills/backend-configuration/SKILL.md @@ -65,11 +65,13 @@ connection. python3 /scripts/preflight.py [--benchmark ] [--containers ] ``` -Which CLIs are installed and at which version, what is running, the -repository's branch and cleanliness, what each `.odd/` store holds, and -a named benchmark's target service and base URLs — every question a +Which CLIs are installed and at which version (the configured stack's +among them), what is running (or that the Docker daemon is +unreachable), the repository's branch and cleanliness, what each +`.odd/` store holds, and a named benchmark's target service and base +URLs — every question a preflight asks the machine rather than the user. It takes no judgment, -so it takes no turns: one call, about a third of a second, instead of a +so it takes no turns: one call instead of a shell command per question. Exit is always 0 — an absent CLI or a missing directory is an answer the steps below act on, not a failure. @@ -138,9 +140,13 @@ is persisted for the configured environment, the plain `` otherwise — whole, never a merge of the two) and its `stack_config` is the values; read them there, never composed from `stack_config` by hand. For a replay's other pair (`odd-verify`, step 1) `effective` -does not apply: resolve the entry yourself from `stack_config` — -`-` when that key is present, else `` — and -say so in the display, next to the divergence. Show that +does not apply: the `baseline` command's `entry:` line names it — say +so in the display, next to the divergence. +An environment configured whose key is the plain `` (or none) is +a **named degradation**, said as such: "no `-` +entry - the runs read the plain entry (none: the CLI's active context), +whose environment is unverified", with the offer `persist ... for + in `. Show that configuration to the user **as-is, no confirmation needed** — it is informative: which instance, tenant, or site the queries are about to hit is exactly what a user wants to see before a run, and what catches a @@ -312,8 +318,9 @@ is a write of the `environment` field and nothing else: `unknown` or `local`), or `{"environment": null}` to clear it. It selects which of the configured stack's entries the missions read (`## Check` step 2's effective entry) and changes no `stack_config` value; the local -stack refuses it. A switch that names both ("switch to cloudwatch in -prod") writes both fields in the one call. Like every path it ends at +stack refuses it, so on local write none and tell the user so. A +switch that names both ("switch to cloudwatch in prod") writes both +fields in the one call. Like every path it ends at verification (step 5). The switch alone touches nothing else: it does not boot, reset, or stop @@ -349,7 +356,9 @@ excepted, it takes no prefix), under the prefixed key: `{"stack_config": {"-": {...}}}`. The key carries the environment, the fields stay the stack's, and the entry is whole: what the missions read for that pair is that entry alone, never the -plain one underneath it. The +plain one underneath it. Creating one, say so, and ask for the +reference's other `What to persist` fields or offer to copy the plain +entry's values into the same write. The payload is merged into that entry and every other entry is left untouched, so a one-value correction is a one-value call. Values are flat scalars (string, number, boolean) and nothing else: identifiers, diff --git a/.apm/skills/backend-configuration/scripts/preflight.py b/.apm/skills/backend-configuration/scripts/preflight.py index 935c049..77d6468 100755 --- a/.apm/skills/backend-configuration/scripts/preflight.py +++ b/.apm/skills/backend-configuration/scripts/preflight.py @@ -10,7 +10,9 @@ It deliberately does NOT resolve the stack or read the backend's configuration: those come from the MCP tools and the backend's own -reference, which is the part of the preflight that is not mechanical. +reference, which is the part of the preflight that is not mechanical. It +reads the configured stack's name only, to probe that stack's CLI too (its +row of the observability-cli-guides skill's builtin-stacks.md). preflight.py preflight.py --benchmark .odd/benchmarks/ --containers llmbench @@ -37,6 +39,13 @@ "k6": ["k6", "version"], "docker": ["docker", "--version"], } +CONFIG = Path.home() / ".oddyssey" / "config.json" +BUILTIN_STACKS = ( + Path(__file__).resolve().parents[2] + / "observability-cli-guides" + / "references" + / "builtin-stacks.md" +) ODD_STORES = ( ".odd/observe-run-reports", ".odd/otel-instrumentation-reports", @@ -60,16 +69,38 @@ def run(args: list[str], cwd: Path | None = None, timeout: int = 20) -> str: return p.stdout.strip() +def stack_cli() -> str | None: + """The configured stack's CLI - the first backticked name of its row's CLI + column in builtin-stacks.md - or None: no configuration, a custom stack, + or a CLI already probed.""" + try: + stack = json.loads(CONFIG.read_text()).get("stack") + table = BUILTIN_STACKS.read_text() + except (OSError, ValueError, AttributeError): + return None + for line in table.splitlines(): + cells = [c.strip() for c in line.split("|")] + if len(cells) > 4 and cells[1] == f"`{stack}`": + m = re.search(r"`([\w.-]+)`", cells[3]) + if m and m.group(1) not in CLIS: + return m.group(1) + return None + + def cli(name: str) -> dict: """Present or absent, and the version string when present. A CLI that prints a JSON object (`gcx version` does) is reduced to its - `version` field: a raw `{...}` in the block reads as a corrupted tool + `version` field, else its first string value (`az version`): a raw `{...}` in the block reads as a corrupted tool result to a run, which then re-runs the script and reads its source. """ if shutil.which(name) is None: return {"present": False} - out = (run(CLIS[name]) or "").strip() + # a stack's CLI: `version` first (`az --version` checks for updates over + # the network, two seconds), `--version` when it refuses (`aws`) + out = (run(CLIS.get(name, [name, "version"])) or "").strip() + if not out and name not in CLIS: + out = (run([name, "--version"]) or "").strip() if out.startswith("{"): try: parsed = json.loads(out) @@ -77,6 +108,9 @@ def cli(name: str) -> dict: parsed = None if isinstance(parsed, dict) and parsed.get("version"): return {"present": True, "version": str(parsed["version"])} + if isinstance(parsed, dict): # `az version`: {"azure-cli": "2.89.1", ...} + first = next((v for v in parsed.values() if isinstance(v, str)), "") + return {"present": True, "version": first} first = out.splitlines() return {"present": True, "version": first[0].strip() if first else ""} @@ -102,10 +136,15 @@ def machine_line(report: dict) -> str: else f"{name} ABSENT" for name, value in report["clis"].items() ) - running = ", ".join(c["name"] for c in report["containers"]) or "nothing" + if report["containers"] is None: + running = "docker daemon unreachable" + else: + running = "running " + ( + ", ".join(c["name"] for c in report["containers"]) or "nothing" + ) r = report["repo"] state = "clean" if r["clean"] else f"dirty ({r['dirty_paths']} paths)" - parts = [clis, f"running {running}", f"repo {r['branch']} {r['head']} {state}"] + parts = [clis, running, f"repo {r['branch']} {r['head']} {state}"] b = report.get("benchmark") if b: if not b["exists"]: @@ -123,14 +162,24 @@ def report_json(report: dict) -> dict: return {**report, "machine": machine_line(report)} -def containers(name_filter: str | None) -> list[dict]: +def containers(name_filter: str | None) -> list[dict] | None: + """The running containers; None when `docker ps` fails - the daemon is + unreachable, which an empty list would misstate as nothing running.""" if shutil.which("docker") is None: return [] args = ["docker", "ps", "--format", "{{.Names}}\t{{.Status}}\t{{.Image}}"] if name_filter: args += ["--filter", f"name={name_filter}"] + try: + p = subprocess.run( + args, capture_output=True, text=True, timeout=20, check=False + ) + except (subprocess.TimeoutExpired, OSError): + return None + if p.returncode != 0: + return None rows = [] - for line in (run(args) or "").splitlines(): + for line in p.stdout.splitlines(): parts = line.split("\t") if len(parts) == 3: rows.append({"name": parts[0], "status": parts[1], "image": parts[2]}) @@ -237,7 +286,9 @@ def render(report: dict) -> str: lines.append(f" {path}") if r.get("dirty_truncated"): lines.append(f" ... and {r['dirty_truncated']} more") - if report["containers"]: + if report["containers"] is None: + lines.append(" containers daemon unreachable") + elif report["containers"]: lines.append(" containers") for c in report["containers"]: lines.append(f" {c['name']} {c['status']}") @@ -276,7 +327,9 @@ def main() -> int: root = Path(args.root).resolve() with ThreadPoolExecutor(max_workers=8) as pool: - f_clis = {name: pool.submit(cli, name) for name in CLIS} + extra = stack_cli() + names = [*CLIS, *([extra] if extra else [])] + f_clis = {name: pool.submit(cli, name) for name in names} f_containers = pool.submit(containers, args.containers) f_repo = pool.submit(repo, root) f_stores = pool.submit(stores, root) diff --git a/.apm/skills/observability-cli-guides/references/azure-monitor.md b/.apm/skills/observability-cli-guides/references/azure-monitor.md index 1f72a54..8fbee16 100644 --- a/.apm/skills/observability-cli-guides/references/azure-monitor.md +++ b/.apm/skills/observability-cli-guides/references/azure-monitor.md @@ -537,9 +537,13 @@ A mission records profiles as a telemetry gap and moves on. ### Display ```bash -python3 /observability-cli-guides/scripts/azure-monitor-context.py check --app --workspace +python3 /observability-cli-guides/scripts/azure-monitor-context.py check --subscription --resource-group --workspace --app ``` +Whole surface of `check`: the four flags above, each the resolved +entry's value, omitted when not persisted (the output names it as +skipped), and `--json`. + Two sources, and every line says which one it came from - the CLI identity and the persisted targeting values are different facts and a mismatch between them is what this display exists to catch. @@ -567,9 +571,9 @@ otherwise - shown next to its field: takes), not its resource name. A field the user did not persist reads "not persisted - the mission will -ask"; an empty resolved entry (`{}`) is all four -unset, a valid state. `app_insights_app` unset is the one exception to -that neutral wording - a **named degradation**: +ask"; an entry with neither `workspace` nor `app_insights_app` is the +proof's `identity only` below. `app_insights_app` unset is the one +exception to that neutral wording - a **named degradation**: > no Application Insights configured - `requests`/`dependencies`/ > `customMetrics`/`traces`/`exceptions` and the Profiler are unavailable, @@ -592,10 +596,16 @@ alone is not a connected verdict when `app_insights_app` is persisted: (never `-g` beside it, never `--subscription`: the data plane needs neither), and with `--workspace` the same against the workspace; about a second each. Skipped - not failed - when `--app` is not given. +- `--subscription` is proved by `az account show --subscription`, and + `--resource-group` by `az group exists` (`false` is `not-found`); under + a second each, verified 2026-09-26. The exit code is the verdict, and the output carries the diagnosis: - **0** - connected (both parts). Verified 2026-09-11. +- **3** with `NOT connected - identity only` - neither `--workspace` nor + `--app` given: nothing the queries read is proven; route to the + switch for the entry's values. Verified 2026-09-26. - **3** - the persisted value does not resolve: an unknown appId (`ApplicationNotFoundError`, az's exit 3) or a value that is not an appId GUID (`The Application Insight is not found. Please check the app @@ -633,6 +643,7 @@ stated above, and the mission proceeds logs-only having said so. - "persist workspace for azure-monitor" - "persist app insights for azure-monitor" - "clear the workspace for azure-monitor" +- "persist workspace for azure-monitor in prod", "target prod" ## What to persist diff --git a/.apm/skills/observability-cli-guides/references/cloudwatch.md b/.apm/skills/observability-cli-guides/references/cloudwatch.md index 1070966..8bac488 100644 --- a/.apm/skills/observability-cli-guides/references/cloudwatch.md +++ b/.apm/skills/observability-cli-guides/references/cloudwatch.md @@ -610,12 +610,12 @@ report's stack-friction section. Two sources, labelled per line - the CLI's effective credentials and the persisted targeting values. -**If the resolved entry persists `profile`, run every -command below (display and connection proof alike) with `--profile -`** - a bare call answers for whatever profile happens to -resolve without a flag, which on an SSO setup with no `default` is -routinely none at all, reporting a degradation on an account that is -configured and working. +**`profile` and `region` are required: run every command below +(display and connection proof alike) with `--profile `** - an +entry missing either is incomplete, the proof's exit 3 - and a bare +call answers for whatever profile happens to resolve without a flag, +which on an SSO setup with no `default` is routinely none at all, or +another account. From the `aws` CLI: @@ -655,9 +655,8 @@ otherwise - shown next to its key: - `xray` - the X-Ray group name the service graph reads (`--xray-group`), when persisted; the default group otherwise. -Every field the user did not persist is listed as "not persisted - the -mission will ask", and a present-but-empty resolved entry -(`{}`) means exactly that for all of them: a valid state, not an error. +Every other field the user did not persist is listed as "not persisted +- the mission will ask": a valid state, not an error. Call out a persisted `region` that differs from the CLI's effective one - the query targets the persisted value. @@ -674,7 +673,8 @@ Never echo: an access key, a session token, the SSO cache. python3 /observability-cli-guides/scripts/cloudwatch-context.py check --profile --region --log-group --metrics-log-group ``` -Whole surface: `--profile`, `--region` (required), `--log-group`, +Whole surface: `--profile`, `--region` (required: either missing, exit +3 - the entry is incomplete, nothing is run), `--log-group`, `--metrics-log-group` (each optional: given, proved to resolve; omitted, skipped and said so), `--json`. Two parts: **identity** - `aws sts get-caller-identity --profile ` (needs no permission: a @@ -726,8 +726,8 @@ information: under (`--profile ` / `AWS_PROFILE`). SSO setups routinely have **no `default` profile at all** - without this, `aws sts get-caller-identity` fails with `NoCredentials` even though the CLI is - configured and working under its named profile. Skip the field only - when a `default` profile truly resolves on its own. + configured and working under its named profile. The proof requires + it: persist `default` when that profile truly resolves on its own. - `log_group` - the CloudWatch Logs group the missions read for **application logs**. When the services follow a convention rather than one fixed group, store the **naming pattern** instead diff --git a/.apm/skills/observability-cli-guides/scripts/azure-monitor-context.py b/.apm/skills/observability-cli-guides/scripts/azure-monitor-context.py index ca70cf6..f5a8708 100755 --- a/.apm/skills/observability-cli-guides/scripts/azure-monitor-context.py +++ b/.apm/skills/observability-cli-guides/scripts/azure-monitor-context.py @@ -1,28 +1,32 @@ #!/usr/bin/env python3 """The connection proof in two parts, and the bounded landing poll of a driven run. - azure-monitor-context.py check --app [--workspace ] + azure-monitor-context.py check --subscription --resource-group --workspace --app azure-monitor-context.py landing --app --identity --expect 110 --from ... --to ... -Whole surface - check: --app (the appId GUID; omitted, the targeting part is -skipped and the output says the run is logs-only), --workspace (the customer -ID GUID; given, it is proved the same way), --json. landing: --app, --identity +Whole surface - check: --subscription, --resource-group, --workspace (the +customer ID GUID), --app (the appId GUID) - each the resolved entry's value, +each proved when given and named as skipped when not (no --app: the run is +logs-only) - and --json. landing: --app, --identity (the run's user agent, matched on customDimensions['user_agent.original']), --expect N (the request count the poll waits for), a window (--from/--to or --since), --dimension (the customDimensions key the identity is matched on, default user_agent.original), --service (repeatable, scopes the count to cloud_RoleName), --every (seconds between polls, default 20), --cap (the -bound, default 3m), --json. Exit codes - check: 0 connected (both parts), -1 identity failure or a rights/network error (the message says what is -yours to do), 3 the persisted value does not resolve (a wrong value: route -to the switch), 2 az could not parse the command. landing: 0 landed, 1 the +bound, default 3m), --json. Exit codes - check: 0 connected (identity and +every part given, a workspace or a component among them), 1 identity +failure or a rights/network error (the message says what is yours to do), +3 a persisted value does not resolve, or neither a workspace nor a +component is given (identity alone proves nothing the queries read): route +to the switch, 2 az could not parse the command. landing: 0 landed, 1 the cap was reached (the last count is in the output), 3/2 as check. Identity is `az account show` (the local profile, no network: a stale token passes here and fails on the targeting part, which is why both run). -Targeting is `print 1` against the component with the appId alone - never --g beside it, never --subscription, the data plane needs neither - and, -with --workspace, against the workspace. The count of the landing poll is +Targeting is `az account show --subscription`, `az group exists`, and +`print 1` against the workspace and the component with the appId alone - +never -g beside it, never --subscription, the data plane needs neither - +run together. The count of the landing poll is read in json at tables[0].rows[0][0]: `-o tsv` would print 1 whatever the value (the number of result rows). """ @@ -51,6 +55,7 @@ render_commands, resolve_window, run_az, + run_many, ) DIAGNOSIS = { @@ -69,6 +74,12 @@ def _last_minutes(n: int) -> tuple[str, str]: return iso(now - timedelta(minutes=n)), iso(now) +def _targeting_code(kind: str) -> int: + if kind in ("not-found", "not-an-appid", "not-a-workspace-id"): + return 3 + return 2 if kind == "usage" else 1 + + def cmd_check(ns) -> tuple[int, dict]: out: dict = {"identity": {}, "targeting": {}, "connected": False, "commands": []} results = [] @@ -92,56 +103,57 @@ def cmd_check(ns) -> tuple[int, dict]: "user_type": (d.get("user") or {}).get("type"), "state": d.get("state"), } + # One bounded call per persisted part, run together; a part not given is + # named as skipped (#657), never left out of the output. + frm, to = _last_minutes(5) + sub = ["--subscription", ns.subscription] if ns.subscription else [] + parts = [ + ("subscription", ns.subscription, ["account", "show", *sub]), + ( + "resource_group", + ns.resource_group, + ["group", "exists", "--resource-group", ns.resource_group or "", *sub], + ), + ("workspace", ns.workspace, la_call(ns.workspace or "", "print 1", frm, to)), + ("component", ns.app, ai_call(ns.app or "", "print 1", frm, to)), + ] + given = [(name, call) for name, value, call in parts if value] + answers = dict(zip([n for n, _ in given], run_many([c for _, c in given]))) code = 0 - if not ns.app: - out["targeting"]["component"] = { - "ok": None, - "note": "no app_insights_app given: skipped, not failed - requests/dependencies/customMetrics/traces/exceptions are unavailable and the run is logs-only; distributed tracing is a telemetry gap", - } - else: - frm, to = _last_minutes(5) - r = run_az(ai_call(ns.app, "print 1", frm, to)) + for name, value, _ in parts: + if not value: + note = "not persisted" + if name == "component": + note += " - no app_insights_app: requests/dependencies/customMetrics/traces/exceptions are unavailable and the run is logs-only; distributed tracing is a telemetry gap" + elif name == "subscription": + note += " - the CLI's active subscription is the one queried" + out["targeting"][name] = {"ok": None, "note": note} + continue + r = answers[name] results.append(r) - value = ai_rows(r.data)[0].get("print_0") if r.ok and ai_rows(r.data) else None - out["targeting"]["component"] = { - "ok": r.ok, - "value": value, - "error": r.error, - "kind": r.kind, - "diagnosis": "" + if r.ok and name == "resource_group" and r.data is not True: + r.ok, r.kind, r.error = False, "not-found", "az group exists answered false" + t: dict = {"ok": r.ok, "error": r.error, "kind": r.kind} + if r.ok and name == "component": + rows = ai_rows(r.data) + t["value"] = rows[0].get("print_0") if rows else None + if r.ok and name == "subscription": + t["differs_from_active"] = (r.data or {}).get("id") != d.get("id") + t["diagnosis"] = ( + "" if r.ok else DIAGNOSIS.get( r.kind, "read the error: connection, proxy, throttling or service error - report it verbatim and retry; never rewrite it as a targeting failure", - ), - } - if not r.ok: - code = ( - 3 - if r.kind in ("not-found", "not-an-appid") - else (2 if r.kind == "usage" else 1) ) - if ns.workspace: - frm, to = _last_minutes(5) - r = run_az(la_call(ns.workspace, "print 1", frm, to)) - results.append(r) - out["targeting"]["workspace"] = { - "ok": r.ok, - "error": r.error, - "kind": r.kind, - "diagnosis": "" - if r.ok - else DIAGNOSIS.get( - r.kind, - "read the error: the customer ID GUID (not the workspace name) is what -w takes; a connection or rights error says so", - ), - } + ) + out["targeting"][name] = t if not r.ok and code == 0: - code = ( - 3 - if r.kind in ("not-found", "not-a-workspace-id") - else (2 if r.kind == "usage" else 1) - ) + code = _targeting_code(r.kind) + if code == 0 and not (ns.app or ns.workspace): + # identity alone proves nothing the queries read (#657) + code = 3 + out["identity_only"] = True out["connected"] = code == 0 out["commands"] = commands(results) out["failed"] = failures(results) @@ -162,13 +174,25 @@ def render_check(o: dict) -> str: for part, t in o["targeting"].items(): if t.get("ok") is None: out.append(f"targeting {part}: skipped - {t['note']}") + elif t["ok"] and part in ("subscription", "resource_group"): + differs = ( + " - differs from the CLI's active one: the queries target the persisted one" + if t.get("differs_from_active") + else "" + ) + out.append(f"targeting {part}: resolves{differs}") elif t["ok"]: out.append(f"targeting {part}: connected (print 1 answered)") else: out.append( f"targeting {part}: FAILED [{t['kind']}] {t['error']}\n {t['diagnosis']}" ) - out.append("connected" if o["connected"] else "NOT connected") + if o.get("identity_only"): + out.append( + "NOT connected - identity only: the resolved entry persists neither a workspace nor an app_insights_app, so nothing the queries read is proven - route to the switch to persist them" + ) + else: + out.append("connected" if o["connected"] else "NOT connected") out += render_commands(o) return "\n".join(out) @@ -248,6 +272,8 @@ def main() -> int: a = sub.add_parser("check") a.add_argument("--app") a.add_argument("--workspace") + a.add_argument("--resource-group") + a.add_argument("--subscription") a.add_argument("--json", action="store_true") b = sub.add_parser("landing") b.add_argument("--app", required=True) diff --git a/.apm/skills/observability-cli-guides/scripts/azure_monitor_az.py b/.apm/skills/observability-cli-guides/scripts/azure_monitor_az.py index 0a11055..6343fd0 100644 --- a/.apm/skills/observability-cli-guides/scripts/azure_monitor_az.py +++ b/.apm/skills/observability-cli-guides/scripts/azure_monitor_az.py @@ -139,7 +139,12 @@ def classify(stderr: str, code: int) -> tuple[str, str]: "persist the workspace's customer ID" ), ) - if code == 3 or "ApplicationNotFoundError" in text or "ResourceNotFound" in text: + if ( + code == 3 + or "ApplicationNotFoundError" in text + or "ResourceNotFound" in text + or re.search(r"Subscription '.*' not found", text) + ): return "not-found", msg or "the resource does not exist (exit 3)" if "AADSTS" in text or "az login" in text or "re-authenticate" in text.lower(): return ( diff --git a/.apm/skills/observability-cli-guides/scripts/cloudwatch-context.py b/.apm/skills/observability-cli-guides/scripts/cloudwatch-context.py index 9081bdc..bf9353f 100755 --- a/.apm/skills/observability-cli-guides/scripts/cloudwatch-context.py +++ b/.apm/skills/observability-cli-guides/scripts/cloudwatch-context.py @@ -4,9 +4,10 @@ cloudwatch-context.py check --profile --region --log-group --metrics-log-group cloudwatch-context.py landing --profile --region --log-group --until 2026-09-11T12:30:00Z --service orders-api -Whole surface - check: --profile, --region (both required), --log-group and ---metrics-log-group (each optional: given, the group is proved to resolve; -omitted, that part is skipped and the output says so), --json. landing: +Whole surface - check: --profile, --region (both required by the proof: +either missing, the entry is incomplete - exit 3, nothing run), --log-group +and --metrics-log-group (each optional: given, the group is proved to +resolve; omitted, that part is skipped and the output says so), --json. landing: --profile, --region, --log-group (required), --metrics-log-group (optional, polled the same way as a lower bound only), --until (the RFC 3339 UTC instant the newest record must reach - the run's end), --service @@ -15,7 +16,8 @@ default 10), --cap (the bound, default 3m), --json. Exit codes - check: 0 connected (identity and every group given), 1 an identity failure or a rights/network error (the message says what is yours to do), 3 a persisted -group does not resolve (a wrong value: route to the switch), 2 aws refused +group does not resolve or the profile or region is missing (route to the +switch), 2 aws refused the command. landing: 0 landed, 1 the cap was reached (the last newest is in the output), 3/2 as check. @@ -80,6 +82,19 @@ def _code(kind: str) -> int: def cmd_check(ns) -> tuple[int, dict]: register_targets(log_group=ns.log_group, metrics_log_group=ns.metrics_log_group) out: dict = {"identity": {}, "targeting": {}, "connected": False, "commands": []} + missing = [f for f in ("profile", "region") if not getattr(ns, f)] + if missing: + # an entry without them would run under whatever the CLI defaults + # to, possibly another account (#657): incomplete, not a usage error + for f in missing: + out["targeting"][f] = { + "ok": False, + "kind": "not-persisted", + "error": "MISSING - not persisted", + "diagnosis": f"the resolved entry is incomplete: route to the switch to persist {f} ({'default when that profile is the one' if f == 'profile' else 'the region the missions query'})", + } + out["failed"] = [] + return 3, out results = [] ident = run_aws(["sts", "get-caller-identity"], ns.profile, ns.region) results.append(ident) @@ -165,7 +180,9 @@ def cmd_check(ns) -> tuple[int, dict]: def render_check(o: dict) -> str: out = [] i = o["identity"] - if not i.get("ok"): + if not i: + pass # an incomplete entry: nothing was run + elif not i.get("ok"): out.append( f"identity NOT connected [{i.get('kind')}] {i.get('error')}\n {i.get('diagnosis')}" ) @@ -181,6 +198,8 @@ def render_check(o: dict) -> str: out.append( f"targeting {fld}: resolves (retention {t.get('retention_days') or 'never expires'} days, {t.get('stored_bytes')} bytes stored)" ) + elif t["kind"] == "not-persisted": + out.append(f"targeting {fld}: {t['error']} - {t['diagnosis']}") else: out.append( f"targeting {fld}: FAILED [{t['kind']}] {t['error']}\n {t['diagnosis']}" @@ -327,8 +346,8 @@ def main() -> int: ap = argparse.ArgumentParser(description=__doc__.splitlines()[0]) sub = ap.add_subparsers(dest="cmd", required=True) a = sub.add_parser("check") - a.add_argument("--profile", required=True) - a.add_argument("--region", required=True) + a.add_argument("--profile") + a.add_argument("--region") a.add_argument("--log-group") a.add_argument("--metrics-log-group") a.add_argument("--json", action="store_true") diff --git a/.apm/skills/odd-memory/references/observe-run-report.md b/.apm/skills/odd-memory/references/observe-run-report.md index b82f507..e76aa96 100644 --- a/.apm/skills/odd-memory/references/observe-run-report.md +++ b/.apm/skills/odd-memory/references/observe-run-report.md @@ -358,7 +358,8 @@ verify`, a report it cannot read, a repository it cannot compare). its end is `drive`; a chain reaching none is an `ask:` for the mode. A drive needs the user's confirmation when the stack or the record's base URL is not local. Its `verifies` line is - what the replay's `new --verifies` takes. + what the replay's `new --verifies` takes; its `entry:` line, the + `stack_config` entry the pair resolves to. - `boundary ` decides **verification or re-measure**: the baseline's `tree_anchor` against `HEAD` of `--repo`, entry by entry; the tree at `revision` when there is no diff --git a/.apm/skills/odd-memory/scripts/odd_report.py b/.apm/skills/odd-memory/scripts/odd_report.py index 2f49e3a..a04e57f 100755 --- a/.apm/skills/odd-memory/scripts/odd_report.py +++ b/.apm/skills/odd-memory/scripts/odd_report.py @@ -49,6 +49,7 @@ from __future__ import annotations import argparse +import json import re import subprocess import sys @@ -2549,6 +2550,50 @@ def is_local_target(url: str) -> bool: return host.lower() in LOCAL_HOSTS +def resolved_entry(stack: str, environment: str | None) -> str: + """The stack_config entry the replay's pair resolves to, read off the + global configuration the MCP server writes - `-` + when present, else the plain `` said as a degradation (#657); + the local stack takes no environment entry.""" + if not stack or stack == "local": + return stack or "none" + try: + config = json.loads((Path.home() / ".oddyssey" / "config.json").read_text()) + keys = config.get("stack_config") or {} + known = set(config.get("custom") or {}) + except (OSError, ValueError, AttributeError): + keys, known = {}, set() + try: # the built-in stacks, from the table the server's STACKS mirrors + table = ( + Path(__file__).resolve().parents[2] + / "observability-cli-guides/references/builtin-stacks.md" + ).read_text() + known |= set(re.findall(r"^\| `([a-z0-9-]+)` \|", table, re.MULTILINE)) + except OSError: + pass + named = f"{environment}-{stack}" if environment else None + # the server's parse-back rule (#656): the key is the pair's only when + # it is no known stack itself and its longest known suffix is the stack + suffix = max( + (k for k in known - {"local"} if named and named.endswith("-" + k)), + key=len, + default=stack, + ) + if named and named in keys and named not in known and suffix == stack: + return named + if not named: + return stack if stack in keys else f"none (no {stack} entry)" + if stack in keys: + return ( + f"{stack} (no {named} entry - the replay reads the plain entry, " + "whose environment is unverified)" + ) + return ( + f"none (no {named} nor {stack} entry - the replay reads the CLI's own " + "active context, whose environment is unverified)" + ) + + def baseline_facts(root: Path, args: argparse.Namespace) -> dict: resolved = resolve_report(root, args.target, args.service, args.stack, args.env) baseline, how = hop_to_baseline(root, resolved, args.own_protocol) @@ -2581,6 +2626,10 @@ def baseline_facts(root: Path, args: argparse.Namespace) -> dict: "environment": ( None if baseline["kind"] == "instrumentation" else fm.get("environment") ), + "entry": resolved_entry( + stack, + None if baseline["kind"] == "instrumentation" else fm.get("environment"), + ), "mode": mode, "mode_why": mode_why, "revision": fm.get("revision"), @@ -2605,6 +2654,7 @@ def render_baseline(facts: dict) -> str: if facts["kind"] == "instrumentation" else str(env or "none") ), + f"entry: {facts['entry']}", f"mode: {facts['mode']} ({facts['mode_why']})", f"revision: {facts['revision'] or 'none'}", f"benchmark: {', '.join(facts['benchmarks']) or 'none named'}", diff --git a/docs/guide/backends.md b/docs/guide/backends.md index 62917cb..9875761 100644 --- a/docs/guide/backends.md +++ b/docs/guide/backends.md @@ -14,9 +14,11 @@ skill; this page restates it, never extends it. Naming a stack in an A remote backend's values can be persisted per deployment environment, and a mission pointed at one: `persist ... for in ` writes that environment's entry of the stack (`prod-cloudwatch` next -to `cloudwatch`, the same fields), `target ` makes the -missions read it - the stack's plain entry where no such entry exists - -and `/odd-observe checkout on cloudwatch in prod` targets it in passing. +to `cloudwatch`, the same fields; a new one asks for the stack's other +values or copies the plain entry's), `target ` makes the +missions read it - the stack's plain entry where no such entry exists, +which the check says - and `/odd-observe checkout on cloudwatch in prod` +targets it in passing. The local stack takes no environment. The agent still detects the environment from the telemetry and stops when the two diverge. @@ -100,6 +102,7 @@ workspace — without one, tracing is reported as a telemetry gap. ```text /odd-config switch to azure-monitor, app insights "checkout-appinsights" /odd-config switch to azure-monitor, subscription "Contoso Prod", resource group "rg-observability", workspace "log-analytics-prod", app insights "checkout-appinsights" +/odd-config persist workspace "log-analytics-prod" for azure-monitor in prod ``` **Persists**: `subscription`, `resource_group`, `workspace`, and @@ -128,7 +131,8 @@ oddyssey. /odd-config switch to cloudwatch in prod, profile "myteam-prod", region "eu-central-1", log group "/ecs/checkout-prod" ``` -**Persists**: `region`, `profile`, `log_group`, and optionally +**Persists**: `region` and `profile` (both required - `default` when +that profile is the one), `log_group`, and optionally `metrics_log_group` and `xray` — `aws` says who you are, never which log groups the missions read. diff --git a/docs/guide/plugin.md b/docs/guide/plugin.md index fb3d928..318b82c 100644 --- a/docs/guide/plugin.md +++ b/docs/guide/plugin.md @@ -523,7 +523,7 @@ each other. | [`package-layout`](../../.apm/skills/package-layout/SKILL.md) | Where this package is installed and what each part of it is - the skills' root, the sibling directories the install carries, and per skill its `SKILL.md`, its references and its scripts; owns the script that answers it (`scripts/layout.py`: nothing in, the installation's map out - read from the script's own location, so no path is hardcoded and nothing is searched for) | Nothing - read by the prompts' preflights for the mission block's `Skills:` line, and by an agent dispatched without one | | [`otel-guides`](../../.apm/skills/otel-guides/SKILL.md) | Curated map of the official OpenTelemetry docs: every supported language plus the cross-language guides (SDK configuration, semantic conventions, generative AI conventions and instrumentation libraries, Collector deployment, profiling) | Nothing - read by `otel-instrumentation-expert`, and by `observe-run` for the generative AI reference | | [`k6-guides`](../../.apm/skills/k6-guides/SKILL.md) | Curated map of the official k6 docs: install, running a script, scripting (checks, thresholds, scenarios), test types, protocols - including driving an MCP server - and which of a benchmark's inputs a human must decide rather than an agent; owns the script that replays a stored benchmark (`scripts/replay_benchmark.py`: the manifest in, the k6 command built and run and the replay's record out - refusing the flags that would silently edit the benchmark) | Nothing | -| [`odd-memory`](../../.apm/skills/odd-memory/SKILL.md) | The `.odd/` memory: the contract every kind shares, and one reference per kind - observation reports, instrumentation reports, the maintainer-ruling ledgers, benchmarks, custom stacks - saying how to persist, recall and show it; owns the five stores, the script that writes, checks, reads, persists and shows an observation report (`scripts/odd_report.py`: the run's values in, the file's path out, then `check`, `read`, `persist`, `synthesis`, `show` on that path; `baseline` and `boundary` for a replay's preflight - the one reader of the report format, imported by `odd_recall.py`, `odd_ledger.py` and `get-status`'s `odd_status.py`), the script that lists the stored reports and benchmarks a recall considers (`scripts/odd_recall.py`: the mission's scope in, the matches newest first out, one line each), and the script that writes the two ruling ledgers (`scripts/odd_ledger.py`: resolve a finding, record a decision or its reversal, classify a tree entry, checked before the row lands - it reads the global configuration's `stack_config` values, read-only and failing open, so a rationale never carries one) | Nothing - read by the three agents at persist and recall time, by the prompts at show time, by `get-status`, by `backend-configuration` for a custom stack; never invoked on its own | +| [`odd-memory`](../../.apm/skills/odd-memory/SKILL.md) | The `.odd/` memory: the contract every kind shares, and one reference per kind - observation reports, instrumentation reports, the maintainer-ruling ledgers, benchmarks, custom stacks - saying how to persist, recall and show it; owns the five stores, the script that writes, checks, reads, persists and shows an observation report (`scripts/odd_report.py`: the run's values in, the file's path out, then `check`, `read`, `persist`, `synthesis`, `show` on that path; `baseline` and `boundary` for a replay's preflight, `baseline` naming the `stack_config` entry the report's pair resolves to - the one reader of the report format, imported by `odd_recall.py`, `odd_ledger.py` and `get-status`'s `odd_status.py`), the script that lists the stored reports and benchmarks a recall considers (`scripts/odd_recall.py`: the mission's scope in, the matches newest first out, one line each), and the script that writes the two ruling ledgers (`scripts/odd_ledger.py`: resolve a finding, record a decision or its reversal, classify a tree entry, checked before the row lands - it reads the global configuration's `stack_config` values, read-only and failing open, so a rationale never carries one) | Nothing - read by the three agents at persist and recall time, by the prompts at show time, by `get-status`, by `backend-configuration` for a custom stack; never invoked on its own | | [`observability-cli-guides`](../../.apm/skills/observability-cli-guides/SKILL.md) | One reference per stack - query surface, configuration display, what to persist - plus the built-in stack list: the local stack, Grafana (gcx), Datadog (Pup), Dynatrace (dtctl), Azure Monitor (az), CloudWatch (aws); the reference contract every stack reference follows - a custom stack's guide and the query scripts it names included - with the script that checks one against it, and the scripts that query a Grafana stack: `grafana-discover.py` (services and a window in, what each signal holds per service out), `grafana-metrics.py` (a metric and a window in, quantiles, settled counter readings or a label's values out), `grafana-traces.py` (services or a TraceQL selector in, per-operation latency with exemplar span trees or deduplicated counts out; `watch` polls a driven run's identity until it has started and ended, resumable from a state file), `grafana-logs.py` (a stream selector in, exact counts, severities, correlation and samples out), `grafana-profiles.py` (a profile selector in, top frames by self and total out), `grafana-context.py` (a gcx context name in, the proved config to query through out - the user's own in place when the context is current, a per-session copy with datasource defaults otherwise); the reference names them, a mission runs them instead of composing gcx calls; and the scripts that query Azure Monitor the same way: `azure-monitor-discover.py` (the component and workspace tables a window holds per service, the environment read off the resource attributes), `azure-monitor-traces.py` (the per-operation table from `requests`, the dependencies joined on `operation_Id`, exemplars, a trace tree; `watch` polls a driven run's identity until it has started and ended, resumable from a state file), `azure-monitor-metrics.py` (`customMetrics` with the temporality probe, platform metrics with their definitions), `azure-monitor-logs.py` (the workspace's `*_CL` tables and the component's `traces`, a `kql` escape hatch), `azure-monitor-context.py` (the two-part connection proof, the landing poll); and the scripts that query CloudWatch and X-Ray: `cloudwatch-discover.py` (the log groups' freshness and records per service, the EMF namespaces and dimension sets, the service graph, the profiling gap), `cloudwatch-traces.py` (one summaries call split client-side into the per-operation table with client-side percentiles, trace trees batched by five, the service graph's histograms; `watch` polls a driven run's identity until it has started and ended, resumable from a state file), `cloudwatch-metrics.py` (the EMF temporality probe and the window's edge diffs or sums per full dimension set, `get-metric-data` through a queries file), `cloudwatch-logs.py` (Logs Insights with the poll inside, `filter-log-events`, the route normalisation, a `query` escape hatch), `cloudwatch-context.py` (the connection proof with its SSO diagnoses, the landing proofs per signal) | Routes the local-stack case to `setup-local-stack` | | [`setup-local-stack`](../../.apm/skills/setup-local-stack/SKILL.md) | Configure gcx against the local stack without touching the user's contexts, with the datasource UIDs and the push-model caveats; owns the script that writes and proves that context (`scripts/gcx_local.py`: the configured ports in, the isolated context written whole and checked out) and the script that inventories the services (`scripts/probe_services.py`: service names and a window in, per-service signal presence, identity and counter baseline out) | `odd_config_get`; `odd_stack_status` / `odd_stack_up` / `odd_stack_reset` | | [`backend-configuration`](../../.apm/skills/backend-configuration/SKILL.md) | The configured backend, in two sections: `## Check` displays the configured stack's CLI context, proves it connected, guides the setup and hands the preflight over to the mission; `## Switch` owns the change - CLI presence with a guided install offer, the switch, the environment, and the `stack_config` values persisted per stack and, when an environment is named or configured, per environment (`-` entries), then `## Check` for the proof; a custom stack's directory is checked against the reference contract before the switch | `observability-cli-guides` (`builtin-stacks.md`, the stack's four preflight sections, the contract check script); `odd-memory` (the `observability-stack` reference, for a custom stack's directory); `odd_config_get`, `odd_config_set`; routes to `setup-local-stack` for the local stack | diff --git a/docs/guide/prompts.md b/docs/guide/prompts.md index c588676..74f1c52 100644 --- a/docs/guide/prompts.md +++ b/docs/guide/prompts.md @@ -479,7 +479,8 @@ custom stack is written by `/odd-instrument-stack`, never here. /odd-config ``` -No arguments: display, then the "Change backend?" choice. +No arguments: display, then the "Change backend?" choice, naming the +environments already persisted for the stack. ```text /odd-config switch to datadog diff --git a/tests/skills/backend-configuration/test_preflight.py b/tests/skills/backend-configuration/test_preflight.py index 9a86484..118e839 100644 --- a/tests/skills/backend-configuration/test_preflight.py +++ b/tests/skills/backend-configuration/test_preflight.py @@ -309,3 +309,85 @@ def test_a_porcelain_entry_resolves_to_the_path_it_is_about(preflight, line, exp """A rename reported by its source would read as documentation when a file actually landed under the code.""" assert preflight.porcelain_path(line) == expected + + +def stub(bin_dir: Path, name: str, body: str) -> None: + bin_dir.mkdir(exist_ok=True) + path = bin_dir / name + path.write_text("#!/bin/sh\n" + body) + path.chmod(0o755) + + +def test_a_failing_docker_ps_is_a_daemon_unreachable_not_nothing_running( + preflight, tmp_path, monkeypatch +): + """#657: with the daemon down, `docker ps` fails and the empty listing read + as "running nothing" - a false fact the handoff copied.""" + stub( + tmp_path / "bin", + "docker", + 'echo "Cannot connect to the Docker daemon" >&2\nexit 1\n', + ) + monkeypatch.setenv("PATH", str(tmp_path / "bin")) + assert preflight.containers(None) is None + report = {**FULL_REPORT, "containers": None} + assert "; docker daemon unreachable; " in preflight.machine_line(report) + assert "running" not in preflight.machine_line(report) + assert " containers daemon unreachable" in preflight.render(report) + + +def run_script(tmp_path: Path, config: dict | None, bin_dir: Path) -> dict: + home = tmp_path / "home" + (home / ".oddyssey").mkdir(parents=True, exist_ok=True) + if config is not None: + (home / ".oddyssey" / "config.json").write_text(json.dumps(config)) + env = {**os.environ, "HOME": str(home), "PATH": str(bin_dir)} + p = subprocess.run( + [sys.executable, str(SCRIPT), "--root", str(tmp_path), "--json"], + capture_output=True, + text=True, + env=env, + check=False, + ) + assert p.returncode == 0, p.stderr + return json.loads(p.stdout) + + +def test_the_configured_stack_s_cli_is_probed_too(tmp_path): + """#657: the probe was fixed to gcx/k6/docker - the configured stack's own + CLI (its row of builtin-stacks.md) never reached the Machine line.""" + bin_dir = tmp_path / "bin" + # `az version` prints a JSON object, `az --version` checks for updates + stub( + bin_dir, + "az", + 'if [ "$1" = version ]; then echo \'{"azure-cli": "2.89.1", "extensions": {}}\'; ' + "else sleep 5; fi\n", + ) + report = run_script(tmp_path, {"stack": "azure-monitor"}, bin_dir) + assert report["clis"]["az"] == {"present": True, "version": "2.89.1"} + assert ", az 2.89.1;" in report["machine"] + # a CLI refusing `version` answers --version + stub( + bin_dir, + "aws", + 'if [ "$1" = --version ]; then echo "aws-cli/2.36.37 Python/3.14.7"; ' + "else exit 252; fi\n", + ) + report = run_script(tmp_path, {"stack": "cloudwatch"}, bin_dir) + assert report["clis"]["aws"]["version"] == "aws-cli/2.36.37 Python/3.14.7" + assert ", aws 2.36.37;" in report["machine"] + (bin_dir / "aws").unlink() + report = run_script(tmp_path, {"stack": "cloudwatch"}, bin_dir) + assert report["clis"]["aws"] == {"present": False} + + +def test_no_extra_cli_for_the_default_a_custom_or_no_configuration(tmp_path): + bin_dir = tmp_path / "bin" + bin_dir.mkdir() + for config in (None, {"stack": "local"}, {"stack": "grafana"}, {"stack": "seq"}): + assert set(run_script(tmp_path, config, bin_dir)["clis"]) == { + "gcx", + "k6", + "docker", + } diff --git a/tests/skills/observability-cli-guides/fixtures/azure-monitor/2d59a3e2998e.json b/tests/skills/observability-cli-guides/fixtures/azure-monitor/2d59a3e2998e.json new file mode 100644 index 0000000..58837de --- /dev/null +++ b/tests/skills/observability-cli-guides/fixtures/azure-monitor/2d59a3e2998e.json @@ -0,0 +1,13 @@ +{ + "args": [ + "account", + "show", + "--subscription", + "Contoso", + "-o", + "json" + ], + "code": 0, + "stdout": "{\n \"environmentName\": \"AzureCloud\",\n \"homeTenantId\": \"cccccccc-0000-0000-0000-00000000000c\",\n \"id\": \"dddddddd-0000-0000-0000-00000000000d\",\n \"isDefault\": true,\n \"managedByTenants\": [],\n \"name\": \"Contoso\",\n \"state\": \"Enabled\",\n \"tenantDefaultDomain\": \"contoso.onmicrosoft.com\",\n \"tenantDisplayName\": \"Default Directory\",\n \"tenantId\": \"cccccccc-0000-0000-0000-00000000000c\",\n \"user\": {\n \"name\": \"example-user@example.com\",\n \"type\": \"user\"\n }\n}\n", + "stderr": "" +} diff --git a/tests/skills/observability-cli-guides/fixtures/azure-monitor/3171f719e513.json b/tests/skills/observability-cli-guides/fixtures/azure-monitor/3171f719e513.json new file mode 100644 index 0000000..731c232 --- /dev/null +++ b/tests/skills/observability-cli-guides/fixtures/azure-monitor/3171f719e513.json @@ -0,0 +1,15 @@ +{ + "args": [ + "group", + "exists", + "--resource-group", + "contoso-rg", + "--subscription", + "Contoso", + "-o", + "json" + ], + "code": 0, + "stdout": "true\n", + "stderr": "" +} diff --git a/tests/skills/observability-cli-guides/fixtures/azure-monitor/5c773e31fe37.json b/tests/skills/observability-cli-guides/fixtures/azure-monitor/5c773e31fe37.json new file mode 100644 index 0000000..ee168c4 --- /dev/null +++ b/tests/skills/observability-cli-guides/fixtures/azure-monitor/5c773e31fe37.json @@ -0,0 +1,13 @@ +{ + "args": [ + "account", + "show", + "--subscription", + "Missing-Sub", + "-o", + "json" + ], + "code": 1, + "stdout": "", + "stderr": "ERROR: Subscription 'Missing-Sub' not found. Check the spelling and casing and try again.\n" +} diff --git a/tests/skills/observability-cli-guides/fixtures/azure-monitor/c66a91fa58c4.json b/tests/skills/observability-cli-guides/fixtures/azure-monitor/c66a91fa58c4.json new file mode 100644 index 0000000..eb48590 --- /dev/null +++ b/tests/skills/observability-cli-guides/fixtures/azure-monitor/c66a91fa58c4.json @@ -0,0 +1,15 @@ +{ + "args": [ + "group", + "exists", + "--resource-group", + "missing-rg", + "--subscription", + "Contoso", + "-o", + "json" + ], + "code": 0, + "stdout": "false\n", + "stderr": "" +} diff --git a/tests/skills/observability-cli-guides/fixtures/azure-monitor/ebf1b306aee5.json b/tests/skills/observability-cli-guides/fixtures/azure-monitor/ebf1b306aee5.json new file mode 100644 index 0000000..30a5923 --- /dev/null +++ b/tests/skills/observability-cli-guides/fixtures/azure-monitor/ebf1b306aee5.json @@ -0,0 +1,13 @@ +{ + "args": [ + "group", + "exists", + "--resource-group", + "contoso-rg", + "-o", + "json" + ], + "code": 0, + "stdout": "true\n", + "stderr": "" +} diff --git a/tests/skills/observability-cli-guides/test_azure_monitor_scripts.py b/tests/skills/observability-cli-guides/test_azure_monitor_scripts.py index 2ea4efd..dba33b7 100644 --- a/tests/skills/observability-cli-guides/test_azure_monitor_scripts.py +++ b/tests/skills/observability-cli-guides/test_azure_monitor_scripts.py @@ -803,6 +803,74 @@ def test_context_check_proves_identity_and_both_targets(fake): assert "--subscription" not in out["commands"][1] +def test_context_check_proves_the_subscription_and_the_resource_group(fake): + """#657: resource_group and subscription are part of the targeting proof - + one bounded call each (captured live 2026-09-26: 0.3 s and 0.7 s).""" + code, out = fake.json( + "context", + "check", + "--subscription", + SUB, + "--resource-group", + RG, + "--workspace", + WS, + "--app", + APP, + ) + assert code == 0 and out["connected"] is True + assert out["targeting"]["subscription"]["ok"] is True + assert out["targeting"]["resource_group"]["ok"] is True + assert ( + "az group exists --resource-group --subscription -o json" + in out["commands"] + ) + assert "Contoso" not in " ".join(out["commands"]) + + +def test_context_check_a_missing_resource_group_or_subscription_exits_three(fake): + result = fake.run( + "context", + "check", + "--subscription", + SUB, + "--resource-group", + "missing-rg", + "--app", + APP, + ) + assert result.returncode == 3 + assert "targeting resource_group: FAILED [not-found]" in result.stdout + assert "NOT connected" in result.stdout + result = fake.run("context", "check", "--subscription", "Missing-Sub", "--app", APP) + assert result.returncode == 3 + assert "targeting subscription: FAILED [not-found]" in result.stdout + + +def test_context_check_names_every_skipped_part(fake): + result = fake.run("context", "check", "--app", APP) + assert result.returncode == 0 + for part in ("subscription", "resource_group", "workspace"): + assert f"targeting {part}: skipped - not persisted" in result.stdout + assert "\nconnected\n" in result.stdout + + +def test_context_check_on_identity_alone_is_not_connected(fake): + """#657: an entry holding neither a workspace nor a component (an empty or + near-empty environment entry) proves nothing the queries read.""" + result = fake.run("context", "check", "--resource-group", RG) + assert result.returncode == 3 + assert "targeting resource_group: resolves" in result.stdout + assert "targeting workspace: skipped - not persisted" in result.stdout + assert "\nconnected" not in result.stdout + assert "NOT connected - identity only" in result.stdout + result = fake.run("context", "check", "--json") + assert result.returncode == 3 + out = json.loads(result.stdout) + assert out["connected"] is False + assert out["targeting"]["component"]["ok"] is None + + def test_context_check_wrong_values_exit_three_with_the_diagnosis(fake): result = fake.run( "context", "check", "--app", "12345678-1234-1234-1234-123456789abc" diff --git a/tests/skills/observability-cli-guides/test_cloudwatch_scripts.py b/tests/skills/observability-cli-guides/test_cloudwatch_scripts.py index 95673ca..262c1f0 100644 --- a/tests/skills/observability-cli-guides/test_cloudwatch_scripts.py +++ b/tests/skills/observability-cli-guides/test_cloudwatch_scripts.py @@ -455,6 +455,23 @@ def test_context_check_wrong_values_and_a_missing_profile_are_diagnosed(fake): ) +def test_context_check_an_entry_missing_profile_or_region_is_incomplete(fake): + """#657: an environment entry persisting a log group and no profile or + region reads as an incomplete entry routed to the switch - never argparse + usage, never a call under whatever profile the CLI defaults to.""" + result = fake.run("context", "check", "--log-group", LG) + assert result.returncode == 3 + assert "targeting profile: MISSING - not persisted" in result.stdout + assert "targeting region: MISSING - not persisted" in result.stdout + assert "NOT connected" in result.stdout + assert "usage:" not in result.stderr + result = fake.run("context", "check", "--profile", "p", "--log-group", LG) + assert result.returncode == 3 + assert "targeting region: MISSING" in result.stdout + assert "targeting profile" not in result.stdout + assert fake.calls() == [] + + def test_context_landing_proves_the_logs_and_states_the_other_two_signals(fake): code, out = fake.json( "context", diff --git a/tests/skills/odd-memory/test_odd_report.py b/tests/skills/odd-memory/test_odd_report.py index f9809fc..9a40816 100644 --- a/tests/skills/odd-memory/test_odd_report.py +++ b/tests/skills/odd-memory/test_odd_report.py @@ -12,6 +12,7 @@ from __future__ import annotations import importlib.util +import json import os import shutil import subprocess @@ -1749,6 +1750,76 @@ def test_baseline_says_when_a_drive_needs_the_users_confirmation(repo): assert got["drive confirmation"] == "not needed (mode observe)" +def test_baseline_names_the_stack_config_entry_the_replay_s_pair_resolves_to( + repo, tmp_path +): + """#657: the replay's pair is resolved by the script - the environment's + entry, else the plain one said as a degradation - never by hand.""" + home = tmp_path / "home" + (home / ".oddyssey").mkdir(parents=True) + config = home / ".oddyssey" / "config.json" + + def entry(stack_config: dict | None) -> str: + if stack_config is None: + config.unlink(missing_ok=True) + else: + config.write_text( + json.dumps({"stack": "local", "stack_config": stack_config}) + ) + proc = subprocess.run( + [sys.executable, str(SCRIPT), "baseline", "--repo", str(repo.root)], + capture_output=True, + text=True, + cwd=repo.root, + env={**os.environ, **GIT_ENV, "HOME": str(home)}, + check=False, + ) + assert proc.returncode == 0, proc.stderr + return lines_of(proc)["entry"] + + stored( + repo, + "2026-08-08-1000-checkout-sweep.md", + stack="azure-monitor", + environment="dev", + ) + assert entry({"dev-azure-monitor": {"workspace": "x"}, "azure-monitor": {}}) == ( + "dev-azure-monitor" + ) + assert entry({"azure-monitor": {"workspace": "x"}}) == ( + "azure-monitor (no dev-azure-monitor entry - the replay reads the plain " + "entry, whose environment is unverified)" + ) + assert entry(None) == ( + "none (no dev-azure-monitor nor azure-monitor entry - the replay reads the " + "CLI's own active context, whose environment is unverified)" + ) + # the server's parse-back rule (#656): a built-in's own entry is never + # another pair's - a custom "monitor" in "azure" does not read azure-monitor + config.write_text( + json.dumps( + { + "custom": {"monitor": {"stack_config_fields": []}}, + "stack_config": {"azure-monitor": {}, "monitor": {}}, + } + ) + ) + stored( + repo, "2026-08-08-1100-checkout-sweep.md", stack="monitor", environment="azure" + ) + proc = subprocess.run( + [sys.executable, str(SCRIPT), "baseline", "--repo", str(repo.root)], + capture_output=True, + text=True, + cwd=repo.root, + env={**os.environ, **GIT_ENV, "HOME": str(home)}, + check=False, + ) + assert lines_of(proc)["entry"].startswith("monitor (no azure-monitor entry") + stored(repo, "2026-08-09-1000-checkout-sweep.md") # local, environment local + assert entry({"local": {"GF_LOG_LEVEL": "debug"}}) == "local" + + def test_baseline_with_nothing_stored_says_so(repo): proc = run(repo, "baseline", "--repo", str(repo.root)) assert proc.returncode == 2 and "nothing to verify" in proc.stderr