From b320eba4d8170827ed04544cab96932e85d133df Mon Sep 17 00:00:00 2001 From: Tanishka SInghal Date: Fri, 11 Sep 2026 11:20:32 +0530 Subject: [PATCH 1/4] Add migrate-tml-copies-to-publishing skill MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds the first guided-procedure skill to the plugin: moving per-tenant Model copies created by TML export/import onto Orgs Publishing, so one governed Model serves every Org instead of a copy per Org. The skill is structured so that the whole assessment — can the governed Model stand in for the tenant's copy — happens offline from two TML exports, before anything is written. Every write step is gated on a confirmation, and the tenant's copy is never deleted, because it is the rollback. - skills/migrate-tml-copies-to-publishing/SKILL.md — the procedure, prerequisites including the Primary-Org privilege model, the errors the platform raises and what each means, and the View and nested dependent traps in repointing - references/assessing-two-exports.md — the offline comparison: table pairing on column sets, which differences are expected, the column map, verdicts, and the multi-join-path hole - CLAUDE.md, README.md — list the skill, split into documentation lookup and guided procedures Co-Authored-By: Claude Opus 5 (1M context) --- CLAUDE.md | 10 +- README.md | 1 + .../migrate-tml-copies-to-publishing/SKILL.md | 169 ++++++++++++++++++ .../references/assessing-two-exports.md | 101 +++++++++++ 4 files changed, 280 insertions(+), 1 deletion(-) create mode 100644 skills/migrate-tml-copies-to-publishing/SKILL.md create mode 100644 skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md diff --git a/CLAUDE.md b/CLAUDE.md index 83a9210..b04389a 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,9 +1,17 @@ # SpotterCode Plugin -This plugin integrates ThoughtSpot developer documentation with Claude Code, giving you semantic search over the Visual Embed SDK, REST API v2, and developer guides. +This plugin integrates ThoughtSpot developer documentation with Claude Code, giving you semantic search over the Visual Embed SDK, REST API v2, and developer guides — plus guided procedures for platform migrations. ## Skills +### Documentation lookup + - **get-visual-embed-sdk-reference** — Guidance for embedding ThoughtSpot content in web applications - **get-rest-api-reference** — Guidance for working with the ThoughtSpot REST API v2 - **get-developer-docs-reference** — Guidance for searching general ThoughtSpot developer documentation + +### Guided procedures + +- **migrate-tml-copies-to-publishing** — Move per-tenant Model copies made by TML export/import onto Orgs Publishing, so one governed Model serves every Org + +Procedure skills write to your instance. They gate every write behind a confirmation and never delete the object being replaced. diff --git a/README.md b/README.md index 887cf85..bdc19e2 100644 --- a/README.md +++ b/README.md @@ -7,6 +7,7 @@ Configuration for integrating the ThoughtSpot developer documentation MCP server - **Get Visual Embed SDK Docs** — Types, interfaces, configuration, events, authentication, and CSS theming - **Get REST API Docs** — Endpoint specs, request/response schemas, authentication, and Java/TypeScript SDK guides - **Get Developer Docs** — Platform features, SSO/SAML, deployment, TML, and general ThoughtSpot guidance +- **Migrate TML Copies to Publishing** — Guided procedure for moving per-tenant Model copies onto Orgs Publishing, assessing from two TML exports before anything is written ## Prerequisites diff --git a/skills/migrate-tml-copies-to-publishing/SKILL.md b/skills/migrate-tml-copies-to-publishing/SKILL.md new file mode 100644 index 0000000..c970012 --- /dev/null +++ b/skills/migrate-tml-copies-to-publishing/SKILL.md @@ -0,0 +1,169 @@ +--- +description: Move per-tenant Model copies created by TML export/import onto ThoughtSpot Orgs Publishing, so one governed Model serves every Org instead of a copy per Org +--- + +# Migrate TML Model copies to Orgs Publishing + +Customers who adopted Orgs before Publishing existed usually have **one Model per tenant**, +each made by exporting TML from the Primary Org and importing it into a tenant's Org. Every +schema change then has to be repeated per Org, and the copies drift. + +This skill converts that estate to the publishing pattern: **one** governed Model in the +Primary Org, published into each tenant Org, with each tenant's Answers and Liveboards +repointed onto it. The tenant's copy is **kept**, not deleted — it is the rollback. + +## When to Use + +Apply this skill when a customer has per-Org Model copies and wants to adopt Orgs +Publishing without rebuilding their content, and when the inputs available are TML exports +rather than access to both Orgs. + +Not this skill: distributing a Model outward to Orgs that hold no copies yet (ordinary +publishing), or moving a tenant into a brand-new Org. + +## The question you are actually answering + +**Can the governed Model stand in for the tenant's copy?** That is structural, not a naming +question. Two Models with different names, different column labels and different warehouse +schemas can be perfect substitutes; two with identical names can be incompatible. + +What decides it is whether every column the tenant's content reads still resolves to the +same warehouse column, and whether the shape carrying those columns agrees. Both are +decidable offline from two TML exports — see +[references/assessing-two-exports.md](references/assessing-two-exports.md). + +## Prerequisites + +- **Orgs enabled**, and Orgs Publishing enabled for your instance. Confirm with ThoughtSpot + — availability varies by release +- **Administrator privilege in the Primary Org.** Objects can only be published from the + Primary Org, and modifying a published object requires administrator privilege *there* — + an Org administrator of a tenant Org cannot do it, even for their own Org's content +- **Membership in every Org involved.** The dependent lookup and the repoint run inside the + tenant Org and take their scope from your session. A missing membership shows up as an + **empty dependent list**, not as a permission error — which reads exactly like "nothing + to migrate" +- Two TML exports per Model: one from the Primary Org, one from the tenant Org + +## Procedure + +Steps 1–5 write nothing. Confirm with the customer before each step from 6 onward. + +| # | Step | Writes? | +|---|---|---| +| 1 | Collect both TML exports per Model | no | +| 2 | Pair the tables and compare structure | no | +| 3 | Review findings and the column map | no | +| 4 | Identify the fields that differ per Org | no | +| 5 | Report the verdict: READY / RENAME / BLOCKED | no | +| 6 | Parameterize the per-Org fields on the governed Model | yes | +| 7 | Re-key the copy's custom object id | yes | +| 8 | Publish the governed Model into the tenant Org | yes | +| 9 | Find what depends on the copy | no | +| 10 | Repoint each dependent | yes | +| 11 | Verify | no | + +**Export both with dependencies**, rooted at the Model, not at a Liveboard: +`export_associated: true`, `export_fqn: true`, `edoc_format: YAML`. + +Two silent mistakes to catch at Step 1: both exports taken from the same Org (they compare +perfectly clean, because they are the same object — compare the root GUIDs), and platform +entries in the zip that will not parse. Never filter zip entries by name to work around the +second; name the entry that failed. + +## Order matters: 6 and 7 before 8 + +**Parameterization is the work, not a preliminary.** A Model whose dependency tree contains +no template variable cannot be published at all — the platform refuses with *"No template +variable node found in the dependency tree."* The per-Org fields found at Step 4 (typically +the warehouse `schema`, sometimes the database or table name) become variables with a value +per Org. + +One trap: that check is an **existence** check across the whole tree, not a completeness +check. If *something* in the tree is parameterized the publish succeeds — even if a table +you meant to parameterize was missed. A partially parameterized publish exposes Primary Org +data to the tenant. Verify every field you intended, per Org, yourself. + +**The copy inherited the original's custom object id.** That id must be unique within an +Org, and publishing makes the governed Model a member of the tenant's Org — so both objects +are then in one Org carrying one id, and the publish is refused. Re-key **the copy**, not +the governed Model, via `POST /api/rest/2.0/metadata/headers/update` from a tenant-Org +session. Always set a new value; never blank it. + +## Let the platform refuse what only it can see + +Publishing validates before writing anything, and the commit is transactional — a refused +publish leaves the estate completely unchanged. So an attempted publish is cheap. It +refuses on: + +| Error | Meaning | +|---|---| +| `Cannot publish/unpublish objects with Cohort Column as dependency` | a Set is in the closure. Sets are invisible in TML, so do not try to detect one from the exports — attempt the publish and read this | +| `No template variable node found in the dependency tree` | Step 6 was skipped or did not land | +| duplicate custom object id | Step 7 was skipped for this Org | +| `Objects can only be published/unpublished from primary org` | wrong Org session | + +Spend the offline effort on what TML *can* settle — correspondence and shape. Let the +platform refuse the rest. + +## Repointing, and its two traps + +A dependent references the **columns** it reads, so what moves it is swapping the Model +reference it resolves through: export its TML, swap `tables[].fqn` from the copy's GUID to +the governed Model's GUID, apply the column map, re-import over the same object. + +**Views are a layer, not an endpoint.** A dependent that is a View or SQL View exposes its +own column names to everything built on it. Rewrite the names it *reads* — `tables[].fqn`, +formulas, column references — and **preserve exactly** the names it *exposes* and its own +name. Get this wrong and every Answer above the View breaks, and none of them appear in the +copy's dependent list, because they depend on the View. Get it right and content on that +View needs no rewrite at all. Do Views first, then re-run the dependent lookup. + +**Some dependents are nested.** A dependency can be carried by a hidden object belonging to +a Liveboard rather than by the Liveboard itself. Repoint the **owner**, and de-duplicate — +several nested objects in one Liveboard collapse to one rewrite. + +Check coverage **before** importing: scan the rewritten document for any surviving reference +to a source column name. A partial rewrite imports cleanly and renders wrong, which the +customer finds rather than you. After importing, confirm the object resolves to the governed +Model. + +## What a TML round trip does not carry + +Presentation, not meaning. Conditional formatting, column widths, chart state, parameter +ids and Liveboard tile layout are not guaranteed to survive a re-import. Which columns an +Answer reads, what it computes and what it filters are. + +So **do not refuse a migration to protect formatting** — that leaves the Answer pointing at +a Model about to be retired, which is worse. Instead: disclose per dependent before writing, +keep the pre-repoint export as both rollback and reference, and verify that the Answer +returns the **same numbers**, never that the documents match. + +## Safety + +Nothing is deleted. At every stage there is a way back: + +| Stopped after | To undo | +|---|---| +| 1–5 | nothing was written | +| 6 | un-parameterize the fields | +| 7 | set the previous custom object id back | +| 8 | unpublish from that Org | +| 10, partly | re-import the pre-repoint TML for the ones already moved | + +**Validate the whole sequence in a non-production Org before running it against a tenant's +live content**, and repoint one dependent at a time rather than in a batch — a failure +part-way leaves the rest still pointing at the copy, which is a safe state. + +## Not covered + +- **Deleting the copy.** Deliberately: it is the rollback. Retiring it is a separate + decision, after verification +- **Answers and Liveboards as publish roots.** They are read as dependents to be repointed +- **Sets / Cohorts.** They block publishing and cannot be seen in TML. Resolve with the + customer +- **Sharing and permissions.** TML carries no sharing information, so nothing here + preserves or migrates it. Check the published Model's grants in each tenant Org explicitly +- **Tenant data isolation.** Publishing one Model to several Orgs does not by itself + separate their data — that comes from row-level security or from the publication + variables. It is a required review, and it is not established by this migration diff --git a/skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md b/skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md new file mode 100644 index 0000000..c2f636b --- /dev/null +++ b/skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md @@ -0,0 +1,101 @@ +# Assessing two TML exports + +Everything on this page is decidable **offline**, from the two export zips, with no +connection to either Org. It is the whole of Steps 1–5, and it is where the judgement in +this migration actually lives. + +## What TML carries, and why that is enough + +A Model export carries more than labels: + +| in the TML | what it settles | +|---|---| +| `db_name`, `schema`, `db_table`, `db_column_name` | the **physical binding** — which warehouse column a Model column reads | +| `model_tables` / `fqn` | which tables the Model is built from | +| each join's `on`, `type`, `cardinality` | the **shape** carrying those columns | +| `formula` | derived columns, comparable as text | +| column `name`, `description` | the tenant-facing labels, which may differ freely | + +So correspondence ("is this the same warehouse column?") and shape ("does the schema graph +agree?") are both computable without touching a cluster. Names are the one thing that may +differ, and the mapping between them is the output. + +## Pair the tables before comparing anything + +Do not pair by name — a copy may have been renamed. Pair on **column sets**: for each table +in the governed export, find the table in the tenant export whose `db_column_name` set +overlaps most. A clean pair overlaps at or near 1.00. + +Report the overlap for every pair. An overlap well below 1.00 that still wins is the +signal that the two Models have diverged, and it deserves a human look rather than a +verdict. + +## Compare, and what may legitimately differ + +Three differences are expected and are **not** findings: + +- **`schema`** (and sometimes `db_name` or `db_table`) — this is exactly what per-Org + parameterization exists to carry. Collect these; they become Step 6's work +- **connection name** — compare the connection *type* instead +- **GUIDs** — every object has its own + +Everything else is a finding. In particular: + +- a column the tenant's content uses that the governed Model does not have — **blocking**, + and no mapping can fix it. Ask the customer; do not guess a near match +- a differing `formula` for columns that otherwise correspond — blocking +- a join present on one side and not the other, or with a different `on` clause + +`cardinality` is **view-relative**: the same join reads `ONE_TO_MANY` or `MANY_TO_ONE` +depending on which table's block it is written under. Normalise before comparing, or every +join looks changed. + +## The column map is the output + +Where corresponding columns carry different names, record +`tenant name -> governed name`. That map drives the repoint later, and it is the artifact +worth reviewing with the customer before any write: + +- an **identity** map means no content rewriting is needed for names at all — the cheapest + possible migration +- a non-empty map means every dependent must be rewritten, and each rewrite is a chance to + get a name wrong + +Review the map before Step 6, not after Step 8. + +## Verdicts + +| verdict | meaning | +|---|---| +| `READY` | corresponds cleanly, identity column map — publish and repoint | +| `RENAME` | corresponds cleanly, but names differ — same path, plus content rewriting | +| `BLOCKED` | something no mapping can fix. Report it and stop | + +Running this across a whole estate first is worth far more than doing it one Model at a +time: a sweep that reports twelve Models `READY` and three `BLOCKED` is one conversation +with the customer instead of fifteen. + +## Two holes to be honest about + +**Multi-join-path.** Where the same table occupies several slots in the schema graph +(role-playing dimensions — an order date and a ship date on the same calendar table), the +column sets are identical and pairing cannot tell the slots apart. Detect the condition — +the same table appearing more than once in the join graph — and escalate rather than +reporting a clean pair. + +**Where a pair binds the same warehouse column under different labels**, the comparison +sees a match on binding and a difference on name, which is exactly the `RENAME` case. That +is correct. But two *different* columns that happen to bind the same warehouse column are +indistinguishable from each other on binding alone. Fall back to name and formula, and +surface both candidates rather than choosing. + +## Two things TML cannot see at all + +- **Sets / Cohorts.** They do not appear in a TML export, and they block publishing. + A clean assessment is not evidence that there is no Set. The publish attempt is the only + test, and it refuses without writing anything +- **Sharing and permissions.** No sharing information is carried, so a clean assessment + says nothing about who can see what afterwards + +A clean assessment means *the governed Model can stand in for the copy*. It does not mean +the migration will succeed. From c55ab6419a2cbf32c42c439ac990f442c40c4fa1 Mon Sep 17 00:00:00 2001 From: Tanishka SInghal Date: Mon, 28 Sep 2026 15:49:08 +0530 Subject: [PATCH 2/4] Replace untested TML-copies skill with the live-tested design Remove the first draft of migrate-tml-copies-to-publishing: it kept the copy, left sharing out and re-keyed only the Model, none of which held up against a live cluster. Add the design (docs/migrate-tml-copies-design.md) and the test plan with results (docs/migrate-tml-copies-test-scenarios.md). The basic scenario - one shared Answer on an identical copy - passed end to end on a live multi-Org cluster: compare, publish with a per-Org variable, carry the copy's sharing, repoint in place, verify the same numbers as a tenant user, and delete the copy with its Tables and connection. The skill itself is rebuilt from this design in the documented build order. Co-Authored-By: Claude Opus 5.5 (1M context) --- CLAUDE.md | 6 - README.md | 1 - docs/migrate-tml-copies-design.md | 1095 +++++++++++++++++ docs/migrate-tml-copies-test-scenarios.md | 274 +++++ .../migrate-tml-copies-to-publishing/SKILL.md | 169 --- .../references/assessing-two-exports.md | 101 -- 6 files changed, 1369 insertions(+), 277 deletions(-) create mode 100644 docs/migrate-tml-copies-design.md create mode 100644 docs/migrate-tml-copies-test-scenarios.md delete mode 100644 skills/migrate-tml-copies-to-publishing/SKILL.md delete mode 100644 skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md diff --git a/CLAUDE.md b/CLAUDE.md index b04389a..7404bf2 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,9 +9,3 @@ This plugin integrates ThoughtSpot developer documentation with Claude Code, giv - **get-visual-embed-sdk-reference** — Guidance for embedding ThoughtSpot content in web applications - **get-rest-api-reference** — Guidance for working with the ThoughtSpot REST API v2 - **get-developer-docs-reference** — Guidance for searching general ThoughtSpot developer documentation - -### Guided procedures - -- **migrate-tml-copies-to-publishing** — Move per-tenant Model copies made by TML export/import onto Orgs Publishing, so one governed Model serves every Org - -Procedure skills write to your instance. They gate every write behind a confirmation and never delete the object being replaced. diff --git a/README.md b/README.md index bdc19e2..887cf85 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,6 @@ Configuration for integrating the ThoughtSpot developer documentation MCP server - **Get Visual Embed SDK Docs** — Types, interfaces, configuration, events, authentication, and CSS theming - **Get REST API Docs** — Endpoint specs, request/response schemas, authentication, and Java/TypeScript SDK guides - **Get Developer Docs** — Platform features, SSO/SAML, deployment, TML, and general ThoughtSpot guidance -- **Migrate TML Copies to Publishing** — Guided procedure for moving per-tenant Model copies onto Orgs Publishing, assessing from two TML exports before anything is written ## Prerequisites diff --git a/docs/migrate-tml-copies-design.md b/docs/migrate-tml-copies-design.md new file mode 100644 index 0000000..88c3ae1 --- /dev/null +++ b/docs/migrate-tml-copies-design.md @@ -0,0 +1,1095 @@ +# Design: migrate TML Model copies to Orgs Publishing + +Status: **Draft — all steps designed.** Open questions are answered by testing on a live +cluster before each step is coded. +Skill: [`skills/migrate-tml-copies-to-publishing`](../skills/migrate-tml-copies-to-publishing/SKILL.md) + +## Context + +Before Orgs Publishing, customers shared a Model with other Orgs by exporting its TML from +the Primary Org and importing it into each secondary Org. Each secondary Org now holds its +own **copy**, with its own Answers and Liveboards built on it. + +The goal is to move each copy onto the publishing flow: publish the governed Model from the +Primary Org into the secondary Org, repoint the content built on the copy at the published +Model, carry over the copy's permissions, and then **delete the copy**. + +**The rollback is the TML files**, not the copy: the copy's export (input 3) restores the +copy, and each dependent's pre-repoint export restores that dependent. Because the copy is +deleted, those files are the only way back, so they are validated and archived before +anything is written. + +**Phase 1 scope: Models only.** Answers and Liveboards are handled only as dependents to be +repointed, not as objects to publish. + +**Steps** + +| Step | What | Writes? | +|---|---|---| +| 1 | Validate and compare the two TMLs | no (offline) | +| 2 | Live check and dependents | no | +| 3 | Publish the governed Model into the secondary Org | yes | +| 4 | Give the published Model the copy's permissions | yes | +| 5 | Repoint dependents | yes | +| 6 | Verify | no | +| 7 | Delete the copy | yes | + +## Summary and build order + +**What the skill does.** Given the governed Model's TML (Primary Org), the copy's TML and +the secondary Org's name, it proves the governed Model can stand in for the copy, publishes +it into the secondary Org, gives it the copy's permissions, repoints everything built on +the copy, verifies, and deletes the copy. The TML files, the sharing snapshots and the +ledger are the rollback. + +**How it runs.** Step 1 is local. Every cluster call goes through spotter-code's +`execute-thoughtspot-code` with `org_identifier` (the `org-aware-code-exec` branch is a +hard dependency). Reads run freely; each write is a separate run behind its own approval. +State between runs lives in `migration///plan.json`. + +**Found live — sign-in always lands in the default Org.** The spotter-code OAuth sign-in +fetches its token with `callosum/v1/v2/auth/token/fetch` and no `org_identifier`, so the +token is scoped to the user's **default** Org whatever Org the browser is in. A multi-Org +admin can therefore never reach a secondary Org through the production tool: +`org_identifier` is a hard requirement, not a convenience. The sign-in also depends on +the cluster's IAM login (`callosum/v1/saml/login`) and fails on a cluster with IAMv2 +disabled. For testing only, the server's `/bearer/mcp` route accepts a ThoughtSpot token +per Org (`Authorization: Bearer `, `x-ts-host: `), obtained with +`auth/token/full` and `org_id`. + +**Live vs downtime.** Steps 1–3 run live. Steps 4–7 run in a downtime window; how it is +enforced for one Org is still open (Q18). + +**What ships in the skill** (`skills/migrate-tml-copies-to-publishing/`): + +| File | Used by | Runs | +|---|---|---| +| `SKILL.md` | the agent | — (the procedure; rewritten from this design) | +| `scripts/fingerprint.mjs` | Steps 1, 2a, 7a | locally **and** in the sandbox — one function, both places | +| `scripts/compare.mjs` | Step 1 | locally (Node; `yaml` package pinned) | +| `scripts/dependents.js` | Steps 2, 6a, 7a | sandbox, read-only | +| `scripts/publish.js` | Step 3 (3a / 3b / 3c as separate modes) | sandbox, write | +| `scripts/grants.js` | Steps 4, 6c, 7a | sandbox, read + write | +| `scripts/repoint.js` | Step 5 (backup / rewrite+import modes) | sandbox, read + write | +| `scripts/verify.js` | Step 6 | sandbox, read-only | +| `scripts/delete.js` | Step 7 | sandbox, write | + +**Build order** — read-only first, and each step only after its open questions are answered +on a live cluster: + +1. `fingerprint.mjs` + `compare.mjs` (Step 1) — offline, testable with any two exports. +2. `dependents.js` (Step 2) — first cluster use, read-only. +3. `publish.js` (Step 3). +4. `grants.js` (Step 4). +5. `repoint.js` (Step 5) — the largest; port the rules and tests from `ts-migrate-orgs` + `rewrite.py`. +6. `verify.js` (Step 6). +7. `delete.js` (Step 7). +8. Rewrite `SKILL.md` around the finished scripts. + +## Running the migration — downtime (to be decided) + +Steps 1–3 run live: none of them changes what secondary-Org users see or use. Steps 4–7 +change the objects users work with, and **run in a downtime window** — relying on users not +to edit during a live run was rejected. + +How downtime is enforced for one Org is open, and is settled by testing on a live cluster. +Options considered, from the REST spec: + +| Option | Per Org? | Status | +|---|---|---| +| Set the Org inactive | — | No endpoint sets it (`orgs/search` only filters on `IN_ACTIVE`) | +| `users/deactivate` | No, cluster-wide | Rejected: reactivation needs a token and a new password per user | +| `orgs/update` `REMOVE` users from the Org | Yes | Likely drops their group memberships and content in the Org — untested | +| `NO_ACCESS` on the copy and its dependents | Yes | Authors and admins keep access; two extra writes per object | +| **Customer blocks access outside ThoughtSpot** (embedded app stops issuing tokens / SSO route off), then the skill ends open sessions with `users/force-logout` on an explicit user list | Yes | **Proposed.** Never call `force-logout` with an empty list — that logs out every user on the cluster | + +Whatever the mechanism, these checks run in Steps 4–7 and **stop the run** if they fire, +because a hit means the downtime leaked: + +1. **At the start of Steps 4–7** — re-run the dependent lookup (2b), the copy's sharing, and + the copy file check (2a), and take the "before" data fingerprints (6b). Phase 1 may be + days old, and fingerprints must be as close to the repoint as possible. +2. **Per object in Step 5** — record the last-modified time at backup; re-read it just + before import. Different means the object was edited after its backup. +3. **Before Step 7** — the copy must have zero dependents, and any share added to the copy + since Step 4 is applied to the published Model first. + +## Step 1 — Validate and compare the two TMLs + +Offline. Reads two files, writes only a local plan file. No cluster access. + +### Inputs + +The customer provides: + +| # | Input | Used for | +|---|---|---| +| 1 | Governed Model TML export, from the Primary Org | Comparison; gives the governed GUID and `obj_id` | +| 2 | Secondary Org name | Recorded in the plan; first used in Step 2 | +| 3 | Copy Model TML export, from the secondary Org | Comparison; gives the copy GUID and `obj_id` | + +Both exports must be taken **with dependencies** (the Model plus its Tables), so each Model +column can be traced to a physical warehouse column. + +**Why TML files rather than exporting from the cluster.** `execute-thoughtspot-code` runs +in one Org per call and returns at most 24,000 characters (`DEFAULT_OUTPUT_MAX_CHARS` in +spotter-code `src/common/code-exec/consts.ts`, overridable by `CODE_EXEC_OUTPUT_MAX_CHARS`). +A Model export with its Tables often exceeds that, and the two sides live in different Orgs, +so they cannot be compared in one run. Local files have no such limit. + +### 1a. Validate the inputs + +Each check stops Step 1 on failure. None is downgraded to a warning, because each failure +leads to a comparison that looks clean but is wrong. + +| Check | Why | On failure | +|---|---|---| +| Every file in each export parses | A skipped file silently drops a Table from the comparison | Stop and name the file. Never filter files by name to get past it | +| Each export contains exactly one Model, its Table files, and the Connection file of every Table | Without Table TML there are no physical bindings to compare; without Connection TML there is no way to tell whether the published connection can reach the copy's data | Stop; ask for a re-export with dependencies (`export_associated: true`) — a UI export with dependencies includes both | +| The two Model GUIDs differ | Equal GUIDs mean both exports came from the same Org, which compares perfectly clean | Stop | +| Every Model table resolves to exactly one Table file (rules below) | A wrong match binds columns to the wrong physical table | Stop; ask for a re-export with FQNs (`export_fqn: true`) | +| Every Table file in the export is used by the Model | An unused file suggests the wrong or a mixed export | Stop and name the file | + +**Resolving a Model's tables to the Table files in its own export** + +| Model table reference | Action | +|---|---| +| Has `fqn` (GUID) | Match to the Table file with that GUID. Preferred | +| No `fqn`; the name matches exactly one Table file in the export | Match by name; note it in the report | +| No `fqn`; the name matches several files, or none | Stop; ask for a re-export with FQNs | + +A name is used only when it is provably unique **within the export**. The export holds only +this Model's own Tables, so uniqueness there is checkable; uniqueness across the cluster is +never assumed. + +### 1b. Build a fingerprint of each side + +Both sides are reduced by the same function, so they are always compared like for like. + +``` +connections: key → { name, type, properties (non-secret only), selected_databases } +tables: key → { db, schema, db_table, connection, variables, + columns: { db_column_name → data_type }, + rls_rules } +model: columns: [{ name, binding, formula, aggregation, column_type }] + joins: [{ left, right, on, type, cardinality }] + filters, parameters +``` + +- `binding` is `.` for a column on a Table; empty for a formula. +- `formula` has its column references rewritten to bindings, so two formulas that read the + same warehouse columns under different column names compare equal. References take the + form `[TABLE::column]`, where `column` is the Table's column **name** — resolve it to + `db_column_name` through the Table TML. Whitespace is normalised before comparing: + found live, the same formula exported as `'… [SALES_FACT::COST] '` on one side and + without the trailing space on the other. +- `cardinality` is normalised to one direction. The same join reads `ONE_TO_MANY` or + `MANY_TO_ONE` depending on which side it is written under. +- `db`, `schema` and `db_table` are kept **raw** as well as resolved. A parameterized + field reads `${variable_name}` in TML; `variables` records which field carries which + variable. This tells Step 3 offline whether the governed Tables are already + parameterized (case A) or not (case B). +- **Connections are compared by content, never by name.** After the migration the Org + reads through the **published** connection, so what matters is whether it can reach the + data the copy's connection reached: its `type` and non-secret `properties` + (`accountName`, `warehouse`, `role`, `user`, …) and `selected_databases`. +- **Secrets are never read, compared or stored.** Exports blank them (`password: ""`); any + property that is secret is dropped from the fingerprint. + +### 1c. Pair the Tables across the two exports + +GUIDs do not help here — the copy's Tables have their own GUIDs. Names do not decide it +either. + +1. Pair on `db_table`. +2. Otherwise pair on the Table whose `db_column_name` set overlaps most. Report the overlap + for every pair; below 1.0 is flagged for review. +3. **Stop** when two candidates tie, or when the same physical Table appears more than once + in a Model (role-playing dimensions such as order date and ship date on one calendar + table). Column sets cannot tell those slots apart; ask rather than pick. + +### 1d. Compare + +**Tables** + +| Difference | Meaning | +|---|---| +| `db`, `schema` or `db_table` | Expected. Becomes the secondary Org's value for the publishing variable | +| Connection name | Informational. The copy's connection is not used after the migration | +| Connection type | Blocker. The published connection cannot reach data on another platform | +| Connection properties (account, warehouse, role, user) | Finding. The published connection reads with the Primary Org's settings unless the Org gets `CONNECTION_PROPERTY` values (3a); the customer chooses which | +| `selected_databases` lacks the copy's `db` | Blocker. The published connection cannot see the copy's database | +| RLS rules | Security finding. After the repoint the Org gets the governed Model's rules | + +The `db` / `schema` / `db_table` values are where the copy's data lives. With identical +connection type and properties, the published connection reaches them for this Org — +decided here, offline. What TML cannot show is whether the connection's role has been +granted the copy's schema in the warehouse; Step 3d confirms it on the cluster by reading +through the published Model. + +**Columns** — matched on binding (formulas on normalised formula), never on name + +| Class | Rule | Effect | +|---|---|---| +| `MATCHED` | Same binding, same name | None | +| `RENAMED` | Same binding, different name | Entry in the column map, `copy name → governed name` | +| `MISSING` | In the copy, not in the governed Model | Blocker if content uses it | +| `CHANGED` | Same binding or name; different formula, aggregation or type | Blocker if content uses it | +| `EXTRA` | In the governed Model only | Informational | + +Where two different columns bind the same warehouse column, binding alone cannot tell them +apart: fall back to name and formula, and report both candidates rather than choosing. + +**Model** — any difference in joins (after normalising cardinality), filters or parameters +is a finding. + +### Output + +A report for the customer, and `migration///plan.json`: + +```json +{ + "secondary_org": "ACME", + "governed": { "guid": "…", "name": "Sales", "obj_id": "sales_model" }, + "copy": { "guid": "…", "name": "Sales", "obj_id": "sales_model" }, + "table_pairs": [{ "governed": "…", "copy": "…", "overlap": 1.0 }], + "verdict_provisional": "RENAME", + "column_map": { "Segment": "STRING_1" }, + "org_values": { "schema": "ACME_PROD" }, + "findings": [], + "needs_rekey": true +} +``` + +| Verdict | Meaning | +|---|---| +| `READY` | Corresponds cleanly; identity column map | +| `RENAME` | Corresponds cleanly; names differ, so content must be rewritten | +| `BLOCKED` | A difference no mapping can fix | + +The verdict is **provisional**: `MISSING` and `CHANGED` block only if content uses those +columns, which Step 2 determines from the copy's dependents. + +`needs_rekey` lists **every** copy object whose `obj_id` equals the governed counterpart's — +the Model, its Tables and its Connection. Found live: a TML import carries all of them +over (`SALES_FACT-40639b92`, `PRODUCT_DIM-9f6ea9ff`, `PRIMARYconn-deaca893`), not only the +Model's. Publishing puts both in one Org with one `obj_id`, which is refused. + +### Archive the inputs + +The copy's export is the rollback for the copy, so Step 1 copies both exports unchanged +into `migration///inputs/` and records a checksum of each in +`plan.json`. Later steps restore from this archive, never from wherever the customer +originally put the files. The validation in 1a doubles as the check that the rollback is +complete: an export missing Tables or with an unparseable file would also fail to restore. + +### What Step 1 cannot see + +- **Sets.** A Set is not in TML. Step 2 finds them in the copy's dependents. +- **Sharing and column security.** Not in TML. Handled in a later step. +- **Whether the exports are current.** Step 2's first call checks the files against the live + objects. + +### Implementation + +A fixed script shipped with the skill, `scripts/compare.mjs`, run with Node. The agent runs +it; it does not reason through the comparison itself, so the same inputs always give the +same result. + +### Open questions — Step 1 + +1. **Parsing YAML.** UI exports are zips of YAML. Proposed: pin the `yaml` npm package in + `scripts/package.json`. Alternatives: require JSON exports (REST, `edoc_format: JSON`), + or a hand-written parser (rejected — formulas and descriptions break naive parsers). +2. **Does a UI export include `fqn` on Model tables by default?** If not, the name fallback + in 1a is the common path, not the exception. +3. **Formula normalisation.** Confirm which reference forms appear in Model formulas + (`[Column]`, `[TABLE::column]`, …) so all of them are rewritten to bindings. +4. **Restoring a deleted copy.** Does re-importing the copy's TML after deletion recreate + it with the **same GUID**? If it gets a new GUID, restored dependents must be repointed + at it, and the rollback becomes a rewrite rather than a plain re-import. +5. **What restoring cannot bring back.** TML carries no sharing, and Sets are not in TML. + The copy's permissions must be snapshotted before deletion (later step), and a copy + carrying a Set is blocked in Step 2 anyway. +6. ~~**The copy's Tables.**~~ **Decided:** deleted after the copy Model, each only if it then + has zero dependents; a Table still in use is kept and reported. +7. ~~**Is the connection published with the Model?**~~ **Answered live:** publishing the + Model also publishes its Tables and connection. The connection is published but not + returned by `metadata/search` in the secondary Org, so a check for it must not rely on + that search. +8. ~~**The copy's connection after deletion.**~~ **Decided:** deleted last, after the copy's + Tables, only if no Table depends on it (so it survives until the last copy Model in the Org + is migrated), with the same identity check. Use `connections/{id}/delete` — + `metadata/delete` has no connection type. Restoring it needs the credentials re-entered; + the report says so before approval. + +## Step 2 — Live check and dependents + +The first step that uses the cluster, through `execute-thoughtspot-code` with +`org_identifier`. **Read-only**: every endpoint it calls is on the code-exec read allowlist +(`READ_OPERATIONS` in spotter-code `src/common/code-exec/security/policy.ts`), so no write +confirmation is ever requested and the step can be re-run freely. + +**Goal:** find everything that depends on the copy, decide what blocks, and turn Step 1's +provisional verdict into a final one. The copy is deleted at the end, so anything left +pointing at it would break — the dependent list must be complete, not approximately right. + +### 2a. Live checks + +| Check | Org | On failure | +|---|---|---| +| Session user is an administrator | Primary and secondary | Stop. Without it, reads return empty lists rather than errors, so "no dependents" would be silently wrong | +| Governed GUID exists and is a Model | Primary | Stop | +| Copy GUID exists, is a Model, and is **not** a published object | secondary | Stop | +| The files match the live objects: export the copy (JSON) in the sandbox, fingerprint it with Step 1's function, compare its hash with `plan.json` | secondary | Stop; ask for a fresh export. A stale file is also a stale rollback | + +The last check requires Step 1's fingerprint function to be plain JS that runs both +locally and in the sandbox — one function, used in both places. + +### 2b. Find the dependents + +```json +POST /api/rest/2.0/metadata/search +{ "metadata": [{ "type": "LOGICAL_TABLE", "identifier": "" }], + "include_dependent_objects": true, + "dependent_objects_record_size": -1, + "dependent_object_version": "V2" } +``` + +- `dependent_objects_record_size: -1` is required; the default truncates the list silently. +- If `dependent_objects` comes back as a **string** (with HTTP 200), the lookup failed. + Treat it as unknown and stop, never as "none". +- **Follow through Views.** A single lookup is one level deep: 200 Answers on 4 Views + appear as 4 dependents. Repeat the lookup on each View found (depth cap 4, skip GUIDs + already seen) and tag each object with `via_view`. That content needs no rewrite, but it + is what breaks if a View is repointed wrongly, so the report must show it. +- Run the same lookup on each of the **copy's Tables**. Content built directly on them is + not moved by repointing the Model. + +| Dependent | Phase 1 action | +|---|---| +| `ANSWER`, `LIVEBOARD` directly on the copy | Repoint (later step). A hidden visualization Answer belonging to a Liveboard is repointed through its Liveboard and listed once | +| View (`LOGICAL_TABLE`, subtype `AGGR_WORKSHEET`) directly on the copy | Repoint, **keeping the column names it exposes** (later step) | +| Anything reached through a View (`via_view` set), including stacked Views | No rewrite: the View it sits on keeps exposing the same names. Listed for visibility | +| `COHORT` (Set) | **Blocker, no override.** For each Set, look up its own GUID as `LOGICAL_COLUMN` to list the content that reads it | +| Content directly on the copy's Tables | Reported, not migrated. See Q6 | +| Any other type | Reported; **blocks deletion of the copy**, not the migration | + +A SQL View (`SQL_VIEW`) is built on a connection, not on a Model, so it never appears as a +dependent of the copy. + +### 2c. Which copy columns content actually uses + +Export the dependents that read the copy directly (Answers, Liveboards, Views) with one +batch `metadata/tml/export` call — up to about 40 per run to stay under the 50-call limit, +split beyond that. The full TML stays in the sandbox; only a compact row comes back: + +``` +{ guid, name, type, uses: ["Segment", "Revenue", …], tml_chars: 48210 } +``` + +- **Answers and Liveboards:** `uses` comes from the `[...]` tokens in each `search_query` + plus formula expressions, counting only visualizations whose `tables[]` point at the + copy. Names of the object's own formulas are excluded. +- **Views:** `uses` comes from the View's own `search_query` and + `view_columns[].search_output_column` — what it reads, not what it exposes. +- Content reached through a View is not scanned; it reads the View's names, which do not + change. +- `tml_chars` sizes the backup the repoint step must take (the 24k return limit applies + there). + +### 2d. The final verdict + +| Step 1 class | Used by a dependent | Not used | +|---|---|---| +| `MISSING` / `CHANGED` | **BLOCKED**, naming the dependents that use it | Warning: it disappears with the copy | +| `RENAMED` | That dependent needs a column rewrite | No effect | +| `MATCHED` | Only the Model reference changes | — | + +Report the effort, which is not the object count: + +``` +23 dependents: 5 need column rewrites, 3 need only the Model reference swapped, +15 sit on 2 Views and need nothing. 0 blocked. +``` + +### 2e. Record the current state + +Read-only, and needed before anything is written: + +- **The copy's sharing** — `security/metadata/fetch-permissions` on the copy (and its + Tables, if Q6 says they are deleted). TML carries no sharing and the copy will be + deleted, so this is the only record of who could access it. A later step re-applies it + to the published Model. +- **The dependents' sharing** — to verify later that it survived the repoint. They keep + their GUIDs, but that is to be confirmed. +- **Data fingerprints are not taken here.** They are taken at the start of the downtime + window (see 6b), so warehouse data has as little time as possible to change before the + repoint. + +### Output + +Added to `plan.json`: + +```json +"live_check": { "copy_matches_export": true, "admin_primary": true, "admin_secondary": true }, +"dependents": [{ "guid": "…", "type": "ANSWER", "name": "…", "via_view": null, + "uses": [], "needs_rewrite": true, "tml_chars": 0 }], +"views": [], "table_dependents": [], "sets": [], "unsupported": [], +"verdict": "RENAME", "blockers": [], +"sharing_snapshot": { "copy": [], "dependents": {} }, +"data_fingerprints": {} +``` + +### Runs + +| # | Org | Does | +|---|---|---| +| 1 | Primary | Administrator check; governed Model check | +| 2 | secondary | Administrator check; copy check; file-match check; dependents of the copy, through Views, and of its Tables | +| 3…n | secondary | Batch export of direct dependents (~40 per run); usage extraction | +| n+1 | secondary | Content reading each Set — only when a Set was found | +| n+2 | secondary | Sharing snapshot | + +### Open questions — Step 2 + +9. **Data fingerprints** — take them in Phase 1 (see 6b), or verify only that each object + resolves to the published Model? +10. **Dependent lookup through a View** — confirm `metadata/search` on a View's GUID + returns its dependents the same way as on a Model. +11. **Hidden visualization Answers** — confirm how V2 dependents report Answers embedded in + a Liveboard, so they collapse to the owning Liveboard. +12. **Admin check** — confirm which field of `auth/session/user` shows administrator + privilege in an Org-scoped session. + +## Step 3 — Publish the governed Model into the secondary Org + +The first step that writes. Request shapes below are from the REST spec shipped with +spotter-code (`src/common/rest-api-sdk-resolved-spec.json`). + +**Preconditions:** Step 2's verdict is not `BLOCKED`, and the customer has approved the +Step 1–2 report, including the column map, the findings, and the variable names proposed +below. + +**One sub-step, one run, one approval.** 3a, 3b and 3c each run as a separate +`execute-thoughtspot-code` call and are never combined, so a write approval +(`confirm_write_operations: true`) covers exactly one change, shown to the customer +beforehand. Each run reads the current state first and writes only what is missing, so a +re-run after a partial failure skips what is done, and a completed sub-step makes no write +and needs no approval. + +### 3a. Variables — point the published Tables at the copy's data (Primary Org) + +Publishing is refused unless something in the Model's dependency tree is parameterized. +More importantly, the variable values decide **what data the secondary Org reads**. The +target values are Step 1's `org_values`: the copy's `db` / `schema` / `db_table`. + +Step 1's fingerprint already shows which case applies; `template/variables/search` (read) +confirms it and reads the current values. + +| Case | Action | +|---|---| +| **A** — the governed Tables are already parameterized (usual when the Model is published to other Orgs) | Add only this Org's value: `template/variables/update-values`, `operation: ADD`, scoped to `org_identifier: ` | +| **A′** — this Org already has a value, and it differs from the copy's | **Stop and ask.** Someone configured it differently; overwriting would move the Org to other data | +| **B** — not parameterized | In this order: `template/variables/create` (`TABLE_MAPPING`, **no `data_type`** — refused for this type) → ADD the **Primary Org's value = the current literal** → ADD the secondary Org's value → `metadata/parameterize` each field | + +```json +POST /api/rest/2.0/template/variables/update-values +{ "variable_assignment": [{ "variable_identifier": "primaryconn_primary_data_schema", + "variable_values": ["ACME_PROD"], "operation": "ADD" }], + "variable_value_scope": [{ "org_identifier": "ACME" }] } +``` + +Rules: + +- **ADD only; never REPLACE or RESET.** A mis-scoped replace can wipe the values of Orgs + already on the Model. +- **In case B, values before parameterizing.** Otherwise the Primary Org briefly reads + through a variable with no value, and its own queries fail. +- **Check coverage yourself.** The publish check is an *existence* check: it passes if any + field in the tree is parameterized. After 3a, confirm that every Table whose values + differ between the two sides is parameterized and resolves to the copy's value for this + Org. A missed Table silently shows the secondary Org the Primary Org's data. +- If Step 1 found connection properties that differ, the Org also needs + `CONNECTION_PROPERTY` values (or the customer accepts the Primary Org's settings). How + those are assigned depends on Q7; 3a stops until it is answered. + +#### Variable naming + +**Case A — no naming.** The name is read from the governed Table TML +(`schema: ${apj_sales_schema}`); 3a only adds a value under it. + +**Case B — one variable per shared value, named so the result does not depend on the +order Models are migrated in.** + +Tables that share a value share a variable: both `SALES_FACT` and `PRODUCT_DIM` in +`PRIMARY_DATA` → one variable. The variable's identity is **(connection, field, current +Primary value)**, and the name is built from exactly that, always: + +| Rule | Example | +|---|---| +| Name = `{connection}_{current Primary value}_{field}`, slugified (lowercase, `_`) | `primaryconn_primary_data_schema` | +| **Reuse first.** A later Model whose Tables have the same connection, field and Primary value uses the existing variable | A second Model on `PRIMARY_DATA` → `primaryconn_primary_data_schema` | +| A different Primary value is a different variable | Tables on `PRIMARY_OTHER` → `primaryconn_primary_other_schema` | +| **Fallback:** Tables share the key, but their copy needs a **different** value in an Org that already has one | `primaryconn_primary_data_{model}_schema` | +| **Stop (case A′):** the *same* Table needs two values in one Org | — a data conflict, not a naming one | + +Recommended fields are `databaseName` and `schemaName`; `tableName` only when Step 1 shows +the table names actually differ. Names are unique across the cluster, not per Org. + +**The name is a label, never the decision.** + +- **Reuse is decided from cluster state**, not from the name: which variable the Tables are + already bound to (their TML), and each candidate variable's Primary value + (`template/variables/search`). A name can go stale — the Primary value is itself a + variable value and may change later. +- **A name that exists but does not match** — its Tables' connection, field or Primary value + differ — is treated as taken: add `_2`, `_3`. Slugging can collide (`PRIMARY-DATA` and + `PRIMARY_DATA` both become `primary_data`); never reuse on a name match alone. +- **The customer may override** the proposed name in the Step 1–2 report, e.g. an existing + standard such as `tenant_schema`. +- Name length and allowed characters are Q35. + +**Why always fold the value in.** `ts-publish-orgs` folds it only when one run sees more +than one value, so migrating Models one at a time would give the first Model a short name +and later ones long names, and `…_schema` would not say which schema it replaces. Tables +already parameterized by `ts-publish-orgs` keep their variable (case A); reuse is decided +by the key, so the two tools still interoperate. + +**Reusing a variable that has no value for this Org yet** adds one; every Table already on +that variable would then read that value in this Org if ever published there. The report +says so before 3a runs. + +The names proposed appear in the Step 1–2 report and are approved before 3a creates +anything. + +**Rollback:** REMOVE this Org's value. In case B, also unparameterize and delete the +variable — only if this run created it. + +### 3b. Re-key the copy's `obj_id`s (secondary Org — only if `needs_rekey`) + +The copy inherited the governed objects' `obj_id`s — the Model's and also its Tables' and +Connection's — and an `obj_id` must be unique within an Org. Every copy object in +`needs_rekey` whose governed counterpart is published into the Org (Q7) is re-keyed, in +one `update-obj-id` call; otherwise 3c is refused. + +```json +POST /api/rest/2.0/metadata/update-obj-id +{ "metadata": [{ "metadata_identifier": "", + "current_obj_id": "sales_model", + "new_obj_id": "sales_model__copy_acme" }] } +``` + +- The endpoint takes `metadata_identifier` **or** `current_obj_id`, not both, so the + stale-value guard is a read of each header before the write: stop if any `obj_id` differs + from Step 2's. +- A connection is re-keyed with `type: DATA_SOURCE` (verified live), in its own call. +- Re-key **the copy**, never the governed Model. The copy is deleted in Step 7, so the new + value only has to last until then. +- Before and after values go to the ledger. + +**Rollback:** set the previous value back — only after 3c is undone, or the clash returns. + +### 3c. Publish (Primary Org) + +```json +POST /api/rest/2.0/security/metadata/publish +{ "metadata": [{ "identifier": "", "type": "LOGICAL_TABLE" }], + "org_identifiers": [""] } +``` + +- **Never send `skip_validation`.** Publishing validates before writing and commits + all-or-nothing, so a refused publish leaves nothing changed. That validation is the + safety net. +- Whether the Tables must be listed or follow the Model depends on Q7. Until then the + script lists the Model and its Tables. +- Already published to this Org (a re-run) → nothing is sent. + +| Error | Cause | +|---|---| +| `Cannot publish/unpublish objects with Cohort Column as dependency` | A Set on the **governed** Model. Step 2 checked only the copy | +| `No template variable node found in the dependency tree` | 3a did not land | +| Duplicate custom object id | 3b skipped or failed | +| `Objects can only be published/unpublished from primary org` | The run was in the wrong Org | + +**Rollback:** `security/metadata/unpublish` from this Org. Possible only before Step 5 — +afterwards repointed content depends on the published Model and unpublishing is refused. + +### 3d. Verify the publish (secondary Org, read-only) + +**Query data; never trust the publish's `204` alone.** Found live: after a successful +publish, Sage's indexing lagged — its org-aware snapshot did not yet include the secondary +Org, so every search on the published Model failed there (`searchdata` 11020 *"Invalid +data source guid"*; UI Ref 10028; Sage log `PERMISSION_DENIED_INACCESSIBLE_TO_ORG`). It was +still failing at least 17 minutes after the publish, then resolved on its own. + +So 3d **waits and retries**: re-query the published Model every few minutes, up to a +limit (Q36). Only if it is still not searchable at the limit does 3d stop, with +"published, but not yet searchable in this Org — retry later or contact support". Steps +4–7 never start before 3d returns data: repointing would break every dependent. + +- The published Model is visible in the secondary Org, and its GUID is the governed GUID. +- **Each published Table resolves to the copy's physical location** — the same `db` / + `schema` / `db_table` Step 1 recorded. This is the tenant-isolation check: if the Org + reads exactly what the copy read, the migration has exposed no other data. How to read + the resolved values is Q13. +- The copy is present and unchanged apart from its `obj_id`. + +### Output + +A **ledger** in `plan.json`, one entry per write. It drives rollback in reverse order and +lets a re-run skip what is done. + +```json +"ledger": [ + { "step": "3a", "action": "variable_value_add", "variable": "primaryconn_primary_data_schema", + "org": "ACME", "value": "ACME_PROD" }, + { "step": "3b", "action": "update_obj_id", "guid": "…", + "before": "sales_model", "after": "sales_model__copy_acme" }, + { "step": "3c", "action": "publish", "guid": "…", "org": "ACME" } +] +``` + +### Open questions — Step 3 + +13. **Reading resolved values.** How to see what a published Table resolves to in a given + Org — `show_resolved_parameters` on `metadata/search`, or exporting the Table's TML in + that Org? +14. **`metadata/parameterize` field names.** The exact `field_name` values + (`databaseName`, `schemaName`, `tableName`?). +15. **ADD on an existing value.** With an `org_identifier` scope, does `update-values` ADD + append a second value when one exists? Decides whether reading first is enough to + detect case A′. + +35. **Variable name limits** — what length and characters does `template/variables/create` + accept? + +36. **Sage indexing lag after a publish** — how long can it take before a published Model is + searchable in the secondary Org? Sets 3d's retry limit. Observed: more than 17 minutes. + +## Step 4 — Give the published Model the copy's permissions + +Runs in the secondary Org, inside the downtime window, **before** any content is repointed, +so no user loses access at any point. Request shapes are from the REST spec shipped with +spotter-code. + +**Preconditions:** Step 3 done and 3d passed; no dependent repointed yet. + +### 4a. Read the current state (read-only) + +1. **The copy's explicit sharing** — `security/metadata/fetch-permissions` on the copy with + `permission_type: DEFINED` (10.3+). `DEFINED` returns only what was explicitly shared; + without it the response lists effective access, including every administrator and one + row per member of each shared group. This fresh read, taken at the start of the + downtime window, is what Step 4 grants from — not the Step 2 snapshot. +2. **Current sharing on the published Model and its Tables** in this Org, so access that + already exists is never lowered. +3. **Principals who see a dependent but not the copy** — from the dependents' sharing. Once + their content sits on the published Model, Strict Object Mode requires access to the + Model too. They are **not** granted automatically: Model access shows them far more + than one Answer. Listed; the customer decides. *(Proposed — pending decision.)* +4. **Column security on the copy** — see 4d. + +### 4b. What gets granted + +Every explicit entry on the copy is carried to the same principal. The copy and the +published Model are in the same Org, so every principal already exists there. + +| On the copy | On the published Model | Why | +|---|---|---| +| `READ_ONLY` | `READ_ONLY` | Same | +| `MODIFY` | `READ_ONLY`, disclosed in the report | A published object is edited from the Primary Org only; edit rights here would mean nothing (Q22) | +| `NO_ACCESS` (explicit) | `NO_ACCESS` | Keeps an exception, e.g. a user blocked inside a group that has access | +| The copy's author / owner | Listed, not shared by default | Authorship is not a share; usually it is the admin who ran the import | + +**Never lower access:** a principal that already has equal or higher access on the +published Model is skipped. + +### 4c. Grant (one run, one approval) + +Bottom-up. Under Strict Object Mode a grant on an object whose source is not shared returns +`HTTP 204` and is **silently dropped** (found live by `ts-migrate-orgs`): + +1. The published **Tables** — the same principals, `READ_ONLY`. +2. The published **Model** — the entries from 4b. + +```json +POST /api/rest/2.0/security/metadata/share +{ "metadata": [{ "type": "LOGICAL_TABLE", "identifier": "" }], + "permissions": [ + { "principal": { "type": "USER_GROUP", "identifier": "acme_analysts" }, "share_mode": "READ_ONLY" }, + { "principal": { "type": "USER", "identifier": "jdoe" }, "share_mode": "NO_ACCESS" } ] } +``` + +- `message` is **required** despite the spec (a share without it is refused, verified live): send `"message": ""` and no `emails`. Whether principals are notified still depends on their notify-on-share setting (Q23). +- The run is scoped to the secondary Org; it cannot change the published Model's sharing + in any other Org. Principals are per Org. + +### 4d. Column security on the copy — detect and stop + +The largest security risk in the migration. If the copy **hid columns** from a group — +column-level sharing on `LOGICAL_COLUMN`, or Column Security Rules +(`security/column/rules/fetch`) — the published Model carries no such restriction in this +Org. After the repoint the group would see columns it could not see before, with no error. + +- **Found → stop before Step 5.** The report names each column and principal. The customer + recreates the restriction on the published Model in this Org (the `ts-security-columns` + flow), or explicitly accepts the change, then re-runs Step 4. *(Proposed — pending + decision.)* +- **None found → continue.** + +Detection is cheap and read-only. Recreating the restriction automatically is Phase 2 +scope. + +### 4e. Verify (read-only) + +- `fetch-permissions` (`DEFINED`) on the published Tables and Model: every entry from 4b is + present at the expected level. +- Read the **`permission`** field, not `shared_permission` — the latter stays `NO_ACCESS` + on a successful share. `HTTP 204` alone does not prove a grant landed. +- Where possible, check as a **non-admin** member of a granted group; an admin sees objects + regardless of sharing. Repeated in Step 6. + +### Output + +Ledger entries, one per principal and object: + +```json +{ "step": "4c", "action": "share", "object": "", + "principal": "acme_analysts", "before": "NONE", "after": "READ_ONLY" } +``` + +**Rollback:** for each entry this run added, restore its `before` value (`NO_ACCESS` where +there was none). Access that already existed is never touched. + +### Open questions — Step 4 + +16. **Principals who see only a dependent** — ask the customer (proposed), or grant them + `READ_ONLY` on the Model automatically? +17. **Column security on the copy** — stop until resolved (proposed), or warn only? +18. **Enforcing downtime for one Org** — see "Running the migration — downtime". +19. **Last-modified time** — which `metadata/search` header field it is, and whether every + save changes it, including a Liveboard layout change. +20. **`force-logout` and embedded sessions** — does it end trusted-auth sessions, or only + UI sessions? +21. **Strict Object Mode** — on for this cluster? Bottom-up granting is correct either way; + it decides whether 4a item 3 matters. +22. **`MODIFY` on a published Model** — can it be granted in a secondary Org, and does it + do anything? +23. **Notifications** — does `share` without `emails` still notify the principals? +24. **Column-level sharing on the copy** — is it read with `fetch-permissions` on + `LOGICAL_COLUMN`, and can one call cover all columns? + +## Step 5 — Repoint dependents + +Runs in the secondary Org, inside the downtime window. The rewrite rules are those of +`ts-migrate-orgs` (thoughtspot-agent-skills `tools/ts-cli/ts_cli/migrate/rewrite.py`), +which were proven against real content; this step runs them through +`execute-thoughtspot-code` instead of a local CLI. + +**Preconditions:** Step 4 done and 4e passed; the start-of-window re-checks passed; no +column security on the copy, or the customer accepted the change. + +**Scope**, from the Step 2 list as refreshed at the start of the window, in this order: + +1. **Views directly on the copy** — first, because everything on them depends on them. +2. **Answers and Liveboards directly on the copy.** A hidden visualization Answer is handled + through its Liveboard, once. + +Content reached through a View (`via_view`) is **not touched**. + +### 5a. Back up each object (read-only) + +Per object: + +1. Record its **last-modified time `t0`**. +2. Export its TML (`export_fqn: true`) and compute a **SHA-256** of it in the sandbox. +3. Return the TML to the agent — **in slices** when it exceeds the 24,000-character return + limit (Step 2's `tml_chars` gives the size in advance). +4. The agent joins the slices, checks the hash, and writes `backup/.tml`. + +**Every backup exists before the first write.** A hash mismatch or a failed write stops +Step 5: that file is the only way back for the object. + +### 5b. The rewrite + +**1. Swap the Model reference.** In `tables[]`, only entries whose `fqn` is the copy's GUID +change: `fqn` → the governed GUID, and `name` / `id` → the governed name. A stale `name` +would let resolution fall back to a name match and bind to the wrong object if the `fqn` +ever failed to resolve. Entries for **other** sources are left alone — rebinding them +imports cleanly and renders wrong. + +**2. Rename columns by the column map** (skipped when the verdict is `READY`: the map is +identity). + +- **Denylist, not allowlist.** Every string is rewritten except known label fields + (`LABEL_PATHS`: visualization titles, `answer.name`, filter display names, chart column + labels). A reference field added by a future release is then covered automatically; the + label list is the one thing to review when the platform changes. +- **Three reference forms:** bare `[Segment]`, qualified `Sales::Segment`, and decorated + `Total Segment` inside `search_output_column`. +- **`client_state_v2`** — a JSON string holding chart state — is **parsed** and rewritten + field by field, never by substring replacement, which corrupts unrelated state. + +**3. Views.** Rewrite what the View **reads** (`search_query`, formulas, +`search_output_column`). **Preserve exactly** what it **exposes**: `view_columns[].name` and +the View's own name. Content on the View then needs no rewrite. + +**4. Column rewrites are scoped to visualizations that read the copy.** `ts-migrate-orgs` +applies the map document-wide. Here it applies only inside visualizations whose `tables[]` +point at the copy: on a Liveboard that also shows another Model, a visualization on that +Model with its own `Segment` column must not be renamed. + +**A Liveboard filter that spans the copy and another source, on a renamed column, stops +that object for manual handling.** It cannot be rewritten for one source without breaking +the other, and a wrong rewrite filters the other Model's visualizations by the wrong column +with no error. Such Liveboards are expected to be rare; the object stays on the copy (still +working) and is listed in the report. + +### 5c. Checks before any import (inside the write run) + +| Check | On failure | +|---|---| +| **Last-modified time is still `t0`** | Stop Step 5 — the downtime leaked | +| **Coverage:** no copy column name survives outside label fields, and no `fqn` still points at the copy | Do not import; report the paths. **Never work around it** — a partial rewrite imports cleanly and shows wrong numbers | +| **The map is injective**, and no target name collides with the object's own formulas | Do not import | +| **`tml/import` with `import_policy: VALIDATE_ONLY`** | Do not import; report the platform's error | + +### 5d. Import and confirm (write) + +- `metadata/tml/import`, `import_policy: ALL_OR_NONE`, `create_new: false`, over the same + GUID. The GUID is unchanged, so sharing, schedules and favourites are expected to stay + (Q29). +- **Confirm in the same run:** re-export and check every former copy reference now points + at the governed GUID. +- Ledger: `MOVED`, with the backup path and hash. + +### Batches and approvals + +A write run costs about five calls per object (modified-time read, export, validate, +import, confirm). Within the 50-call and 150-second limits that is **up to about eight +objects per run**. + +- **One approval per batch.** The customer sees the batch's objects and, for each, whether + it needs a column rewrite or only the Model reference swapped. +- Views go in their own batch(es), first. +- **Stop at the first failure.** Objects already moved work on the published Model; the + rest still work on the copy. Both are safe states. Investigate, then re-run — the ledger + skips what is done. + +### What a TML round trip may not keep + +Conditional formatting, column widths, some chart state and Liveboard tile layout are not +guaranteed to survive a re-import (found by `ts-migrate-orgs`). Which columns an object +reads, what it computes and what it filters do survive. + +- **Disclosed per object** in the report before the batch is approved. +- **Not a reason to refuse** — leaving content on a copy about to be deleted is worse. +- Step 6 checks that the **numbers** match, not that the documents match. + +### Rollback + +Re-import `backup/.tml` over the same GUID. Straightforward while the copy exists +(before Step 7); after the copy is deleted it depends on Q4. + +### Output + +```json +{ "step": "5d", "action": "repoint", "guid": "…", "type": "LIVEBOARD", + "backup": "backup/.tml", "sha256": "…", "column_rewrite": true, "t0": "…" } +``` + +Objects stopped for manual handling are recorded as `MANUAL`, with the reason. + +### Open questions — Step 5 + +25. **Import format and mode** — does `tml/import` accept the JSON export, and update in + place with `create_new: false`? +26. **`tables[].name` / `id`** — must they equal the governed name, or is `fqn` enough? +27. **`VALIDATE_ONLY`** — does it catch a reference to a column that does not exist, or only + malformed TML? +28. **`LABEL_PATHS`** — is the `ts-migrate-orgs` list still complete on the current release? +29. **In-place import** — do sharing, schedules, favourites and embed links survive? +30. **Hashing in the sandbox** — is `crypto.subtle` available in the code-exec engine? + +## Step 6 — Verify + +Read-only, in the secondary Org, inside the downtime window. **Step 6 is the gate for +Step 7: the copy is deleted only after it passes.** + +### 6a. Everything resolves to the right Model + +- **The copy's dependents** — re-run the lookup (2b). Only objects recorded as `MANUAL` in + Step 5 may remain. Anything else is a failure: missed, or created during the window. +- **Every `MOVED` object** appears as a dependent of the **published** Model in this Org, and + its re-exported TML holds no reference to the copy. +- **Views** — each repointed View depends on the published Model, and content on it still + resolves through it. + +### 6b. The numbers match + +The "before" fingerprints are taken **at the start of the downtime window**, with the other +start-of-window checks — as late as possible, so warehouse data has little time to change. +Whether fingerprints are taken at all is Q9. + +- **Answers** — `metadata/answer/data`: row count, plus a hash of the first N rows **sorted + first**, since row order is not guaranteed. +- **Liveboards** — `metadata/liveboard/data` per visualization, batched within the 50-call + and 150-second limits. +- **Content on Views** — sampled; it was not rewritten. +- **A mismatch is not rolled back automatically.** It is marked `REVIEW` and blocks Step 7. + The customer compares, then accepts it or rolls that object back from its backup. + +Administrators bypass RLS, but both readings come from the same admin session, so the +comparison is like for like. It proves the numbers did not change — **not** what a tenant +user sees. That is 6d. + +### 6c. Access is right + +- `fetch-permissions` (`DEFINED`) on the published Model and Tables matches what Step 4 + granted. Read `permission`, not `shared_permission`. +- **Each dependent's sharing** is unchanged from the start-of-window reading. This answers + Q29 on real data. + +### 6d. As a real tenant user (manual) + +The skill cannot do this: code-exec runs as the connected admin, and acting as another user +would mean minting a token (a gated write that needs trusted authentication). The report +ends with a checklist the customer completes, logged in as a **non-admin** member of the +Org: + +1. Open two or three moved Liveboards and Answers — they load and show data. +2. Row counts match what that user saw before — RLS still applies. +3. Columns hidden from them (if any were recreated after 4d) are still hidden. +4. They still see the copy as a second Model in the data picker — expected until Step 7. + +The customer confirms the checklist before Step 7. + +### Result + +| Result | Meaning | Next | +|---|---|---| +| **PASS** | All checks pass; customer confirms 6d | Step 7 may run | +| **REVIEW** | Number mismatches, or `MANUAL` objects remaining | Step 7 blocked until each item is accepted or resolved | +| **FAIL** | Wrong Model reference, a missed dependent, or grants missing | Step 7 blocked; fix, or roll back the affected objects | + +Recorded in `plan.json` as `verification`, with a status per object. + +### Column names change for users (Phase 1: disclosed) + +With a `RENAME` verdict, tenant users **see the governed Model's column names** after the +repoint — `Segment` becomes `STRING_1` in their Answers, search and Spotter. The data is +right; the labels change. + +`ts-migrate-orgs` restores the tenant's names with **per-Org column aliases** on the +Primary Org's Model (`ts alias`, once per wave, with an `--expect-org` guard so aliases for +Orgs already cut over are not wiped). That writes to the governed Model in the **Primary** +Org and affects every Org, and it is the one step that tool calls catastrophic if done +wrong. + +**Phase 1 discloses instead:** the Step 1–2 report lists every renamed column as "users will +now see X instead of Y", so the customer knows before approving. Aliases are a later phase. +Copies whose names match the governed Model (`READY`) are unaffected. + +### Open questions — Step 6 + +31. **Large Liveboards** — how does `liveboard/data` behave within the limits; is sampling + visualizations enough? +32. **Aliases on published Models** — do per-Org aliases render in Answers on the published + Model in a secondary Org? Needed only when aliases are added. + +## Step 7 — Delete the copy + +The only step a re-import of an object's backup cannot undo, so it checks the most before +acting. Secondary Org. + +**Preconditions:** Step 6 is **PASS**, or every `REVIEW` item has been accepted; the +customer confirmed the 6d checklist; **no `MANUAL` objects remain** — each still depends on +the copy and would break. + +### 7a. Final sync + +| Check | On failure | +|---|---| +| **The copy has zero dependents** (lookup 2b) | Do not delete; list what remains | +| **The copy is unchanged since the archive** — re-fingerprint and compare with `inputs/` | Stop; the rollback file no longer matches | +| **The copy's sharing, read again** — any share added since Step 4 | Apply it to the published Model first (a separate `share` run and approval, logged like Step 4), then continue | +| **The rollback set is complete** — the copy's TML, every `MOVED` backup (hashes checked), both sharing snapshots, the ledger | Stop; never delete without a complete way back | + +Some uses of the copy are not dependents: embed code or scripts using its GUID or `obj_id`, +and Spotter conversations on it. The customer confirmed the `obj_id` case at 3b and +confirms the rest here. + +### 7b. Delete (write — its own approval, never combined with other writes) + +```json +POST /api/rest/2.0/metadata/delete +{ "metadata": [{ "type": "LOGICAL_TABLE", "identifier": "" }] } +``` + +**Identity check in the same run, immediately before the call.** Since Step 3 the governed +Model is visible in this Org too, so a wrong GUID could target the published Model. The +script re-reads the object and deletes only if **all** hold: + +- the GUID equals `plan.copy.guid` and is **not** `plan.governed.guid`; +- it is **not** a published object; +- its `obj_id` is the value set in 3b (e.g. `sales_model__copy_acme`). + +No `delete_disabled_objects`. Exactly one object per call. + +**The copy's Tables (Q6 — decided)** exist only for the copy, so they are deleted **after the +copy Model**, each one only if it then has **zero dependents**. A Table something else still +reads — content built on it directly, another Model — is kept and listed in the report with +its dependents. Each Table gets the same identity check as the Model (the GUID is the +copy's Table, it is not a published object, its `obj_id` is the 3b value). The copy's +connection follows the same rule, last (Q8). + +### 7c. Confirm (read-only) + +- A search for the copy's GUID returns nothing. +- The published Model is present in the Org, and its dependents match Step 6. +- A spot re-run of 6a on a sample of moved objects. + +The downtime window then ends and the customer re-opens access. + +### Timing + +Default: delete in the same downtime window. The customer may **defer** it — for example to +observe for a week. That is safe because 7a re-checks everything when it runs, including +content built on the copy in the meantime; the cost is users seeing two Models until then. + +### Full rollback after deletion + +Only the files and the ledger remain. In a downtime window, in this order: + +1. **Recreate the copy** from `inputs/`, with the **re-keyed `obj_id`** — the original + clashes with the published Model, still in the Org. Whether it gets its original GUID + back is Q4 / Q34. +2. **Restore the copy's sharing** from the snapshot; TML carries none. +3. **Re-import each Step 5 backup.** They reference the copy's GUID; if the GUID changed, + rewrite that reference first. +4. **Undo Step 4** — restore each `before` value in the ledger. +5. **Unpublish** from this Org — now possible, as nothing depends on the published Model. +6. **Restore the copy's `obj_id`** (undo 3b), then **remove this Org's variable value** + (undo 3a). + +Phase 1 documents this as a runbook, not an automated command. + +### Output + +```json +{ "step": "7b", "action": "delete", "guid": "", "obj_id": "sales_model__copy_acme", + "rollback_set": { "copy_tml": "inputs/…", "backups": 23, "sharing_snapshot": "…" } } +``` + +The final report: moved / `MANUAL` objects, grants, the copy deleted, the copy's Tables and +connection left in place, and where the rollback set is. + +### Open questions — Step 7 + +33. **Delete with dependents** — does `metadata/delete` refuse when dependents exist, or + remove them with the copy? A second safety net if it refuses; critical to know if it + cascades. +34. **Restore with a different `obj_id`** — can the copy be re-imported with its original + GUID but the re-keyed `obj_id`? Pairs with Q4. diff --git a/docs/migrate-tml-copies-test-scenarios.md b/docs/migrate-tml-copies-test-scenarios.md new file mode 100644 index 0000000..bc872f5 --- /dev/null +++ b/docs/migrate-tml-copies-test-scenarios.md @@ -0,0 +1,274 @@ +# Test scenarios: migrate TML Model copies to Orgs Publishing + +Companion to [migrate-tml-copies-design.md](migrate-tml-copies-design.md). Ordered from +basic to complicated: each level assumes the one before it passes. `Q` numbers refer to the +design's open questions; a scenario lists the ones it answers. + +Status: **In progress** — basic scenario (1.2 + 1.4) **passed** 2026-09-28; see Results. + +## Fixtures + +Built once and reused. Every scenario starts from a fresh copy of these unless it says +otherwise. + +| Fixture | Contents | +|---|---| +| **Warehouse** | Database `TS_MIGRATION_DEMO`, three schemas with the same two tables, `SALES_FACT` (fact) and `PRODUCT_DIM` (dimension): `PRIMARY_DATA` (read by the Primary Org), `ORG1_DATA` (read by org1), `SHARED_DATA` (read by both — Level 3 shared-data scenarios). The schemas must hold **different rows**, so reading the wrong schema shows as wrong numbers | +| **Primary Org** | Connection; Tables on `PRIMARY_DATA`; governed Model **`PRIMARYmodel`** (GUID `faea95b6-6006-4a63-ae38-5937339d1c1f`, `obj_id` `PRIMARYmodel-faea95b6`, on `nebula-test-27sep`), `SALES_FACT` → `PRODUCT_DIM`, 12 columns (below), with an `obj_id`. **Add one formula**, `Margin = [Revenue] - [Cost]`, for the formula scenarios | +| **Secondary Org `org1`** | Connection; the copy of `PRIMARYmodel`, made by exporting from Primary and importing with the schema changed to `ORG1_DATA` | +| **Principals in org1** | Group `org1_analysts` with non-admin member `org1_user`; a second non-admin user `org1_viewer`, **not** in the group; the migrating admin | +| **Exports** | Governed and copy TML exported from the UI **with dependencies**, and again via REST with `export_fqn: true` | + +**`PRIMARYmodel` columns** + +| Column | Source | Type | +|---|---|---| +| Sale Id | `SALES_FACT.SALE_ID` | MEASURE (SUM) | +| Product Id | `SALES_FACT.PRODUCT_ID` | MEASURE (SUM) | +| Sale Date | `SALES_FACT.SALE_DATE` | ATTRIBUTE (DATE) | +| Region | `SALES_FACT.REGION` | ATTRIBUTE | +| Sales Region | `SALES_FACT.SALES_REGION` | ATTRIBUTE | +| Segment | `SALES_FACT.SEGMENT` | ATTRIBUTE | +| Customer Email | `SALES_FACT.CUSTOMER_EMAIL` | ATTRIBUTE — the sensitive column for Level 5 | +| Tenant Code | `SALES_FACT.TENANT_CODE` | ATTRIBUTE — the RLS column | +| Revenue | `SALES_FACT.REVENUE` | MEASURE (SUM) | +| Cost | `SALES_FACT.COST` | MEASURE (SUM) | +| Product Name | `PRODUCT_DIM.PRODUCT_NAME` | ATTRIBUTE | +| Category | `PRODUCT_DIM.CATEGORY` | ATTRIBUTE | + +**Renames used from Level 2 on:** in the copy, `SEGMENT` is named **`Customer Segment`** +and `REVENUE` is named **`Sales Amount`**. The expected column map is +`Customer Segment → Segment`, `Sales Amount → Revenue`. + +## Level 1 — Happy path, smallest possible + +Goal: prove the seven steps work end to end before adding any variation. + +| # | Scenario | Setup | Expected | Answers | +|---|---|---|---|---| +| 1.1 | **Copy with no dependents** | Fixtures only | Step 1 `READY`; Step 2 zero dependents; publish; grants; nothing to repoint; delete | Q1, Q2, Q7, Q13, Q14, Q33 (no deps) | +| 1.2 | **One Answer on the copy** | + Answer on the copy, shared READ_ONLY to `org1_analysts` | Repointed in place; same GUID; same numbers; `org1_user` still sees it | Q25, Q26, Q27, Q29, Q30 | +| 1.3 | **One Liveboard** (2 visualizations, 1 Liveboard filter) | + Liveboard on the copy | Repointed as one object; filter works; layout disclosed | Q11, Q31 | +| 1.4 | **Copy shared to a group** | Copy shared READ_ONLY to `org1_analysts` | Published Model gets the same grant; `org1_user` can search it | Q21, Q23 | + +## Level 2 — Column differences + +Goal: the comparison and the rewrite. Uses 1.2 / 1.3 content. + +| # | Scenario | Setup | Expected | Answers | +|---|---|---|---|---| +| 2.1 | **Renamed columns** | Copy uses `Customer Segment` and `Sales Amount`; an Answer and a Liveboard use both | `RENAME`; column map of 2; content rewritten; same numbers; users now see governed names (disclosed) | Q28 | +| 2.2 | **Renamed column inside a formula** | Answer formula `sum([Sales Amount]) / count([Sale Id])` | Formula rewritten; coverage gate clean | Q3 | +| 2.3 | **Qualified and decorated references** | Liveboard filter `::Customer Segment`; `search_output_column` `Total Sales Amount` | Both forms rewritten | Q28 | +| 2.4 | **Chart state** | Chart with custom series colours and column properties on a renamed column | `client_state_v2` rewritten field by field; colours kept | — | +| 2.5 | **Label that matches a column name** | Visualization titled `Customer Segment` | Title **not** renamed | Q28 | +| 2.6 | **Missing column, unused** | Copy has an extra column no content uses | Warning only; migration proceeds | — | +| 2.7 | **Missing column, used** | Answer uses a copy-only column | `BLOCKED`, naming the Answer; nothing written | — | +| 2.8 | **Changed formula** | Copy's `Margin` is `[Sales Amount] - [Cost] * 1.1`; an Answer uses it | `BLOCKED` | Q3 | +| 2.10 | **Swapped names** | Copy names `SALES_REGION` **`Region`** and `REGION` **`Area`**. Map: `Region → Sales Region`, `Area → Region` | Renames applied **simultaneously**, not one after another — a sequential pass turns `Area` into `Region` and then into `Sales Region`. Same numbers per region | Q28 | +| 2.9 | **Extra columns in governed** | Governed has columns the copy lacks | Informational only | — | + +## Level 3 — Model and publishing variations + +| # | Scenario | Setup | Expected | Answers | +|---|---|---|---|---| +| 3.1 | **Copy has a different name** | Copy named `PRIMARYmodel – org1` | Paired by inputs, not name; `tables[].name` updated on repoint | Q26 | +| 3.2 | **Copy without the inherited `obj_id`** | Clear the copy's `obj_id` | `needs_rekey` false; 3b skipped; publish succeeds | — | +| 3.3 | **Case B — governed not parameterized** | Fresh governed Model | Variable created with the naming convention; Primary value set first; **Primary queries unaffected** | Q14, Q15 | +| 3.4 | **Case A — already published to another Org** | Governed already published to Org `org2` | Only org1's value added; org2 unaffected | Q15 | +| 3.5 | **Case A′ — conflicting value** | org1 already has a different value for the variable | Stop and ask; nothing written | Q15 | +| 3.6 | **Different table names** | Copy reads `ORG1_DATA.SALES_FACT_ORG1` instead of `SALES_FACT` *(needs a renamed copy of the table)* | `tableName` parameterized | Q14 | +| 3.7 | **Shared dimension** | Governed: `SALES_FACT` in `PRIMARY_DATA`, `PRODUCT_DIM` in `SHARED_DATA`. Copy: `SALES_FACT` in `ORG1_DATA`, `PRODUCT_DIM` in `SHARED_DATA` | Only the fact's schema differs → one variable for the fact; the dimension stays unparameterized and both Orgs read `SHARED_DATA` | Q13 | +| 3.8 | **Split variable needed** | Governed: both tables in `PRIMARY_DATA`. Copy: `SALES_FACT` in `ORG1_DATA`, `PRODUCT_DIM` in `SHARED_DATA` | One governed schema value maps to two copy values → case B: split into two variables; case A: stop | — | +| 3.8a | **Copy reads the same data as Primary** | Governed and copy both on `SHARED_DATA` | No values differ; reported as "org1 reads the same data as Primary, as the copy did" — no new exposure. Publish still needs a variable | Q13 | +| 3.9 | **Joins differ** | Copy changes a join's `on` clause; separately, same join written from the other side | First is a finding; second is **not** (cardinality normalised) | — | +| 3.10 | **Role-playing dimension** | `SALES_FACT` joins a date table twice (order date, ship date) *(needs a `DATE_DIM` table and two date keys)* | Stop at Step 1 | — | +| 3.11 | **RLS differs** | RLS rule on `TENANT_CODE` on the copy's `SALES_FACT`, not on governed | Security finding in the report | — | +| 3.12 | **Different warehouse account** | Copy's connection points at another account | 3a stops until connection-property handling is known | Q7, Q8 | + +## Level 4 — Dependent topology + +| # | Scenario | Setup | Expected | Answers | +|---|---|---|---|---| +| 4.1 | **View on the copy** | View on the copy + 2 Answers on the View | View repointed with exposed names kept; Answers untouched and still return data | Q10 | +| 4.2 | **Stacked Views** | View on a View on the copy | Only the bottom View rewritten; all content still works | Q10 | +| 4.3 | **Content directly on the copy's table** | Answer on the copy's `SALES_FACT` table | Reported, not migrated; copy's tables kept | Q6 | +| 4.4 | **Mixed Liveboard** | Liveboard with visualizations on the copy and on another Model that also has a `Customer Segment` column | Only the copy's visualizations rewritten | — | +| 4.5 | **Mixed Liveboard with a shared filter on a renamed column** | 4.4 + Liveboard filter on `Customer Segment` across both | Liveboard marked `MANUAL`; copy **not** deleted | — | +| 4.6 | **Set on the copy** | Set on the copy used by one Answer | `BLOCKED`; the Answer named; nothing written | — | +| 4.7 | **Set on the governed Model** | Set on governed | Publish refused with the Cohort error; nothing changed | — | +| 4.8 | **Large Liveboard** | Liveboard TML > 24,000 characters | Backup returned in slices; hash matches | Q30 | +| 4.9 | **Many dependents** | 50+ Answers on the copy | Batched export (~40 per run) and repoint (~8 per run) | — | + +## Level 5 — Permissions and security + +| # | Scenario | Setup | Expected | Answers | +|---|---|---|---|---| +| 5.1 | **Mixed grants** | Copy: `org1_analysts` READ_ONLY, explicit NO_ACCESS for `org1_user` (inside the group), `org1_viewer` MODIFY (user-level) | Published Model: group READ_ONLY, `org1_user` NO_ACCESS, `org1_viewer` READ_ONLY (downgrade disclosed). `org1_user` cannot see the Model despite the group | Q22 | +| 5.2 | **User sees an Answer, not the copy** | Copy shared to `org1_analysts`; the Answer also shared to `org1_viewer`, who is outside the group | Listed for the customer; with Strict Object Mode on, check what `org1_viewer` sees after the repoint | Q16, Q21 | +| 5.3 | **Access not lowered** | Published Model already MODIFY for a principal | Not lowered | — | +| 5.4 | **Column-level sharing on the copy** | `Customer Email` hidden from `org1_analysts` | Stop before Step 5 | Q17, Q24 | +| 5.5 | **Column Security Rule on the copy** | CSR on `Customer Email` on the copy | Stop before Step 5 | Q17 | +| 5.6 | **Non-admin verification** | Log in as `org1_user` after Step 5 | Content opens; org1 rows only | — | +| 5.7 | **Dependents' sharing survives** | Answer shared to `org1_viewer` before repoint | Same sharing after | Q29 | + +## Level 6 — Wrong inputs, failures, rollback + +| # | Scenario | Setup | Expected | Answers | +|---|---|---|---|---| +| 6.1 | **Both exports from the same Org** | Pass the governed export twice | Stop at 1a (equal GUIDs) | — | +| 6.2 | **Export without tables** | Model-only export | Stop at 1a | — | +| 6.3 | **Unparseable file** | Corrupt one file in the zip | Stop, naming the file | Q1 | +| 6.4 | **Stale export** | Edit the copy after exporting | Stop at 2a | — | +| 6.5 | **Admin in Primary only** | Migrating user not admin in org1 | Stop at 2a, not "zero dependents" | Q12 | +| 6.6 | **Edit during the window** | Edit an Answer between backup and import | Modified-time check stops Step 5 | Q19 | +| 6.7 | **Import fails mid-batch** | Make one object's import fail | Batch stops; earlier objects moved, later on copy; re-run resumes from the ledger | — | +| 6.8 | **Re-run everything** | Run the full migration twice | Second run writes nothing | — | +| 6.9 | **Rollback after Step 3** | Stop after publish | Unpublish, restore `obj_id`, remove value; estate as before | — | +| 6.10 | **Rollback after partial Step 5** | Stop after half the objects | Re-import backups; all on copy again | — | +| 6.11 | **Full rollback after Step 7** | Complete, then run the runbook | Copy restored with re-keyed `obj_id`; content back on it | Q4, Q34 | +| 6.12 | **Delete with dependents** (throwaway objects) | Call delete on a Model that still has an Answer | Record whether it refuses or cascades | Q33 | + +## Level 7 — Several Orgs and downtime + +| # | Scenario | Setup | Expected | Answers | +|---|---|---|---|---| +| 7.1 | **Two secondary Orgs in sequence** | Migrate org1, then org2 onto the same governed Model | org2 uses case A; org1 unaffected throughout | — | +| 7.2 | **Primary unaffected** | Run Primary content before and after 3.3 | Same numbers in Primary | — | +| 7.3 | **Downtime enforcement** | Try the proposed mechanism: block logins, `force-logout` an explicit list | Sessions ended for the listed users only, UI and embedded | Q18, Q20 | + +## First test to run + +**1.2 by hand**, before any script exists: one Answer on an identical copy. It exercises +every step once with the fewest moving parts, and answers the questions most likely to +change the design (Q2, Q7, Q14, Q25, Q26, Q29). Run each call through +`execute-thoughtspot-code`; `org_identifier` needs the `org-aware-code-exec` branch +deployed, otherwise use a session logged into each Org. + +## Results + +### 1.2 — Step 1 (2026-09-28, `nebula-test-27sep`, exports in `~/Documents/Migration_tool_demo`) + +| Check | Result | +|---|---| +| Parse; manifests `OK` | Pass | +| One Model + Tables + Connection per export | Pass | +| Model GUIDs differ | Pass — `faea95b6-…` vs `fae510f4-…` (same first three characters; compare whole GUIDs) | +| Model tables carry `fqn` | **Yes (Q2)** — UI exports reference Tables by GUID | +| Table pairing | `SALES_FACT`, `PRODUCT_DIM`; overlap 1.0 | +| Differences | `schema` only: `PRIMARY_DATA` → `ORG1_DATA`, both Tables. Governed not parameterized (case B) | +| Connection | Names differ (`PRIMARYconn` / `ORG1conn`) — informational. Type `RDBMS_SNOWFLAKE`, account, user, role, warehouse, `selected_databases` identical → published connection reaches `ORG1_DATA` (role grant confirmed at 3d) | +| Columns | 12 + `Margin` matched; identity column map | +| Verdict | **READY** | +| `obj_id` | Model, both Tables and the Connection all inherited → re-key all (design updated) | +| Formula | Governed `Margin` has a trailing space the copy lacks — must normalise whitespace (design updated; Q3 partly answered) | + +### Step 2 first pass (2026-09-28, org1 via `/bearer/mcp`) + +| Check | Result | +|---|---| +| Sessions | `spottercode-org1` → org1 (2137347761), `spottercode-primary` → Primary (0); `tsadmin`, admin in both | +| Copy / governed `obj_id` | Both `PRIMARYmodel-faea95b6` — confirms Step 1 | +| Copy's dependents | **None** | +| Copy's Tables' dependents | Only the copy itself — the Table lookup must exclude the copy from "content on the copy's Tables" | +| Copy's sharing (`DEFINED`) | **None** | +| Fixtures missing on this cluster | `org1_analysts`, `org1_user`, `org1_viewer` | + +### Basic scenario (1.2 + 1.4) — Step 2 (2026-09-28, org1) + +| Check | Result | +|---|---| +| Copy | Model (`WORKSHEET`), not a published object, `obj_id` `PRIMARYmodel-faea95b6`, `modified` present in the header (ms epoch — Q19 partly) | +| File match | Live export agrees with the Step 1 file: 13 columns, same `Margin` formula, both Tables on `ORG1_DATA`. The REST JSON export omits `obj_id` and Connection files — the shared fingerprint must use only fields present in both | +| Dependents | 1: Answer **Revenue by Region** (`7e60f180-992d-4c86-82cd-a787552d328c`). V2 lookup groups it under **`QUESTION_ANSWER_BOOK`**, not `ANSWER` — the classifier must map internal type names | +| Usage | `search_query` `[Revenue] [Margin] by [Region]` → `Revenue`, `Margin`, `Region`; all `MATCHED` → only the Model reference changes | +| Answer TML | 5,831 characters — backup fits in one return | +| Sharing (`DEFINED`) | Copy: **`org1_user` (USER) READ_ONLY** — not the group. Answer: `org1_analysts` (GROUP) READ_ONLY. Author of both: `tsadmin` | +| Data fingerprint | 2 rows — EAST 236036 / 123268, WEST 259223 / 128171 (Revenue / Margin); SHA-256 `0b07f697…a22` | +| `crypto.subtle` in the sandbox | **Works (Q30)** | +| Verdict | **READY**; 1 object, Model reference swap only | + +### Basic scenario — Step 3b (2026-09-28, org1, write) + +| Object | GUID | Before | After | +|---|---|---|---| +| Copy Model | `fae510f4-…` | `PRIMARYmodel-faea95b6` | `PRIMARYmodel-faea95b6__copy_org1` | +| Copy `SALES_FACT` | `25154482-…` | `SALES_FACT-40639b92` | `SALES_FACT-40639b92__copy_org1` | +| Copy `PRODUCT_DIM` | `bbbb1ecd-…` | `PRODUCT_DIM-9f6ea9ff` | `PRODUCT_DIM-9f6ea9ff__copy_org1` | +| Copy connection `ORG1conn` | `679586c3-…` | `PRIMARYconn-deaca893` | `PRIMARYconn-deaca893__copy_org1` | + +Both `update-obj-id` calls returned `204`; read-back confirmed. **`type: DATA_SOURCE` re-keys a connection.** +The endpoint takes `metadata_identifier` **or** `current_obj_id`, not both — the stale-value guard is a +read before the write. Rollback: the same calls with the old values. + +### Basic scenario — Step 3a (2026-09-28, Primary, write) + +| Action | Result | +|---|---| +| `variables/create` `TABLE_MAPPING` **with** `data_type` | `400` — *"Data type is not applicable for TABLE_MAPPING type of variable"* | +| `variables/create` without `data_type` | `200` — `primaryconn_primary_data_schema` (`a991fd59-c025-4711-a610-91562b74713e`); the platform also gives it an `obj_id` | +| ADD `PRIMARY_DATA` scoped to Primary; ADD `ORG1_DATA` scoped to org1 | `204`, `204`; read-back shows one value per Org | +| `parameterize` `schemaName` on both Tables | `204`, `204` — **Q14: the field name is `schemaName`** | +| Tables' TML after | `schema: ${primaryconn_primary_data_schema}` | +| Primary data after | Unchanged: NORTH 488997 / 292626, SOUTH 460567 / 274567 | + +Rollback: unparameterize both Tables, then delete the variable. + +### Basic scenario — Steps 3c–3d (2026-09-28) + +| Check | Result | +|---|---| +| Publish (Model only, to org1, from Primary) | `204` | +| **Q7 — what follows the Model** | The Model, both Tables **and the connection** are published to org1. The connection is published but **not returned by `metadata/search` in org1** — confirmed in Atlas: `PRIMARYconn` has `orgId [0, 2137347761]`, `ownerOrgId 0`. The published Tables' `dataSourceId` is `PRIMARYconn` (`deaca893-…`) | +| Re-key (3b) | Needed for the Model, the Tables **and the connection** — the copy's `ORG1conn` carried `PRIMARYconn-deaca893` | +| Published Tables' TML in org1 | Exported live after the publish: `schema: ${primaryconn_primary_data_schema}`. TML shows the definition, not the per-Org value, so TML cannot answer Q13 — only a data read can | +| Copy | Present, `obj_id` `PRIMARYmodel-faea95b6__copy_org1` | +| `searchdata` on the published Model / Table in org1 | ❌ `400`, code 11020, *"Unable to fetch data: Invalid data source guid"* — possibly because the connection is not visible to org1-scoped lookups | +| `searchdata` by name in org1 | `409` `DUPLICATE_OBJECT_FOUND` — copy and published Model share the name; GUIDs only | +| UI search in org1 | ❌ Ref 10028. HAR: `AddColumns` → **Sage errorCode 5 FAILURE**, empty message. The session had both the published Model and the copy selected; the failing column was the published Model's `Cost` | +| Retry via API, later | ❌ 11020 again (incident `ee7eb3d9-8b98-486b-b779-e8e0e05653cf`). **Control: the copy, same query, same run → EAST / WEST ✅** | +| Blocker (temporary) | The published Model could not be queried in org1, by API or UI, while it worked in Primary. Incidents `316c7dd9-…`, `35157087-…`, `ee7eb3d9-…`. Steps 4–7 held; state stayed safe | +| **Root cause (Sage log)** | `sage/auto_complete/…WARNING…`, 05:25 UTC: `org_aware_metadata_snapshot.cpp:196] org_id: 2137347761 doesn't have access to table=faea95b6-…` → `request_validator.cpp:1198] … PERMISSION_DENIED_INACCESSIBLE_TO_ORG`. Sage's org-aware snapshot does not list org1 for the published Model. The Sage process started 02:59:40 UTC; the publish was ~05:14 | +| Atlas | Model `faea95b6-…` and `SALES_FACT` `40639b92-…`: `orgId [0, 2137347761]`, `ownerOrgId 0`, `modifiedMs` 05:08:23 UTC — the publish **was** recorded. So Sage's org-aware snapshot is stale: still refusing at 05:25, 17 minutes after | +| Resolution | **No restart.** Sage's indexing caught up on its own | +| **After indexing caught up** | Published Model in org1 → **EAST 236036 / 123268, WEST 259223 / 128171** — identical to the copy (control). Primary still NORTH / SOUTH. **Q13 answered: the variable resolves per Org; 3d passes** | +| **Finding** | **Sage indexing lags the publish.** Until it catches up, every search on the published Model in the secondary Org fails (`PERMISSION_DENIED_INACCESSIBLE_TO_ORG`; API 11020; UI Ref 10028). Observed: still failing ≥17 min after the publish, then resolved without intervention. 3d must wait and retry | + +### Basic scenario — Step 4 (2026-09-28, org1) + +| Check | Result | +|---|---| +| Copy's sharing (fresh) | `org1_analysts` (group) READ_ONLY | +| Published Model / Tables before | No sharing (admin-only) | +| Column Security Rules | Feature disabled on this cluster (`security/column/rules/fetch` → 403, "Column Security rule feature is disabled"); request needs `tables: [{identifier}]` | +| `share` without `message` | `400` — *"Variable $message of required type String! was not provided"* — nothing written | +| `share` with `message: ""` | `204` Tables, then `204` Model | +| Read-back (`DEFINED`) | Model, `SALES_FACT`, `PRODUCT_DIM`: `org1_analysts` `permission READ_ONLY`, `shared_permission NO_ACCESS` (as expected — read `permission`) | + +### Basic scenario — Steps 5–6 (2026-09-28, org1) + +| Check | Result | +|---|---| +| 5a Backup | Answer TML (YAML, 5,910 chars) returned in one piece; SHA-256 `a2a5d5b0…` matched locally; saved to `~/Documents/Migration_tool_demo/migration/org1/PRIMARYmodel/backup/7e60f180-….answer.tml` with `t0` 1790569266486 | +| 5b Rewrite | One change: `tables[0].fqn` copy → governed; `id` / `name` unchanged (same Model name); no column map | +| 5c Checks | `modified` = `t0`; copy GUID left 0×, governed GUID 1×; `VALIDATE_ONLY` → `OK` | +| 5d Import | `ALL_OR_NONE`, `create_new: false` → `OK`; same GUID; `fqn` now `faea95b6-…` (**Q25: YAML import updates in place with `create_new: false`**; **Q26: `fqn` swap alone is enough when names match**) | +| 6a Resolution | Copy's dependents: **none**. Published Model's dependents: the Answer | +| 6b Numbers | EAST 236036 / 123268, WEST 259223 / 128171 — SHA-256 equals Step 2's `0b07f697…` | +| 6c Sharing | Answer still `org1_analysts` READ_ONLY — **Q29 (sharing) survives an in-place import** | +| 6d Manual | **Pass** — `org1_user` opened the Answer: EAST / WEST, same numbers | + +### Basic scenario — Step 7 (2026-09-28, org1) + +| Check | Result | +|---|---| +| 7a Final sync | Copy: zero dependents; unchanged since the archive (13 columns, same `Margin`, both Tables on `ORG1_DATA`); sharing already carried | +| Rollback set | `migration/org1/PRIMARYmodel/inputs/` (both zips, `SHA256SUMS`) + `backup/7e60f180-….answer.tml` | +| 7b Model | Identity check passed (copy GUID, not published, `obj_id` `…__copy_org1`); `metadata/delete` → `204` | +| 7b Tables | Each then had zero dependents and passed the identity check; `metadata/delete` → `204`, `204` | +| 7b′ Connection | `ORG1conn`: zero dependents, identity check passed; **`connections/{id}/delete` → `204`** (`metadata/delete` has no connection type) | +| 7c Confirm | All four copy objects gone; published Model's dependents = the Answer; Answer data EAST 236036 / 123268, WEST 259223 / 128171 | + +**Result: the basic scenario passes end to end.** `org1_user` sees the same numbers on the +published Model, the Answer kept its GUID and sharing, and the copy with its Tables and +connection is gone. diff --git a/skills/migrate-tml-copies-to-publishing/SKILL.md b/skills/migrate-tml-copies-to-publishing/SKILL.md deleted file mode 100644 index c970012..0000000 --- a/skills/migrate-tml-copies-to-publishing/SKILL.md +++ /dev/null @@ -1,169 +0,0 @@ ---- -description: Move per-tenant Model copies created by TML export/import onto ThoughtSpot Orgs Publishing, so one governed Model serves every Org instead of a copy per Org ---- - -# Migrate TML Model copies to Orgs Publishing - -Customers who adopted Orgs before Publishing existed usually have **one Model per tenant**, -each made by exporting TML from the Primary Org and importing it into a tenant's Org. Every -schema change then has to be repeated per Org, and the copies drift. - -This skill converts that estate to the publishing pattern: **one** governed Model in the -Primary Org, published into each tenant Org, with each tenant's Answers and Liveboards -repointed onto it. The tenant's copy is **kept**, not deleted — it is the rollback. - -## When to Use - -Apply this skill when a customer has per-Org Model copies and wants to adopt Orgs -Publishing without rebuilding their content, and when the inputs available are TML exports -rather than access to both Orgs. - -Not this skill: distributing a Model outward to Orgs that hold no copies yet (ordinary -publishing), or moving a tenant into a brand-new Org. - -## The question you are actually answering - -**Can the governed Model stand in for the tenant's copy?** That is structural, not a naming -question. Two Models with different names, different column labels and different warehouse -schemas can be perfect substitutes; two with identical names can be incompatible. - -What decides it is whether every column the tenant's content reads still resolves to the -same warehouse column, and whether the shape carrying those columns agrees. Both are -decidable offline from two TML exports — see -[references/assessing-two-exports.md](references/assessing-two-exports.md). - -## Prerequisites - -- **Orgs enabled**, and Orgs Publishing enabled for your instance. Confirm with ThoughtSpot - — availability varies by release -- **Administrator privilege in the Primary Org.** Objects can only be published from the - Primary Org, and modifying a published object requires administrator privilege *there* — - an Org administrator of a tenant Org cannot do it, even for their own Org's content -- **Membership in every Org involved.** The dependent lookup and the repoint run inside the - tenant Org and take their scope from your session. A missing membership shows up as an - **empty dependent list**, not as a permission error — which reads exactly like "nothing - to migrate" -- Two TML exports per Model: one from the Primary Org, one from the tenant Org - -## Procedure - -Steps 1–5 write nothing. Confirm with the customer before each step from 6 onward. - -| # | Step | Writes? | -|---|---|---| -| 1 | Collect both TML exports per Model | no | -| 2 | Pair the tables and compare structure | no | -| 3 | Review findings and the column map | no | -| 4 | Identify the fields that differ per Org | no | -| 5 | Report the verdict: READY / RENAME / BLOCKED | no | -| 6 | Parameterize the per-Org fields on the governed Model | yes | -| 7 | Re-key the copy's custom object id | yes | -| 8 | Publish the governed Model into the tenant Org | yes | -| 9 | Find what depends on the copy | no | -| 10 | Repoint each dependent | yes | -| 11 | Verify | no | - -**Export both with dependencies**, rooted at the Model, not at a Liveboard: -`export_associated: true`, `export_fqn: true`, `edoc_format: YAML`. - -Two silent mistakes to catch at Step 1: both exports taken from the same Org (they compare -perfectly clean, because they are the same object — compare the root GUIDs), and platform -entries in the zip that will not parse. Never filter zip entries by name to work around the -second; name the entry that failed. - -## Order matters: 6 and 7 before 8 - -**Parameterization is the work, not a preliminary.** A Model whose dependency tree contains -no template variable cannot be published at all — the platform refuses with *"No template -variable node found in the dependency tree."* The per-Org fields found at Step 4 (typically -the warehouse `schema`, sometimes the database or table name) become variables with a value -per Org. - -One trap: that check is an **existence** check across the whole tree, not a completeness -check. If *something* in the tree is parameterized the publish succeeds — even if a table -you meant to parameterize was missed. A partially parameterized publish exposes Primary Org -data to the tenant. Verify every field you intended, per Org, yourself. - -**The copy inherited the original's custom object id.** That id must be unique within an -Org, and publishing makes the governed Model a member of the tenant's Org — so both objects -are then in one Org carrying one id, and the publish is refused. Re-key **the copy**, not -the governed Model, via `POST /api/rest/2.0/metadata/headers/update` from a tenant-Org -session. Always set a new value; never blank it. - -## Let the platform refuse what only it can see - -Publishing validates before writing anything, and the commit is transactional — a refused -publish leaves the estate completely unchanged. So an attempted publish is cheap. It -refuses on: - -| Error | Meaning | -|---|---| -| `Cannot publish/unpublish objects with Cohort Column as dependency` | a Set is in the closure. Sets are invisible in TML, so do not try to detect one from the exports — attempt the publish and read this | -| `No template variable node found in the dependency tree` | Step 6 was skipped or did not land | -| duplicate custom object id | Step 7 was skipped for this Org | -| `Objects can only be published/unpublished from primary org` | wrong Org session | - -Spend the offline effort on what TML *can* settle — correspondence and shape. Let the -platform refuse the rest. - -## Repointing, and its two traps - -A dependent references the **columns** it reads, so what moves it is swapping the Model -reference it resolves through: export its TML, swap `tables[].fqn` from the copy's GUID to -the governed Model's GUID, apply the column map, re-import over the same object. - -**Views are a layer, not an endpoint.** A dependent that is a View or SQL View exposes its -own column names to everything built on it. Rewrite the names it *reads* — `tables[].fqn`, -formulas, column references — and **preserve exactly** the names it *exposes* and its own -name. Get this wrong and every Answer above the View breaks, and none of them appear in the -copy's dependent list, because they depend on the View. Get it right and content on that -View needs no rewrite at all. Do Views first, then re-run the dependent lookup. - -**Some dependents are nested.** A dependency can be carried by a hidden object belonging to -a Liveboard rather than by the Liveboard itself. Repoint the **owner**, and de-duplicate — -several nested objects in one Liveboard collapse to one rewrite. - -Check coverage **before** importing: scan the rewritten document for any surviving reference -to a source column name. A partial rewrite imports cleanly and renders wrong, which the -customer finds rather than you. After importing, confirm the object resolves to the governed -Model. - -## What a TML round trip does not carry - -Presentation, not meaning. Conditional formatting, column widths, chart state, parameter -ids and Liveboard tile layout are not guaranteed to survive a re-import. Which columns an -Answer reads, what it computes and what it filters are. - -So **do not refuse a migration to protect formatting** — that leaves the Answer pointing at -a Model about to be retired, which is worse. Instead: disclose per dependent before writing, -keep the pre-repoint export as both rollback and reference, and verify that the Answer -returns the **same numbers**, never that the documents match. - -## Safety - -Nothing is deleted. At every stage there is a way back: - -| Stopped after | To undo | -|---|---| -| 1–5 | nothing was written | -| 6 | un-parameterize the fields | -| 7 | set the previous custom object id back | -| 8 | unpublish from that Org | -| 10, partly | re-import the pre-repoint TML for the ones already moved | - -**Validate the whole sequence in a non-production Org before running it against a tenant's -live content**, and repoint one dependent at a time rather than in a batch — a failure -part-way leaves the rest still pointing at the copy, which is a safe state. - -## Not covered - -- **Deleting the copy.** Deliberately: it is the rollback. Retiring it is a separate - decision, after verification -- **Answers and Liveboards as publish roots.** They are read as dependents to be repointed -- **Sets / Cohorts.** They block publishing and cannot be seen in TML. Resolve with the - customer -- **Sharing and permissions.** TML carries no sharing information, so nothing here - preserves or migrates it. Check the published Model's grants in each tenant Org explicitly -- **Tenant data isolation.** Publishing one Model to several Orgs does not by itself - separate their data — that comes from row-level security or from the publication - variables. It is a required review, and it is not established by this migration diff --git a/skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md b/skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md deleted file mode 100644 index c2f636b..0000000 --- a/skills/migrate-tml-copies-to-publishing/references/assessing-two-exports.md +++ /dev/null @@ -1,101 +0,0 @@ -# Assessing two TML exports - -Everything on this page is decidable **offline**, from the two export zips, with no -connection to either Org. It is the whole of Steps 1–5, and it is where the judgement in -this migration actually lives. - -## What TML carries, and why that is enough - -A Model export carries more than labels: - -| in the TML | what it settles | -|---|---| -| `db_name`, `schema`, `db_table`, `db_column_name` | the **physical binding** — which warehouse column a Model column reads | -| `model_tables` / `fqn` | which tables the Model is built from | -| each join's `on`, `type`, `cardinality` | the **shape** carrying those columns | -| `formula` | derived columns, comparable as text | -| column `name`, `description` | the tenant-facing labels, which may differ freely | - -So correspondence ("is this the same warehouse column?") and shape ("does the schema graph -agree?") are both computable without touching a cluster. Names are the one thing that may -differ, and the mapping between them is the output. - -## Pair the tables before comparing anything - -Do not pair by name — a copy may have been renamed. Pair on **column sets**: for each table -in the governed export, find the table in the tenant export whose `db_column_name` set -overlaps most. A clean pair overlaps at or near 1.00. - -Report the overlap for every pair. An overlap well below 1.00 that still wins is the -signal that the two Models have diverged, and it deserves a human look rather than a -verdict. - -## Compare, and what may legitimately differ - -Three differences are expected and are **not** findings: - -- **`schema`** (and sometimes `db_name` or `db_table`) — this is exactly what per-Org - parameterization exists to carry. Collect these; they become Step 6's work -- **connection name** — compare the connection *type* instead -- **GUIDs** — every object has its own - -Everything else is a finding. In particular: - -- a column the tenant's content uses that the governed Model does not have — **blocking**, - and no mapping can fix it. Ask the customer; do not guess a near match -- a differing `formula` for columns that otherwise correspond — blocking -- a join present on one side and not the other, or with a different `on` clause - -`cardinality` is **view-relative**: the same join reads `ONE_TO_MANY` or `MANY_TO_ONE` -depending on which table's block it is written under. Normalise before comparing, or every -join looks changed. - -## The column map is the output - -Where corresponding columns carry different names, record -`tenant name -> governed name`. That map drives the repoint later, and it is the artifact -worth reviewing with the customer before any write: - -- an **identity** map means no content rewriting is needed for names at all — the cheapest - possible migration -- a non-empty map means every dependent must be rewritten, and each rewrite is a chance to - get a name wrong - -Review the map before Step 6, not after Step 8. - -## Verdicts - -| verdict | meaning | -|---|---| -| `READY` | corresponds cleanly, identity column map — publish and repoint | -| `RENAME` | corresponds cleanly, but names differ — same path, plus content rewriting | -| `BLOCKED` | something no mapping can fix. Report it and stop | - -Running this across a whole estate first is worth far more than doing it one Model at a -time: a sweep that reports twelve Models `READY` and three `BLOCKED` is one conversation -with the customer instead of fifteen. - -## Two holes to be honest about - -**Multi-join-path.** Where the same table occupies several slots in the schema graph -(role-playing dimensions — an order date and a ship date on the same calendar table), the -column sets are identical and pairing cannot tell the slots apart. Detect the condition — -the same table appearing more than once in the join graph — and escalate rather than -reporting a clean pair. - -**Where a pair binds the same warehouse column under different labels**, the comparison -sees a match on binding and a difference on name, which is exactly the `RENAME` case. That -is correct. But two *different* columns that happen to bind the same warehouse column are -indistinguishable from each other on binding alone. Fall back to name and formula, and -surface both candidates rather than choosing. - -## Two things TML cannot see at all - -- **Sets / Cohorts.** They do not appear in a TML export, and they block publishing. - A clean assessment is not evidence that there is no Set. The publish attempt is the only - test, and it refuses without writing anything -- **Sharing and permissions.** No sharing information is carried, so a clean assessment - says nothing about who can see what afterwards - -A clean assessment means *the governed Model can stand in for the copy*. It does not mean -the migration will succeed. From 7047ecd4bf61d77134e212cf523131d4999eb5ab Mon Sep 17 00:00:00 2001 From: Tanishka SInghal Date: Mon, 28 Sep 2026 16:02:37 +0530 Subject: [PATCH 3/4] Add migrate-tml-copies-to-publishing skill v0 from the tested flow The procedure that passed on a live multi-Org cluster, written for an agent: compare the two TML exports, re-key the copy's obj_ids (Model, Tables, connection), parameterize the schema with a per-Org variable, publish, wait for the published Model to be searchable, carry the copy's sharing, back up and repoint each Answer in place, verify, and delete the copy with its Tables and connection. v0 covers one Org, a copy identical column for column, and Answers; anything else is refused as not yet supported. Co-Authored-By: Claude Opus 5.5 (1M context) --- CLAUDE.md | 6 + README.md | 1 + .../migrate-tml-copies-to-publishing/SKILL.md | 208 ++++++++++++++++++ 3 files changed, 215 insertions(+) create mode 100644 skills/migrate-tml-copies-to-publishing/SKILL.md diff --git a/CLAUDE.md b/CLAUDE.md index 7404bf2..44b6027 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,3 +9,9 @@ This plugin integrates ThoughtSpot developer documentation with Claude Code, giv - **get-visual-embed-sdk-reference** — Guidance for embedding ThoughtSpot content in web applications - **get-rest-api-reference** — Guidance for working with the ThoughtSpot REST API v2 - **get-developer-docs-reference** — Guidance for searching general ThoughtSpot developer documentation + +### Guided procedures + +- **migrate-tml-copies-to-publishing** — Replace a Model copy made by TML export/import in a secondary Org with the governed Model published into that Org, carrying the copy's sharing and repointing its Answers (v0: one Org, identical columns, Answers only) + +Procedure skills write to your instance. Every write step needs a confirmation, backs up what it changes, and deletes the replaced copy only after verification. diff --git a/README.md b/README.md index 887cf85..04cc8d7 100644 --- a/README.md +++ b/README.md @@ -7,6 +7,7 @@ Configuration for integrating the ThoughtSpot developer documentation MCP server - **Get Visual Embed SDK Docs** — Types, interfaces, configuration, events, authentication, and CSS theming - **Get REST API Docs** — Endpoint specs, request/response schemas, authentication, and Java/TypeScript SDK guides - **Get Developer Docs** — Platform features, SSO/SAML, deployment, TML, and general ThoughtSpot guidance +- **Migrate TML Copies to Publishing** — Guided procedure that replaces a secondary Org's TML-imported Model copy with the published governed Model, carrying sharing and repointing Answers (v0) ## Prerequisites diff --git a/skills/migrate-tml-copies-to-publishing/SKILL.md b/skills/migrate-tml-copies-to-publishing/SKILL.md new file mode 100644 index 0000000..577d591 --- /dev/null +++ b/skills/migrate-tml-copies-to-publishing/SKILL.md @@ -0,0 +1,208 @@ +--- +description: Move a Model copy that was made by TML export/import into a secondary Org onto ThoughtSpot Orgs Publishing — publish the governed Model into that Org, carry the copy's sharing, repoint the Answers built on the copy, and delete the copy. Phase 1 (v0) supports one Org and one Model whose copy matches the governed Model column for column. +--- + +# Migrate a TML Model copy to Orgs Publishing (v0) + +Before Publishing, customers shared a Model with a secondary Org by exporting its TML from +the Primary Org and importing it there. This skill replaces that **copy** with the governed +Model, published into the Org, without breaking the content built on the copy. + +Design and live-test record: `docs/migrate-tml-copies-design.md`, +`docs/migrate-tml-copies-test-scenarios.md` in this plugin's repository. + +## Scope of v0 — refuse everything else + +Supported, and tested end to end on a live cluster: + +- one governed Model (Primary Org) and its copy in **one** secondary Org +- the copy matches the governed Model **column for column** (verdict `READY`); only the + warehouse `schema` differs +- dependents on the copy are **Answers** +- the copy's sharing is object-level (users and groups) + +**Stop and say "not supported in this version yet"** when any of these appear: renamed or +missing columns, a changed formula or join, a View on the copy, a Liveboard, a Set +(`COHORT`) dependent, content built directly on the copy's Tables, column-level sharing or +Column Security Rules on the copy, or a different database, table name or connection type. + +## Requirements + +- `execute-thoughtspot-code` with an **`org_identifier`** parameter. The browser sign-in + always lands in the user's default Org, so without it the secondary Org is unreachable. + If the tool has no `org_identifier`, stop and say so. +- The user is an administrator in the Primary Org **and** the secondary Org. +- Inputs: the governed Model's TML export (Primary Org), the copy's TML export (secondary + Org) — both exported **with dependencies**, from the UI or REST — and the secondary Org's + name. + +## Rules for every step + +- **GUIDs only.** The copy and the published Model share a name in the secondary Org, and a + lookup by name returns `DUPLICATE_OBJECT_FOUND`. +- **Name the Org on every call.** A read in the wrong Org returns an empty list, not an error. +- **One write step per run, one approval per run.** Show the exact code before each write. + Read the current state first and write only what is missing, so a re-run is safe. +- **Stop at the first failed check.** Every step leaves a state that works. +- Keep `migration///` locally: `inputs/` (both exports + checksums), `backup/`, + and a `ledger.json` of every write with its before and after values. + +## Steps + +| # | Step | Org | Writes | +|---|---|---|---| +| 1 | Compare the two exports | — (local) | no | +| 2 | Live checks and dependents | secondary | no | +| 3b | Re-key the copy's `obj_id`s | secondary | yes | +| 3a | Variable for the schema | Primary | yes | +| 3c | Publish | Primary | yes | +| 3d | Verify the publish | secondary | no | +| 4 | Carry the copy's sharing | secondary | yes | +| 5 | Back up and repoint each Answer | secondary | yes | +| 6 | Verify | secondary | no | +| 7 | Delete the copy, its Tables and connection | secondary | yes | + +3b runs before 3a so both secondary-Org steps come before the switch; it only has to +precede 3c. + +### 1. Compare the two exports (local) + +Unzip both. Each must hold one Model, its Table files and their Connection file, and every +file must parse. **Stop** if the two Model GUIDs are equal (both exports came from one Org). +Resolve each Model table through its `fqn` to a Table file. + +Pair Tables by `db_table`, then compare: + +- **Columns** — by physical binding (`
.`), then name, `column_type`, + `aggregation`. Formulas: resolve `[TABLE::column]` references and **normalise whitespace** + before comparing (exports differ by trailing spaces). +- **Joins** — `on`, `type`, `cardinality`. +- **Tables** — only `schema` may differ. Record the copy's value per Table. +- **Connections** — by content, never name: `type`, non-secret properties (account, user, + role, warehouse), `selected_databases`. Ignore secrets. +- **`obj_id`s** — list every copy object (Model, Tables, Connection) whose `obj_id` equals + the governed one's. They all need re-keying. + +Verdict `READY` only if everything matches except `schema`. Otherwise stop (out of scope). +Archive both zips unchanged in `inputs/` with SHA-256 checksums. + +### 2. Live checks and dependents (secondary Org, read-only) + +- Admin in both Orgs; the governed GUID exists in Primary; the copy's GUID exists in the + secondary Org and is **not** a published object (`include_only_published_objects: true` + returns nothing). +- Dependents of the copy: `metadata/search` on the copy's GUID with + `include_dependent_objects: true`, `dependent_objects_record_size: -1`, + `dependent_object_version: "V2"`. Answers arrive under `QUESTION_ANSWER_BOOK`. A string in + `dependent_objects` means the lookup failed — stop. Any other dependent type → stop. +- The same lookup on each of the copy's Tables must return **only the copy**. Anything else + is content built directly on a Table → stop. +- Per Answer: export its TML; the `[…]` tokens in `search_query` must all be copy columns. +- Sharing snapshot: `security/metadata/fetch-permissions` with `permission_type: "DEFINED"` + on the copy and on each Answer. +- Data fingerprint per Answer: `metadata/answer/data`, rows sorted, SHA-256. + +### 3b. Re-key the copy's `obj_id`s (secondary Org) + +For each object in the Step 1 list, read its `obj_id` first and stop if it changed. Then +`metadata/update-obj-id` — identify by GUID, **not** also by `current_obj_id`: + +```json +{ "metadata": [{ "metadata_identifier": "", "type": "LOGICAL_TABLE", "new_obj_id": "__copy_" }] } +``` + +The copy's Tables use `LOGICAL_TABLE`; its connection uses `DATA_SOURCE`, in a separate call. +Publishing brings the governed connection into the Org too (hidden from search), so the +connection clashes as well. + +### 3a. Variable for the schema (Primary Org) + +If the governed Tables already use a variable (`schema: ${…}` in their TML), only add this +Org's value. Otherwise, in this order: + +1. `template/variables/create` — `{ "type": "TABLE_MAPPING", "name": "" }`, **no + `data_type`** (refused for this type). + Name: `{connection}_{current Primary value}_{field}`, slugified, e.g. + `primaryconn_primary_data_schema`. Tables sharing a value share the variable. Reuse is + decided from the Tables' bindings and the variable's values, never from the name. +2. `template/variables/update-values` — `operation: "ADD"` only, never `REPLACE` or `RESET`. + **Primary's value first** (the current literal), then the secondary Org's (the copy's + schema), each with `variable_value_scope: [{ "org_identifier": "" }]`. +3. `metadata/parameterize` per Table — `metadata_type: "LOGICAL_TABLE"`, + `field_type: "ATTRIBUTE"`, `field_name: "schemaName"`. +4. Read back the values, and confirm Primary's data is unchanged. + +If the secondary Org already has a different value → stop and ask. + +### 3c. Publish (Primary Org) + +`security/metadata/publish` — `{ "metadata": [{ "identifier": "", "type": "LOGICAL_TABLE" }], "org_identifiers": [""] }`. +Never `skip_validation`. The Tables and connection follow the Model. + +### 3d. Verify the publish (secondary Org, read-only) + +Query the published Model with `searchdata` and compare with the copy's numbers for the same +query. **Sage indexing lags the publish**: until it catches up, searches fail with 11020 +*"Invalid data source guid"* (UI: Ref 10028). Retry every few minutes; if it never +succeeds, stop — "published, but not yet searchable in this Org". Never start Step 4 +before the published Model returns the copy's numbers. + +### 4. Carry the copy's sharing (secondary Org) + +Re-read the copy's `DEFINED` sharing now. Grant each principal on the published **Tables +first**, then the **Model** (a grant on an object whose source is not shared can be silently +dropped): + +```json +{ "metadata": [{ "type": "LOGICAL_TABLE", "identifier": "" }], + "permissions": [{ "principal": { "type": "USER_GROUP", "identifier": "" }, "share_mode": "READ_ONLY" }], + "message": "" } +``` + +`message` is required. `MODIFY` on the copy becomes `READ_ONLY` — say so. Never lower +access a principal already has. Read back: check `permission`, not `shared_permission`. + +### 5. Back up and repoint each Answer (secondary Org) + +Per Answer, two runs: + +- **Backup (read)** — record `metadata_header.modified` as `t0`; export YAML + (`export_fqn: true`); return it with its SHA-256 (in slices if over the output limit); + save it to `backup/.answer.tml` and check the hash. +- **Repoint (write)** — stop if `modified` ≠ `t0`. Export again; replace the copy's GUID + with the governed GUID in `tables[].fqn`; the copy's GUID must appear **0** times and the + governed GUID once. `metadata/tml/import` with `import_policy: "VALIDATE_ONLY"`, then + `"ALL_OR_NONE"`, `create_new: false`. Re-export: same GUID, new `fqn`. Read its data. + +Rollback: re-import the backup over the same GUID. + +### 6. Verify (secondary Org, read-only) + +The copy has no dependents; each Answer depends on the published Model; each Answer's +fingerprint equals Step 2's; each Answer's sharing is unchanged. Then ask the user to open +an Answer as a **non-admin** member of a granted group and confirm the numbers. + +### 7. Delete the copy (secondary Org) + +Only after Step 6 passes and the user confirms. Re-check: zero dependents, the copy +unchanged since the archive, no new sharing. Then, each after an identity check (the copy's +GUID, not a published object, `obj_id` ending `__copy_`): + +1. the copy Model — `metadata/delete` +2. each copy Table, if it now has zero dependents — `metadata/delete` +3. the copy's connection, if no Table depends on it — `connections/{id}/delete` + (restoring it later needs its credentials re-entered; say so first) + +Confirm all are gone and the Answers still return their numbers. + +## Rollback + +| Stopped after | Undo | +|---|---| +| 1–2 | nothing written | +| 3b | `update-obj-id` back to the old values | +| 3a | unparameterize the Tables, delete the variable (if created) | +| 3c | `security/metadata/unpublish` from the Org — before Step 5 only | +| 4 | share `NO_ACCESS` for each principal added | +| 5 | re-import each backup | +| 7 | re-import the copy from `inputs/` with the re-keyed `obj_id`s, then the backups | From 289a55ccfcf700444f4784e00d8e943e32dc7669 Mon Sep 17 00:00:00 2001 From: Tanishka SInghal Date: Tue, 29 Sep 2026 12:13:00 +0530 Subject: [PATCH 4/4] Update migrate-tml-copies skill from the mig_1_2 test run The first run from SKILL.md alone (Org mig_1_2) passed end to end and exposed gaps in the instructions: - Requirements: allow one MCP server per Org (token per Org) when the tool has no org_identifier; confirm each server's current_org first. - Step 2: a Table's V2 dependents include content built on the Model; only stop when a dependent's TML references the Table directly. Stop if inaccessible dependents are not returned. Save raw responses. - Step 3a: split into case A (variable exists: add this Org's value only) and case B. Give the exact update-values body: variable_value_scope is top-level, not inside the assignment. Full parameterize body. Primary check query after either case. - Step 7: check the copy is unchanged by comparing TML with inputs/, ignoring obj_id; the 3b re-key changes its modified time. - Add a Report step that writes report.md for the customer. - Ignore migration/ (exports, backups, ledgers). Co-Authored-By: Claude Opus 5.5 (1M context) --- .gitignore | 3 + .../migrate-tml-copies-to-publishing/SKILL.md | 74 ++++++++++++++----- 2 files changed, 60 insertions(+), 17 deletions(-) diff --git a/.gitignore b/.gitignore index e43b0f9..b36f714 100644 --- a/.gitignore +++ b/.gitignore @@ -1 +1,4 @@ .DS_Store + +# Migration runs: exports, backups, ledgers (customer data) +migration/ diff --git a/skills/migrate-tml-copies-to-publishing/SKILL.md b/skills/migrate-tml-copies-to-publishing/SKILL.md index 577d591..bf8edfb 100644 --- a/skills/migrate-tml-copies-to-publishing/SKILL.md +++ b/skills/migrate-tml-copies-to-publishing/SKILL.md @@ -28,9 +28,14 @@ Column Security Rules on the copy, or a different database, table name or connec ## Requirements -- `execute-thoughtspot-code` with an **`org_identifier`** parameter. The browser sign-in - always lands in the user's default Org, so without it the secondary Org is unreachable. - If the tool has no `org_identifier`, stop and say so. +- A way to run `execute-thoughtspot-code` in **each** Org. The browser sign-in always lands + in the user's default Org, so use one of: + - the tool's **`org_identifier`** parameter, if it has one; or + - **one MCP server per Org**, each signed in with a token for that Org (the user names + which server is which Org). Before the first call, confirm each server's Org with + `GET /api/rest/2.0/auth/session/user` (`current_org`), and stop if one is wrong. + + If neither is available, stop and say so. - The user is an administrator in the Primary Org **and** the secondary Org. - Inputs: the governed Model's TML export (Primary Org), the copy's TML export (secondary Org) — both exported **with dependencies**, from the UI or REST — and the secondary Org's @@ -40,7 +45,8 @@ Column Security Rules on the copy, or a different database, table name or connec - **GUIDs only.** The copy and the published Model share a name in the secondary Org, and a lookup by name returns `DUPLICATE_OBJECT_FOUND`. -- **Name the Org on every call.** A read in the wrong Org returns an empty list, not an error. +- **Name the Org on every call** (or use that Org's server). A read in the wrong Org returns + an empty list, not an error. - **One write step per run, one approval per run.** Show the exact code before each write. Read the current state first and write only what is missing, so a re-run is safe. - **Stop at the first failed check.** Every step leaves a state that works. @@ -61,6 +67,7 @@ Column Security Rules on the copy, or a different database, table name or connec | 5 | Back up and repoint each Answer | secondary | yes | | 6 | Verify | secondary | no | | 7 | Delete the copy, its Tables and connection | secondary | yes | +| — | Write the report | — (local) | no | 3b runs before 3a so both secondary-Org steps come before the switch; it only has to precede 3c. @@ -95,12 +102,17 @@ Archive both zips unchanged in `inputs/` with SHA-256 checksums. `include_dependent_objects: true`, `dependent_objects_record_size: -1`, `dependent_object_version: "V2"`. Answers arrive under `QUESTION_ANSWER_BOOK`. A string in `dependent_objects` means the lookup failed — stop. Any other dependent type → stop. -- The same lookup on each of the copy's Tables must return **only the copy**. Anything else - is content built directly on a Table → stop. + If `hasInaccessibleDependents` is true and `areInaccessibleDependentsReturned` is not, + some dependents are missing from the list → stop. +- The same lookup on each of the copy's Tables. A Table can also list content built on the + copy (column-level, indirect). For each Table dependent other than the copy, read its TML: + if `tables[]` references only the copy Model, it is indirect — continue. If it references + the Table's GUID, it is content built directly on a Table → stop. - Per Answer: export its TML; the `[…]` tokens in `search_query` must all be copy columns. - Sharing snapshot: `security/metadata/fetch-permissions` with `permission_type: "DEFINED"` on the copy and on each Answer. - Data fingerprint per Answer: `metadata/answer/data`, rows sorted, SHA-256. +- Save the raw dependents responses in the migration folder. ### 3b. Re-key the copy's `obj_id`s (secondary Org) @@ -117,8 +129,28 @@ connection clashes as well. ### 3a. Variable for the schema (Primary Org) -If the governed Tables already use a variable (`schema: ${…}` in their TML), only add this -Org's value. Otherwise, in this order: +**Case A — the governed Tables already use a variable** (`schema: ${…}` in their TML, e.g. +already published to another Org). Create nothing and parameterize nothing: + +1. Re-read the variable's values. This Org already has the copy's schema → skip the step; + a different value → stop and ask. +2. `template/variables/update-values` — one `ADD` of the copy's schema, scoped to this Org + only: + + ```json + { "variable_assignment": [{ "variable_identifier": "", "variable_values": [""], "operation": "ADD" }], + "variable_value_scope": [{ "org_identifier": "" }] } + ``` + + `variable_value_scope` is **top-level**, next to `variable_assignment` — not inside it + (400: *"Field variable_value_scope is not defined by type VariableUpdateAssignmentInput"*). + One call per Org. +3. Read back all values: the other Orgs' values are unchanged. +4. Run a Primary check query on the governed Model: Primary's numbers are unchanged. + +Undo: remove this Org's value. + +**Case B — the governed Tables use a literal schema.** In this order: 1. `template/variables/create` — `{ "type": "TABLE_MAPPING", "name": "" }`, **no `data_type`** (refused for this type). @@ -127,12 +159,11 @@ Org's value. Otherwise, in this order: decided from the Tables' bindings and the variable's values, never from the name. 2. `template/variables/update-values` — `operation: "ADD"` only, never `REPLACE` or `RESET`. **Primary's value first** (the current literal), then the secondary Org's (the copy's - schema), each with `variable_value_scope: [{ "org_identifier": "" }]`. -3. `metadata/parameterize` per Table — `metadata_type: "LOGICAL_TABLE"`, - `field_type: "ATTRIBUTE"`, `field_name: "schemaName"`. -4. Read back the values, and confirm Primary's data is unchanged. - -If the secondary Org already has a different value → stop and ask. + schema), one call each, same body as case A step 2. +3. `metadata/parameterize` per Table: + `{ "metadata_type": "LOGICAL_TABLE", "metadata_identifier": "
", "field_type": "ATTRIBUTE", "field_name": "schemaName", "variable_identifier": "" }`. +4. Read back the values, and run a Primary check query on the governed Model: Primary's + numbers are unchanged. ### 3c. Publish (Primary Org) @@ -184,8 +215,9 @@ an Answer as a **non-admin** member of a granted group and confirm the numbers. ### 7. Delete the copy (secondary Org) -Only after Step 6 passes and the user confirms. Re-check: zero dependents, the copy -unchanged since the archive, no new sharing. Then, each after an identity check (the copy's +Only after Step 6 passes and the user confirms. Re-check: zero dependents, no new sharing, +and the copy unchanged since the archive — compare its live TML with `inputs/`, ignoring +`obj_id` lines. Don't use the modified time: the 3b re-key changes it. Then, each after an identity check (the copy's GUID, not a published object, `obj_id` ending `__copy_`): 1. the copy Model — `metadata/delete` @@ -195,13 +227,21 @@ GUID, not a published object, `obj_id` ending `__copy_`): Confirm all are gone and the Answers still return their numbers. +### Report + +Write `migration///report.md` and give the user its path. Include: the Orgs and +Models (names and GUIDs); the Step 1 verdict; each step's result; the numbers before and +after for every Answer, and Primary's check query; the sharing carried; what was deleted; +anything disclosed (renamed columns, `MODIFY` lowered to `READ_ONLY`); and the rollback +files (`inputs/`, `backup/`, `ledger.json`). + ## Rollback | Stopped after | Undo | |---|---| | 1–2 | nothing written | | 3b | `update-obj-id` back to the old values | -| 3a | unparameterize the Tables, delete the variable (if created) | +| 3a | case A: remove this Org's value; case B: unparameterize the Tables, delete the variable | | 3c | `security/metadata/unpublish` from the Org — before Step 5 only | | 4 | share `NO_ACCESS` for each principal added | | 5 | re-import each backup |