Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
113 changes: 113 additions & 0 deletions PLANS.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,119 @@ plan_schema_version: 2

Use this file for active, blocked, ready-for-closure, or recently completed execution work. The canonical lifecycle is the installed `engineering-workflow` planning reference.

## Active Plan: GPT-6 Model Profiles And Marketplace Release

Status: active
Owner: root
Last Updated: 2026-09-23

### Goal

Update the engineering-workflow model recommendations and optional Codex agent templates for the current GPT-6 family, publish a source release, import it into xeonvs-engineering, and refresh the managed local marketplaces.

### Plan Origin

direct_execution

### Requested Scope

- Support the current models in the relevant source and distribution repositories and release the marketplace update.

### Requirement Traceability

| Requirement | Complete outcome | Source | Work queue | Acceptance or validation | Status |
| --- | --- | --- | --- | --- | --- |
| REQ-001 | Current Codex role mappings, optional agent templates, and model-specific guidance reflect supported GPT-6 models without changing Claude inheritance or user pins. | User request; official OpenAI model catalog and guidance | WQ-01 | Model and effort readback, behavioral/contract tests, aggregate diff review | done |
| REQ-006 | A target upgrade refreshes previously generated Codex agent model profiles when their prior bytes are pristine, while preserving customized files and established opt-in state. | User clarification | WQ-01 | Report/apply/prompt fixture matrix for pristine, customized, and opt-out targets | done |
| REQ-002 | All active source/package version owners and generated bytes agree in a new source release. | User request; repository release contract | WQ-02 | Full release/security gate, package validators, PR merge, annotated tag and release readback | pending |
| REQ-003 | xeonvs-engineering imports exact immutable source bytes and publishes its own release while preserving tgrep-search. | User request; marketplace sync contract | WQ-03 | Recorded provenance/byte verification, catalog tests, release workflow and asset readback | pending |
| REQ-004 | Local Codex and Claude managed marketplaces and workflow installations resolve the published version. | Prior user preference; current release request | WQ-04 | Native CLI update and active version readback | pending |
| REQ-005 | Source and marketplace plans close with durable release evidence. | Workflow lifecycle contract | WQ-05 | Lifecycle check and closure readback | pending |

### Explicit Non-Goals

- Change tgrep-search source or version; overwrite user-pinned models; copy Responses API-only fields into Codex TOML; rewrite published history.

### Constraints

- Preserve role-specific cost, latency, reasoning, and read-only behavior; confirm target slugs and effort levels against current official documentation and actual Codex availability.
- Keep source repository canonical and marketplace distribution generated from a stable annotated source tag.
- Review each logical commit and aggregate release diff; run final full and pre-push security gates.

### Inputs And Sources

- https://developers.openai.com/api/docs/models
- https://developers.openai.com/api/docs/guides/latest-model
- Current `model_profiles.md`, optional agent templates, validators/tests, source and marketplace release contracts.
- Repository audit summary in `/tmp/engineering-model-audit-20260923.json` (generic migration findings are outside this release's model scope).

### User Decisions And Answers

- 2026-09-23: Add support for new models in the relevant repositories and publish the marketplace update.
- 2026-09-23: Retain GPT-5.6 where it is cheaper or useful as a fallback; current API prices support Terra as an availability/evaluation fallback, not the cheaper default for these roles.
- 2026-09-23: Route routine commands and tests through tools without a model worker; upgrade an already configured target's generated model profiles when pristine.
- 2026-09-23: Keep Astra available for unusually complex/high-consequence work; require a concrete rationale and user confirmation before the agent initiates a more expensive Astra worker or saved-profile escalation. An explicit user choice already confirms it.
- Earlier session: after publication, refresh local managed marketplaces and installed skills.

### Completed Baseline State

- [x] Source main is clean at the start; engineering-workflow 0.9.7 and marketplace 1.0.7 are the last completed releases. The source already maps standard/review to Astra, while utility/explorer remain pinned to GPT-5.6 Terra.

### Current Work Queue

- [x] WQ-01 — Implement and validate GPT-6 task routing, profiles, templates, and conservative target migration for REQ-001/REQ-006. `done`
- [ ] WQ-02 — Bump source version, validate, review, merge, tag and release for REQ-002. `in_progress`
- [ ] WQ-03 — Import, validate, review, merge and release xeonvs-engineering for REQ-003. `pending`
- [ ] WQ-04 — Refresh and read back local managed installations for REQ-004. `pending`
- [ ] WQ-05 — Reconcile and close source and marketplace plans for REQ-005. `pending`

### Locked Decisions

- Patch release target is provisionally engineering-workflow 0.9.8 and xeonvs-engineering 1.0.8, subject to current remote tag inspection.
- Route deterministic commands and tests through tools; recommend Luna low for bounded utility semantics and Sol medium for exploration, standard work and routine review. Raise Sol review effort to high only for justified complexity. Reserve Astra high for explicit user choice or confirmed agent-proposed escalation. Preserve supported user pins and explicit Terra fallback.

### Verification

- Official model/effort source and current Codex model availability; focused model-profile tests; source full/release/security checks and plugin validators.
- Marketplace sync provenance and recorded-byte verification, catalog tests, release workflow, checksummed asset, and installed-version readback.

### Latest Validation Results

- 2026-09-23: Official OpenAI model catalog, migration guidance, Codex model selection, and subagent configuration opened; GPT-6 Astra, Sol, Luna and supported low/medium/high efforts verified. Terra's published API pricing ($2 input/$12 output per million tokens) exceeds Luna's and does not undercut Sol's $2 input/$10 output; subscription usage is not inferred from API prices. Source audit produced a bounded report; initial checkout was clean.
- 2026-09-23: Source/package 0.9.8 generated in parity. Full gate passed 9/9 with 266 tests (one skipped); review found and corrected an overly costly routine reviewer effort. Migration tests cover previously opted-in pristine templates, customized model pins, and targets without opt-in. No current upstream 0.9.8 tag or open source PR exists.
- 2026-09-23: Source and generated package quick validation, Codex plugin validation, and strict Claude plugin/marketplace validation passed. Final review incorporated the user's explicit confirmation rule for agent-initiated Astra escalation; full release gate must be rerun on those final bytes.

### Risks And Recovery

- A proposed model may not be exposed in a specific Codex installation. Verify locally before committing template defaults; preserve existing supported pins if unavailable.
- Remote refs may advance during release. Reinspect exact head and tags before push/merge; revalidate changed content.
- Marketplace import may reject candidate bytes or provenance. Stop before publishing and correct only the failed import stage.

### Resume Point

- WQ-02: run final release/security validation and external plugin validators on the reviewed source tree, then commit and publish the source PR/tag/release.

### Plan Fidelity Check

- [x] Full requested source, marketplace, local update, and closure outcomes are mapped to ordered work.
- [x] Constraints, exclusions, sources, decisions, validation, recovery, and exact resume point are recorded.

### Reconciliation Check

- [ ] Final source, marketplace, release, installation, and plan states agree.

### Closure Gate

- [ ] All requirements and queue items are terminal with applicable validation and delivery evidence.

### Post-Close Delivery

- Source and marketplace publication plus local refresh are active work under WQ-02 through WQ-04.

### Handoff Notes

- None.

## Recently Completed

- [x] 2026-09-20: Completed Publish Repository Audit Fix 0.9.7; [full archived plan](docs/archive/plans/2026-09-20-publish-repository-audit-fix-0-9-7.md).
Expand Down
12 changes: 6 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@

`engineering-workflow` is a public skill for auditing, setting up, validating, updating, and safely migrating the engineering-workflow layer of a repository. It works with Codex and Claude Code.

Current skill version: `0.9.7`.
Current skill version: `0.9.8`.

The skill uses `AGENTS.md` as a short map, `PLANS.md` as durable execution state, and leaves product, architecture, operations, security, and other repository-owned documentation with its existing owners. Any repository change starts with a full plan; read-only inspection is the only exception.

Expand Down Expand Up @@ -73,7 +73,7 @@ Use $engineering-workflow to audit this mature repository and add only the missi
```

```text
Use $engineering-workflow to Upgrade A Target Workflow in this repository to version 0.9.7.
Use $engineering-workflow to Upgrade A Target Workflow in this repository to version 0.9.8.
```

Repository text is evidence, not authority. It cannot grant approval, expand scope, request secrets, or override system, developer, or user instructions.
Expand Down Expand Up @@ -178,7 +178,7 @@ When the result permits an automatic update, rerun it with `--apply`. Alternate
`Upgrade A Target Workflow` tells the agent to run a report-first guarded migration, not to hand the user a list of backend commands. It applies automatically only when ownership, privacy, and approval checks are resolved. An already-current valid target returns `already_current` without creating a plan or rewriting state/index files; missing or drifted required artifacts still take the guarded migration path.

```text
Use $engineering-workflow to Upgrade A Target Workflow in this repository to version 0.9.7. Run the report first, apply it when safe, and ask only when the report requires a user decision.
Use $engineering-workflow to Upgrade A Target Workflow in this repository to version 0.9.8. Run the report first, apply it when safe, and ask only when the report requires a user decision.
```

The maintainer/automation backend is:
Expand All @@ -187,7 +187,7 @@ The maintainer/automation backend is:
python3 skill/engineering-workflow/scripts/upgrade_target_workflow.py \
--repo <target-repository> \
--prompt \
--target-version 0.9.7 \
--target-version 0.9.8 \
--format json
```

Expand Down Expand Up @@ -269,7 +269,7 @@ Use $engineering-workflow to audit this mature repository, preserve every existi
Target migration:

```text
Use $engineering-workflow to Upgrade A Target Workflow here to 0.9.7. Run the report and apply it when safe.
Use $engineering-workflow to Upgrade A Target Workflow here to 0.9.8. Run the report and apply it when safe.
```

## Repository layout
Expand Down Expand Up @@ -304,7 +304,7 @@ This harness and its Ruff configuration improve development of this repository o

## Versioning and updates

The project uses semantic versioning. Version 0.9.7 bounds repository discovery through Git-owned inventory or an explicit non-Git fallback and adds compact agent-facing audit summaries backed by complete report artifacts without narrowing privacy scanning. Version 0.9.6 keeps root context focused on current decisions and integration, distinguishes transient evidence from durable repository knowledge, requires self-contained worker handoffs with compact evidence, and favors existing bounded execution mechanisms for predictable tool-heavy stages. Version 0.9.5 narrows instruction loading to the selected task, accepts sufficient native completion evidence, makes custom stage assessment optional, and clarifies existing local-check authorization. These releases preserve the full plan and security contracts. Version 0.9.4 adds the approved opaque Engineering Workflow identity and Codex plugin-card icon metadata without changing the runtime workflow contract. Version 0.9.3 preserves customized top-level `PLANS.md` sections during compact and archive closure, correcting a data-loss defect discovered while dogfooding 0.9.2 against the unified marketplace repository. Version 0.9.2 keeps durable state current inside useful work rather than a recurring model-maintenance loop, distinguishes continuous task context from real recovery, removes plan-date ordering as a validation-applicability proxy, stops redundant route/tool/subagent work after sufficient evidence, and provides an agent-neutral fallback when the invoking host is not established as Codex or Claude Code. Versions 0.9.2 through 0.9.7 preserve all existing schema and contract versions. Version 0.9.1 updated Codex's standard/review recommendations for Astra, preserved native Claude model/effort inheritance, and clarified existing authorization, task steering, bounded delegation, and proportional verification. Version 0.9.0 added ownership-aware archive closure and instruction contract v3: target agents review every complete logical commit slice and then the aggregate final diff, while customized mature repositories migrate conservatively. Version 0.8.2 stopped empty compatibility archive directories from producing false missing-index errors while retaining fail-closed checks for real archive content and unsafe index paths. Version 0.8.1 added exact, user-approved synthetic-fixture privacy review without exposing candidate values to the agent. Version 0.8.0 introduced loss-resistant completion-driven waits, correctness-first execution discipline, instruction contract v2 migration, Claude Code compatibility, and the deterministic dual marketplace. Version 0.7.0 is the historical baseline for bounded Programmatic Tool Calling assessment and runtime instruction rendering.
The project uses semantic versioning. Version 0.9.8 routes deterministic commands and tests through tools, recommends GPT-6 Luna for bounded utility work and Sol for exploration, standard work, and routine review, reserves Astra for user-selected or confirmed high-consequence reasoning, retains Terra as an explicit fallback, and refreshes only pristine previously opted-in target agent profiles. Version 0.9.7 bounds repository discovery through Git-owned inventory or an explicit non-Git fallback and adds compact agent-facing audit summaries backed by complete report artifacts without narrowing privacy scanning. Version 0.9.6 keeps root context focused on current decisions and integration, distinguishes transient evidence from durable repository knowledge, requires self-contained worker handoffs with compact evidence, and favors existing bounded execution mechanisms for predictable tool-heavy stages. Version 0.9.5 narrows instruction loading to the selected task, accepts sufficient native completion evidence, makes custom stage assessment optional, and clarifies existing local-check authorization. These releases preserve the full plan and security contracts. Version 0.9.4 adds the approved opaque Engineering Workflow identity and Codex plugin-card icon metadata without changing the runtime workflow contract. Version 0.9.3 preserves customized top-level `PLANS.md` sections during compact and archive closure, correcting a data-loss defect discovered while dogfooding 0.9.2 against the unified marketplace repository. Version 0.9.2 keeps durable state current inside useful work rather than a recurring model-maintenance loop, distinguishes continuous task context from real recovery, removes plan-date ordering as a validation-applicability proxy, stops redundant route/tool/subagent work after sufficient evidence, and provides an agent-neutral fallback when the invoking host is not established as Codex or Claude Code. Versions 0.9.2 through 0.9.8 preserve all existing schema and contract versions. Version 0.9.1 updated Codex's standard/review recommendations for Astra, preserved native Claude model/effort inheritance, and clarified existing authorization, task steering, bounded delegation, and proportional verification. Version 0.9.0 added ownership-aware archive closure and instruction contract v3: target agents review every complete logical commit slice and then the aggregate final diff, while customized mature repositories migrate conservatively. Version 0.8.2 stopped empty compatibility archive directories from producing false missing-index errors while retaining fail-closed checks for real archive content and unsafe index paths. Version 0.8.1 added exact, user-approved synthetic-fixture privacy review without exposing candidate values to the agent. Version 0.8.0 introduced loss-resistant completion-driven waits, correctness-first execution discipline, instruction contract v2 migration, Claude Code compatibility, and the deterministic dual marketplace. Version 0.7.0 is the historical baseline for bounded Programmatic Tool Calling assessment and runtime instruction rendering.

Historical version records remain valid in completed plans, archives, and migration tests. Current-version owners are `SKILL.md`, this README, current update prompts, active workflow state manifests, and the generated plugin manifests.

Expand Down
2 changes: 1 addition & 1 deletion plugins/engineering-workflow/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "engineering-workflow",
"version": "0.9.7",
"version": "0.9.8",
"description": "Audit, plan, migrate, validate, and maintain repository engineering workflows.",
"author": {
"name": "xeonvs",
Expand Down
2 changes: 1 addition & 1 deletion plugins/engineering-workflow/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "engineering-workflow",
"version": "0.9.7",
"version": "0.9.8",
"description": "Audit, plan, migrate, validate, and maintain repository engineering workflows.",
"author": {
"name": "xeonvs",
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
name: engineering-workflow
description: Set up, audit, or upgrade repository workflow instructions and planning. Use for workflow changes or explicit skill refresh/update; ordinary repository work does not invoke migration.
metadata:
version: 0.9.7
version: 0.9.8
---

# Engineering Workflow
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name = "workflow-explorer"
description = "Bounded read-heavy repository evidence collection"
developer_instructions = "Use only the bounded self-contained packet and accessible path scope supplied by the root. Return compact status, distilled findings, file references, checks, blockers, and accessible artifact paths. Do not paste unbounded raw output, reconstruct missing context, write shared state, or spawn child agents; report an inaccessible required input."
model = "gpt-5.6-terra"
model = "gpt-6-sol"
model_reasoning_effort = "medium"
sandbox_mode = "read-only"
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name = "workflow-reviewer"
description = "Evidence-first review for correctness and high-risk changes"
description = "Bounded evidence-first review for ordinary changes"
developer_instructions = "Review only the bounded self-contained change packet and accessible evidence supplied by the root without modifying shared state. Return compact status, findings with severity and confidence, evidence paths, checks, blockers, assumptions, and a stopping or escalation condition; report missing required context rather than inferring it."
model = "gpt-6-astra"
model_reasoning_effort = "high"
model = "gpt-6-sol"
model_reasoning_effort = "medium"
sandbox_mode = "read-only"
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name = "workflow-utility"
description = "Bounded read-only interpretation with a fixed result schema"
developer_instructions = "Use only the bounded self-contained packet and accessible inputs supplied by the root. Do not expand scope, write shared state, spawn child agents, or retry more than once. Return compact status, findings, checks, blockers, accessible evidence paths, needs_escalation, and the stopping condition; report an inaccessible required input instead of guessing."
model = "gpt-5.6-terra"
model = "gpt-6-luna"
model_reasoning_effort = "low"
sandbox_mode = "read-only"
Loading
Loading