diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 1744624..95fa560 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -89,7 +89,7 @@ and route tests do not establish production or human acceptance. ## Coding Conventions -These conventions come from the project's `CLAUDE.md` and must be followed in all contributions: +These conventions must be followed in all contributions: - **Type hints on all functions** — parameters and return types, always. Use `from __future__ import annotations` at the top of each module. - **f-strings** — use f-strings for string interpolation, not `%` formatting or `.format()`. @@ -123,7 +123,6 @@ if TYPE_CHECKING: class YourDimensionAnalyzer(BaseAnalyzer): name = "your_dimension" - weight = 0.05 # fraction of overall completeness score def analyze( self, @@ -160,7 +159,7 @@ ALL_ANALYZERS = [ ### Step 3 — Add a weight in the scorer -Open `src/github_repo_auditor/scorer.py` and add your dimension name to the `WEIGHTS` dict. Weights must sum to `1.0` after adding the new entry, so adjust existing weights proportionally. +For a completeness-scored dimension, open `src/github_repo_auditor/scorer.py` and add your dimension name to the `WEIGHTS` dict. Interest is scored separately; advisory dimensions such as `description` remain unweighted. Weights must sum to `1.0` after adding the new entry, so adjust existing weights proportionally. ### Step 4 — Write tests @@ -180,7 +179,7 @@ Before opening a PR, verify: - [ ] `make lint` reports no errors. - [ ] `make type-check` reports no errors (or pre-existing errors only — do not introduce new ones). - [ ] No hardcoded GitHub usernames or API tokens anywhere in the diff. -- [ ] New analyzer (if any) is registered in `ALL_ANALYZERS` and has a weight in `WEIGHTS`. +- [ ] New analyzer (if any) is registered in `ALL_ANALYZERS` and, if completeness-scored, has a weight in `WEIGHTS`. - [ ] New tests added for any new public functions or analyzer logic. - [ ] `CHANGELOG.md` updated under `## [Unreleased]` with a brief description of the change. - [ ] Commit messages follow conventional commits: `feat:`, `fix:`, `chore:`, `refactor:`, `test:`, `docs:`. diff --git a/README.md b/README.md index 86069b2..e441f7b 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Case study — [Operator OS: a multi-agent control plane over a repo portfolio](CASE-STUDY.md).** How this auditor's truth layer anchors six local services and two coordinated coding agents (Claude Code + Codex), with real portfolio metrics and a [90-second demo plan](DEMO-PLAN.md). -GitHub Repo Auditor is a portfolio audit and operator tool for developers with a lot of repositories. It clones every repo on your GitHub account, runs 12 analyzers across completeness and interest dimensions, assigns letter grades and achievement badges, preserves historical state, and generates actionable dashboards you can actually use to decide what to work on next. Built for developers who ship fast, start often, and need a system to manage the sprawl. +GitHub Repo Auditor is a portfolio audit and operator tool for developers with a lot of repositories. It clones every repo on your GitHub account, runs 13 analyzers across completeness, interest, and advisory dimensions, assigns letter grades and achievement badges, preserves historical state, and generates actionable dashboards you can actually use to decide what to work on next. Built for developers who ship fast, start often, and need a system to manage the sprawl. Today the project is best understood as a GitHub portfolio operating system: @@ -131,7 +131,7 @@ Treat campaign/writeback, GitHub Projects, Notion sync, catalog overrides, score ## Features -- **12 Analyzers** — README quality, test coverage, CI/CD, dependency freshness, commit patterns, bus factor, code complexity, security controls, license, build readiness, GraphQL signals, and more +- **13 Analyzers** — README quality, test coverage, CI/CD, dependency freshness, commit patterns, bus factor, code complexity, security controls, license, build readiness, GraphQL signals, and more - **Dual-Axis Scoring** — Completeness (does this project have what shipped software should?) and Interest (is this worth anyone's time?) scored independently on 0.0–1.0 scales - **Letter Grades + Tier Classification** — A–F grades with Shipped / Functional / WIP / Skeleton / Abandoned tiers; 15 achievement badges ("Fully Tested", "CI Champion", "Zero Debt", etc.) - **Quick Wins Engine** — For each repo, shows exactly which single action moves it to the next tier and how far it is from getting there @@ -214,8 +214,8 @@ Expected outputs include `output/demo/demo-report.json`, `output/demo/operator-control-center-demo.json`, `output/demo/operator-control-center-demo.md`, `output/demo/portfolio-truth-latest.json`, -`output/demo/weekly-command-center--.json`, -`output/demo/security-burndown--.json`, +`output/demo/weekly-command-center-demo.json`, +`output/demo/security-burndown-demo.json`, `output/demo/pending-proposals.json`, and `output/demo/portfolio-warehouse.db`. To browse the same fixture in the local web UI: @@ -279,7 +279,7 @@ audit report --portfolio-truth # Semantic search across the portfolio index audit triage --ask "Python projects with no tests" -# Weekly operator briefing (requires Anthropic API key) +# Weekly operator briefing (LLM suggestions are optional) audit run --briefing # Deep Dive — targeted repo rerun merged into the latest baseline @@ -302,7 +302,7 @@ and `Safe to Defer`, and writes `operator-control-center--.json` `.md`. `audit triage --approval-center` is also read-only. It loads the latest approval history, -groups work into `Needs Re-Approval`, `Ready For Review`, `Approved But Manual`, and +groups work into sections including `Needs Re-Approval`, `Ready For Review`, `Approved But Manual`, and `Blocked`, and writes `approval-center--.json` plus `.md`. Local approval capture stays separate from writeback apply. @@ -350,7 +350,7 @@ smokes. [docs/release-gates.md](docs/release-gates.md) retains the release gates ## Architecture -The auditor follows a pipeline architecture: fetch repo list via GitHub API → shallow-clone each repo → run all 12 analyzers in sequence → aggregate scores → generate outputs. Analyzers are pluggable via `--analyzers-dir` for custom extensions. The scoring engine computes completeness and interest independently, applies configurable scoring profiles, and derives letter grades from the combined result. All output writers (Excel, HTML, JSON, Markdown, Notion) are isolated from the analysis layer and consume the same scored result object. Workbook ranking and trend views always use the full filtered portfolio baseline, even for targeted or incremental reruns. +The auditor follows a pipeline architecture: fetch repo list via GitHub API → shallow-clone each repo → run all 13 analyzers in sequence → aggregate scores → generate outputs. Analyzers are pluggable via `--analyzers-dir` for custom extensions. The scoring engine computes completeness and interest independently, applies configurable scoring profiles, and derives separate letter grades from the completeness and interest scores. All output writers (Excel, HTML, JSON, Markdown, Notion) are isolated from the analysis layer and consume the same scored result object. Workbook ranking and trend views always use the full filtered portfolio baseline, even for targeted or incremental reruns. Partial reruns now require a compatible full-baseline report, not just any previous report. The stored baseline contract tracks the audit-affecting portfolio context used to produce the last trustworthy baseline, and targeted or incremental reruns will fail closed if that contract no longer matches the current request. @@ -423,7 +423,7 @@ That command generates stable sample `standard` and `template` workbooks, valida After that manual desktop Excel check, record the outcome back into the gate artifacts: ```bash -make workbook-signoff ARGS="--reviewer yourname --outcome passed --check excel-open-no-repair=passed --check visible-tabs-present=passed --check normal-zoom-readable=passed --check chart-placement-clean=passed --check filters-work=passed" +make workbook-signoff ARGS="--reviewer yourname --outcome passed --check excel-open-no-repair=passed --check visible-tabs-present=passed --check normal-zoom-readable=passed --check chart-placement-clean=passed --check filters-work=passed --check core-navigation-links-work=passed --check operator-story-consistent=passed --check repo-detail-selector-works=passed --check run-changes-readable=passed" ``` ## Managed Campaigns and Governance @@ -456,7 +456,7 @@ The daily operator loop is now: - Run `make workbook-signoff ...` after the manual Excel-open check for workbook-facing changes - Browse [http://127.0.0.1:8080/](http://127.0.0.1:8080/) after `audit serve` to review the dashboard -Scheduled automation stays artifact-first. The weekly workflow now runs the audit, generates a control-center artifact plus a scheduled handoff summary, uploads `output/`, opens or updates one canonical GitHub issue only when blocked or urgent operator findings cross a meaningful threshold, and closes that same issue cleanly when later runs return to a quiet state. The handoff now also calls out whether the queue is getting better, worse, or staying stuck, what was tried most recently, whether that intervention actually helped, whether recovery is only quiet for now or confirmed resolved, whether recent high-confidence guidance has been validating or turning noisy, what trust policy now applies to the live recommendation (`act-now`, `act-with-review`, `verify-first`, or `monitor`), whether a soft exception or recent policy-flip drift should make the operator treat that recommendation more cautiously, and whether recent soft caution is still earning trust or has become cautious enough to recover toward a stronger policy. +Scheduled automation stays artifact-first. The manually dispatched workflow now runs the audit, generates a control-center artifact plus a scheduled handoff summary, uploads `output/`, opens or updates one canonical GitHub issue only when blocked or urgent operator findings cross a meaningful threshold, and closes that same issue cleanly when later runs return to a quiet state. The handoff now also calls out whether the queue is getting better, worse, or staying stuck, what was tried most recently, whether that intervention actually helped, whether recovery is only quiet for now or confirmed resolved, whether recent high-confidence guidance has been validating or turning noisy, what trust policy now applies to the live recommendation (`act-now`, `act-with-review`, `verify-first`, or `monitor`), whether a soft exception or recent policy-flip drift should make the operator treat that recommendation more cautiously, and whether recent soft caution is still earning trust or has become cautious enough to recover toward a stronger policy. In newer follow-through phases, that same weekly story also carries whether a recommendation is escalating, recovering, rebuilding, re-acquiring confidence, or aging back down. The important product principle is still the same: workbook, HTML, Markdown, and review-pack surfaces should tell the same story in different formats. diff --git a/docs/architecture.md b/docs/architecture.md index 4d6116b..d0a9614 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -36,7 +36,7 @@ That means the architecture is intentionally split between raw state assembly an ### `audit` -`src/github_repo_auditor/cli.py` remains the single entrypoint. It keeps one flag-based command surface and packages the product around four guidance modes: +`src/github_repo_auditor/cli.py` remains the single entrypoint. It exposes workflow and gate subcommands alongside the legacy flag-based command surface and packages the product around four guidance modes: - `First Run` - `Weekly Review` @@ -449,7 +449,7 @@ Visible terminology can be cleaned up, but stored keys and historical loading pa The most important current paths are: ```text -src/ +src/github_repo_auditor/ cli.py reporter.py review_pack.py diff --git a/docs/audit-cli-migration.md b/docs/audit-cli-migration.md index 52a5454..7201e3b 100644 --- a/docs/audit-cli-migration.md +++ b/docs/audit-cli-migration.md @@ -36,7 +36,7 @@ Flags that belong here: `--repos`, `--skip-forks`, `--skip-archived`, `--skip-cl `--scoring-profile`, `--watch`, `--resume`, `--vuln-check`, `--reindex`, `--embedder`. -Global flags available in all subcommands: `--token`, `--output-dir`, `--config`, +Global flags available in the four workflow subcommands: `--token`, `--output-dir`, `--config`, `--verbose`. ### audit triage diff --git a/docs/audit-serve.md b/docs/audit-serve.md index 082ad6a..6ba76d3 100644 --- a/docs/audit-serve.md +++ b/docs/audit-serve.md @@ -3,7 +3,7 @@ `audit serve` starts a local FastAPI + progressively enhanced web interface over your latest audit output. It is a read-mostly operator tool: you can browse portfolio state, per-repo history, run history, and the approval queue, and you can trigger new audit runs through a form. -It binds to `127.0.0.1` only and requires no authentication — treat it as a local-only +It binds to `127.0.0.1` by default and requires no authentication — treat it as a local-only tool for solo operator use. ## Installation @@ -45,9 +45,9 @@ Full flag reference (`audit serve --help`): | `--port PORT` | `8080` | Port to listen on | | `--host HOST` | `127.0.0.1` | Interface to bind (do not change to `0.0.0.0`) | | `--output-dir DIR` | `./output` | Directory where audit output files live | -| `--config PATH` | `./audit-config.yaml` | Path to audit config file | -| `--verbose` | off | Print detailed output | -| `--token TOKEN` | `$GITHUB_TOKEN` | GitHub token forwarded to triggered runs | +| `--config PATH` | unset | Accepted CLI option; not forwarded to triggered runs | +| `--verbose` | off | Accepted CLI option; not used by the web launcher | +| `--token TOKEN` | `$GITHUB_TOKEN` or `gh auth token` | Accepted CLI option; not forwarded to triggered runs, which resolve their own credentials | Once started, open `http://127.0.0.1:8080/` in your browser. The server runs until you press Ctrl-C. @@ -86,8 +86,7 @@ run timestamp, username, repo count, portfolio grade, and any run-level notes. ### `GET /approvals` -Approval queue. Reads the latest approval-center state and renders open items grouped by -status (`needs-reapproval`, `ready-for-review`, `approved-manual`, `blocked`). Approve +Approval queue. Reads persisted approval records from the warehouse and renders them in a table. Approve and reject buttons submit via HTMX and record intent locally — they do not trigger writeback automatically. @@ -135,7 +134,7 @@ completed, failed, cancelled, disconnected, and recovered states. - **No authentication.** The UI is designed for single-user local use only. Do not expose it on a non-loopback interface or behind a shared reverse proxy without adding your own auth layer. -- **Binds to `127.0.0.1` only.** The default host is intentionally loopback. Changing +- **Defaults to `127.0.0.1`.** The default host is intentionally loopback. Changing `--host` to `0.0.0.0` is unsupported and not recommended. - **Not for multi-user environments.** The approval intent log and run session registry are in-memory or local-file only; there is no multi-user isolation. diff --git a/docs/demo-proof/public-fixture/README.md b/docs/demo-proof/public-fixture/README.md index 1110bf1..4de674a 100644 --- a/docs/demo-proof/public-fixture/README.md +++ b/docs/demo-proof/public-fixture/README.md @@ -33,7 +33,10 @@ Expected generated artifacts: The truth artifacts are regenerated on every run: the schema version comes from the producer constant and the timestamp is computed at generation time, so the demo always reflects the current contract. `validate_proof_package.py` fails the -package if either drifts. +package if either drifts. The committed proof manifest still declares schema +`0.11.0`, while the producer emits `0.12.0`; `make demo` does not update that +manifest. Validation requires its schema metadata to be reconciled with the +current producer as well as fresh generated artifacts. ## Desktop Demo diff --git a/docs/extending-analyzers.md b/docs/extending-analyzers.md index 0a5dcf7..405f381 100644 --- a/docs/extending-analyzers.md +++ b/docs/extending-analyzers.md @@ -29,19 +29,19 @@ An analyzer should inspect one aspect of a repo and return: In addition to existing README quality fields, `ReadmeAnalyzer` now produces: - `readme_last_touched_days` — days since the README file was last modified, based on Git history -- `code_last_touched_days` — days since any non-README file in the repo was last modified -- `readme_staleness_ratio` — `readme_last_touched_days / code_last_touched_days`; higher means the README is aging faster than the code -- `readme_stale` — boolean; `true` when `readme_staleness_ratio > 5.0` AND `code_last_touched_days < 90`, i.e., the README is more than five times older than the code and the code is still being actively touched +- `code_last_touched_days` — days since a code file matching `_CODE_GLOBS` in `readme.py` was last committed +- `readme_staleness_ratio` — `readme_last_touched_days / max(code_last_touched_days, 1)`; higher means the README is aging faster than the code +- `readme_stale` — boolean or null when either age is unknown; `true` when `readme_staleness_ratio > 5.0` AND `code_last_touched_days < 90`, i.e., the README is more than five times older than the code and the code is still being actively touched Excel and control-center surfacing for these staleness fields is wired via S2.4. ### ActivityAnalyzer -`ActivityAnalyzer` now produces release signal fields via `GithubClient.get_releases()`: +`ActivityAnalyzer` now produces release signal fields via `GitHubClient.get_releases()`: - `has_any_release` — boolean; whether the repo has at least one published release - `release_count` — total number of releases fetched (capped at 10 per run) -- `releases_available` — whether the releases endpoint was reachable +- `releases_available` — false on HTTP 404; other HTTP errors leave it true with an empty release list - `latest_release_age_days` — days since the most recent release was published - `latest_release_is_prerelease` — boolean; whether the most recent release is marked as a pre-release @@ -63,11 +63,12 @@ Inputs that are **not** stable (do not include): wall-clock time, run-specific I ### Current opt-ins -Three analyzers currently opt in: +Four analyzers currently opt in: - `DependenciesAnalyzer` — hashes lockfile bytes - `ReadmeAnalyzer` — hashes README content + git timestamps - `StructureAnalyzer` — hashes sorted directory listing + primary language +- `DescriptionAnalyzer` — hashes description + primary language + sorted topics ### Validating correctness with `--reconcile-cache` diff --git a/docs/modes.md b/docs/modes.md index 846b211..50fe7c4 100644 --- a/docs/modes.md +++ b/docs/modes.md @@ -4,7 +4,8 @@ GitHub Repo Auditor now works best when you think about it as one operating syst four product modes. The flags stay the same underneath; this guide is the shared map for the docs, CLI help, workbook, HTML, Markdown, review-pack, and scheduled-handoff wording. -As of Arc F Sprint 4.3 the CLI has four subcommands (`run`, `triage`, `report`, `serve`). +The CLI has four workflow subcommands (`run`, `triage`, `report`, `serve`) plus +`security-burndown`, `security-gate`, and `pr-evidence`. Examples below show the subcommand form first. The flat form (`audit --flag`) still works and shows a deprecation warning. See [docs/audit-cli-migration.md](audit-cli-migration.md) for the full mapping. @@ -224,7 +225,7 @@ Each section should answer the same three questions quickly: ## Subcommand flag reference -The four subcommands group flags by workflow. All subcommands accept the shared globals +The four workflow subcommands group flags by workflow and accept the shared globals `--token`, `--output-dir`, `--config`, and `--verbose`. ### audit run diff --git a/docs/operator-troubleshooting.md b/docs/operator-troubleshooting.md index 414edb0..037ba3f 100644 --- a/docs/operator-troubleshooting.md +++ b/docs/operator-troubleshooting.md @@ -151,7 +151,7 @@ That command generates canonical sample `standard` and `template` workbooks, val After the manual Excel-open check, record the signoff: ```bash -make workbook-signoff ARGS="--reviewer --outcome passed --check excel-open-no-repair=passed --check visible-tabs-present=passed --check normal-zoom-readable=passed --check chart-placement-clean=passed --check filters-work=passed" +make workbook-signoff ARGS="--reviewer --outcome passed --check excel-open-no-repair=passed --check visible-tabs-present=passed --check normal-zoom-readable=passed --check chart-placement-clean=passed --check filters-work=passed --check core-navigation-links-work=passed --check operator-story-consistent=passed --check repo-detail-selector-works=passed --check run-changes-readable=passed" ``` ## Scheduled handoff issues diff --git a/docs/proof-package-contract.md b/docs/proof-package-contract.md index 0b318da..44f6997 100644 --- a/docs/proof-package-contract.md +++ b/docs/proof-package-contract.md @@ -95,6 +95,7 @@ Run: python scripts/validate_proof_package.py docs/proof-packages//proof-package.json ``` -The validator checks structure and local file references. It intentionally does -not judge whether a claim is true; the package author must still choose honest +The validator checks structure, local file references, and portfolio-truth schema +and freshness against the current producer contract. It does not otherwise +judge whether a claim is true; the package author must still choose honest claim statements and bounded evidence.