Give Codex a map of the tools and skills it already has.
A useful skill can be buried in your setup. JevCompass discovers installed candidates locally, narrows them with a reviewed catalog, and offers a short suggestion when a task has a clear match. It works with Codex Desktop and CLI hooks, without replacing your agent or gating everyday commands.
- Start fast: try explicit advice from the terminal before installing any hooks.
- Keep control: suggestions never run tools, grant permissions, or override required checks.
- Choose your privacy level: local advice works without a key; optional Jev ranking sees only coarse candidate metadata through OpenRouter.
Get started · See real output · Understand the boundary · Check the evidence · Contribute
Requirements: Python 3.11 or newer, pipx, and a Codex installation with hooks enabled. Install the published package on PyPI:
pipx install jevcompass
jevcompass doctorStart with jevcompass doctor to see which representative tasks have enough locally available candidates for a recommendation. Before hook installation it may exit nonzero because the hooks are not registered yet; the capacity report is still useful. Then try a manual request only when the report shows a useful choice for that task; for example, run jevcompass recommend --category coding --domain python when coding_python reports decision candidates or local candidates. A clean profile may report low signal skip or silent, in which case that request can correctly return no recommendation.
To install from source instead, clone this repository and run pipx install . from the checkout. For the verified release, use pipx install jevcompass==0.1.23; check releases for the latest version. Version 0.1.23 publishes the three explicit coding-decision commands (strategy choose, tests discover/rank, and triage with diagnostic costs), the opt-in typed-decision cache, and a fifth bundled skill; per-change details and evidence boundaries are in CHANGELOG.md. The release revision passes the full offline suite (858 Python tests on Linux), and the release pull request and publish workflow re-run the suite on Python 3.11 with exact-tag distribution checks. One earlier isolated CLI session on published 0.1.22 reported exec_command and unittest before first tool; no broader host-level impact is measured and no speed or quality benefit is established.
The supplemental CI-guided coding pilot exposed a measurement gap: older receipts cannot distinguish exact, equivalent or chained test commands. A corrected parser separately records the exact CI invocation and the same suite without -v; in a new pair JevCompass ran the exact CI command while baseline ran the equivalent suite. The advised arm was slower, so no efficiency gain is established. Pilot evidence has the blind receipts and limits.
When a coding task already states an exact test command, the advisor skips duplicate test-selection hints. For a Python repository whose CI clearly runs unittest, the catalog can offer that runner for coding tasks without a prescribed test command. This relevance change is not a measured speed gain. Pilot evidence records the control pair.
The published advice wording remains unchanged. An experimental compact variant shortened one synthetic context but was slower than no advisor in one quality-tied P01 pair; it is not a recommended efficiency setting. Pilot evidence records the limits.
The recommendation is environment-dependent: it can show local advice, an unranked shortlist, or no recommendation. We are testing whether advice improves completion time or token usage versus the same Codex task without it. A single keyless CLI P08 surrogate pair tied on quality and saw a shorter treatment turn. In a separate simple P01/P03 repeat, both arms completed each task and required command; treatment with local advice was slower on P01, while treatment without advice was faster on P03. These small mixed results do not prove an efficiency gain. Pilot evidence separates observed process time, root-turn token counters, and human-rated first productive action; no billing savings are claimed. The manual command uses explicit category/domain metadata and never takes a task prompt.
To enable automatic advice, first inspect what would change, then install:
jevcompass install --dry-run
jevcompass install
jevcompass doctorWhat changes: jevcompass install backs up an existing Codex hooks.json and registers two advisory hooks. Review and trust the registrations in Codex's /hooks screen, then start a fresh session. The default config is under ~/.codex; set CODEX_HOME to use a different Codex home. A successful doctor check confirms local registration, not that Codex loaded or delivered advice. Recent hook metrics are matched to this Codex profile using a short local hash; records without a profile marker do not count as evidence for a fresh profile.
Experimental and explicitly opt-in: versions before 0.1.16 do not include this command. If doctor reports coding_python: low signal skip, you can preview the optional bundled skills before deciding whether to install them:
jevcompass skills install --dry-runThe dry run does not install anything, and installing skills does not guarantee a recommendation. Install only if you choose to:
jevcompass skills install
jevcompass doctorThe installer copies five skills: jevcompass-focused-tests, jevcompass-regression-review, jevcompass-plan-implementation, jevcompass-plan-cutover, and jevcompass-coding-workflow. Published v0.1.23 contains all five; the published v0.1.22 package contained the original two coding skills, so check jevcompass skills install --dry-run for the version you installed. The planning skills cover an ordinary implementation breakdown and a migration or cutover with rollback, respectively; they are distinct optional candidates, and Jev can compare them only when both are installed. With the default profile it uses ~/.agents/skills; with CODEX_HOME set it uses that profile's skills/ directory. It never changes hooks, executes scripts, calls Jev, or overwrites a different skill with the same name. Identical repeats are no-ops. Read the bundled SKILL.md files before enabling them, and start a fresh Codex session if they do not appear. Remove the installed skill directories yourself after checking their contents if you no longer want them.
Suggestions for these skills are conditional on the task matching their guidance in SKILL.md; a suggestion is not a requirement. This remains an experiment with no demonstrated speed or quality benefit. An installed v0.1.22 CLI surrogate for project setup delivered local exec_command/git advice before the first tool, but tied on blind scaffold quality. A subsequent paired CLI repeat retained a recognized successful unittest in both arms; it does not verify the planned Desktop case or show a performance benefit. See pilot evidence. A separate API-contract skill prototype was withheld after a opt-in blinded P07 trial in which the baseline authored the requested contract and the advised arm did not; see pilot evidence. It is not part of the installer. In one source-built C05 pair, the agent read the suggested create-plan skill, but the blinded scores tied 6/7. This single result does not establish causality or effectiveness. That C05 pair did not test Desktop. Separate named-role Desktop worker smokes confirmed advice delivery and one actual focused-skill read after opt-in installation, without measuring a speed or quality gain. The skills and installer are included since v0.1.16 but are not installed unless you opt in. They are absent from v0.1.15.
JevCompass can attach conservative pretask strategy guidance to UserPromptSubmit when you explicitly install it with jevcompass install --strategy-advice. Reinstalling without a flag preserves the setting; use --disable-strategy-advice to turn it off. Ordinary installs remain unchanged. The hook considers only coding prompts with narrow, explicit request wording. These signals describe what the prompt asks for, not verified repository or test facts; prompts, paths and source text are not sent to the strategy API. The strategy selector is the only remote-choice call on this route; catalog advice remains local. Cached selections are labeled and do not make a new request. If the prompt is outside scope, local classification is uncertain, or selection fails, the existing advice path remains available. This opt-in delivery has not been evaluated for coding benefit.
For a substantial coding task with genuinely competing approaches, use only coarse signals you verified locally. Skip this command for small, obvious edits: one blinded simple-task pair tied on observable quality while the explicit local strategy call added 6.81 seconds and more tokens; Jev did not select a strategy in that pair. A richer retry-contract pair received a real Jev choice, but treatment scored lower and took longer; see PILOT.md. Keep strategy selection experimental. Version 0.1.23 exposes jevcompass strategy choose --kind coding --signal existing_symbol --signal behavior_change --json. It returns up to two reviewed strategies and labels whether Jev selected the first (remote-choice) or a deterministic local order applied (no-remote-choice). It accepts no prompt, path, source, or log text. With one clear strategy, unavailable Jev, or uncertain output, it avoids a remote choice. This explicit command is published since v0.1.23; no paired task gain has been measured.
If verified local evidence already determines the approach, continue directly rather than calling a selector to confirm it. Integrations that already carry that local decision can pass resolved_strategy to the Python API or --resolved-strategy to the CLI. For example, jevcompass strategy choose --kind coding --signal existing_symbol --signal behavior_change --resolved-strategy define_contract_then_implement --json returns only that eligible reviewed strategy, without constructing a decision client or making an API request. Unknown or ineligible explicit resolutions abstain locally. This option does not establish that the caller's evidence is correct, grant permissions or waive tests; do not use it to disguise unresolved decisions.
The API also accepts contract_evidence, and the CLI accepts --contract-evidence with consistent, conflicting, absent, partial, or unknown. Supply this only after inspecting the relevant contract, tests and callers locally. For coding with exactly existing_symbol and behavior_change, consistent evidence selects inspection locally; conflicting or absent evidence selects contract definition locally. These routes make no API request. Partial evidence can accompany an already eligible unresolved choice; it does not force a remote call. Unknown evidence preserves the existing fallback. An explicit resolved strategy that contradicts a deterministic evidence result abstains. These are conservative routing rules, not proof that the agent has interpreted the contract correctly or that the task will finish faster.
The latest compliant supervisor-prepared CLI comparison confirmed strategy advice before the first observed tool and complete focused/full validation, but treatment was 0.754 seconds slower with identical final code. See VCR-257-H in PILOT.md. Initial-prompt injection in that harness does not verify native Desktop hook delivery.
With a verified dependency_change signal, source inspection advice also includes a short local migration checklist: compare API/signature, preserve the caller contract, and check exceptions and resource ownership. This adds no provider call and is not a measured quality or speed guarantee.
JevCompass can discover Python focused checks from staged, unstaged, and untracked changes: jevcompass tests discover --required 'python -m unittest discover -s tests -v' --json from the Git root. Supply the actual required command from your repository's instructions or CI; JevCompass does not invent that gate. Discovery bounds the scan, rejects symlinks and ambiguous filename mappings, and abstains when it cannot safely associate code and tests. Unit and integration tests for the same changed module become distinct alternatives. For other languages or curated candidates, jevcompass tests rank accepts explicit local JSON metadata. It prints an order; it does not execute tests. This command is published since v0.1.23; no speed or quality improvement has been established.
{"surface":"api","candidates":[{"id":"unit","kind":"unit","command":"python -m unittest tests.test_api_unit","relevance":0.5},{"id":"contract","kind":"contract","command":"python -m unittest tests.test_api_contract","relevance":0.5}],"required":[{"id":"ci","command":"python -m unittest discover -s tests -v"}]}Save this as test-order.json, then run jevcompass tests rank --input test-order.json --json. status: remote-choice means Jev ranked genuine competing test kinds; no-remote-choice uses stable local relevance order when one choice is obvious, Jev is unavailable, or its answer is uncertain. The JSON and text output also expose a bounded decision_reason: no_choice_needed, local_resolution, invalid_response, unknown_choice, insufficient_confidence, provider_error, or accepted. Older manually constructed results may leave it unknown. This distinguishes fallback paths without exposing backend prose or changing confidence thresholds. The required list is returned unchanged and must still be run. Commands and caller IDs stay local: only coarse surface, test kinds, generic descriptors and opaque IDs may reach OpenRouter. Do not put secrets in command strings; this local file and CLI output are readable on your machine. Python filename-based discovery is published since v0.1.23; support for other layouts and matched outcome trials remain open.
Optional test-order metadata can describe verified signals (public_contract_changed, boundary_mapping_changed, internal_logic_changed) and each candidate's coverage (direct, indirect, unknown) and runtime (fast, slow, unknown). Use coverage evidence and comparable local runtime measurements; leave uncertain facts unknown. Each candidate may also include coverage_targets, a list from the same three change-signal values, after you verify which behaviors the test actually asserts. This distinguishes internal logic from a public mapping without sending test names or source. Omitted targets remain unspecified; unknown or malformed targets abstain locally without reaching OpenRouter. Targets do not establish complete coverage or make a remote call mandatory. Discovery does not invent these observations. A direct fast check against only indirect slow alternatives is chosen locally; mixed tradeoffs can remain eligible for Jev. Test commands, caller IDs and raw durations stay local, and the required validation list remains unchanged. These inputs improve the information available for selection; matched productivity gains remain unverified.
After a real test fails and you have at least two evidence-backed explanations, classify them locally and ask for a first diagnostic step. For example: jevcompass triage --exit-code 1 --kind import --hypothesis import_module_missing --hypothesis import_path_changed --json. The command accepts enums only; never pass a log, source snippet, path, or exception text. It prints at most two locally authored steps and the observed failing exit code. It does not rerun tests, execute a fix, or turn failure into success. Run a cheap read-only discriminator before requesting a remote ranking. For import failures, --import-observation package_present --import-observation target_module_absent --import-observation replacement_module_present accepts only fixed enum facts from a local check; together they rule out the missing-package hypothesis and skip the Jev call. Contradictory observations abstain. Without an ambiguous choice or confident Jev response it uses local order. This command is published since v0.1.23 and has no repeatable measured task benefit yet.
For assertion failures, the CLI also accepts verified --assertion-observation enum facts. If the expected-behavior contract is underspecified, include --hypothesis confirm_behavior_contract --assertion-observation contract_underspecified: the next step is a local authoritative-policy check, not a remote guess or an automatic repair. A legacy_fixture_conflict alone can leave competing diagnostics for Jev. Contradictory contract_confirmed and contract_underspecified observations abstain. Confirm intended semantics before editing an expectation or implementation; defer unsupported changes if no authoritative policy can be established. If your local review already established that prerequisite, perform the policy check directly and skip the extra CLI call: the first matched contract case tied on observable diagnosis quality while the explicit advisor arm took 2.501 seconds longer. See PILOT.md for its scope and limitations.
For timeouts, the CLI also accepts repeatable --timeout-observation enum facts verified locally. An unsatisfiable wait condition selects the existing wait-condition check locally; observed contention together with progress and a satisfiable wait selects the resource-contention check. Conflicting facts abstain. Partial facts can inform an already eligible ambiguous choice; they do not create an API call by themselves. These are next-check suggestions, not confirmed causes, and the original failing exit remains unchanged. Example: jevcompass triage --exit-code 1 --kind timeout --hypothesis timeout_contention --hypothesis timeout_nonterminating --timeout-observation wait_condition_unsatisfiable --json.
For unresolved diagnostics, optionally add caller-verified relative check costs, for example --diagnostic-cost timeout_contention=high --diagnostic-cost timeout_nonterminating=low. Allowed costs are low, medium, high, and unknown, keyed only to supplied hypotheses. Omit costs you cannot establish locally. Costs inform which check to try next; they do not establish causal likelihood, override local resolutions, execute checks, or remove mandatory validation. This is published since v0.1.23 with no measured speed benefit yet.
Triage JSON now includes a fixed decision_reason and each step's selection_source: a remotely preferred next action, locally resolved guidance, or an unranked local fallback. hypothesis_ranking_status remains not_established: choosing the next diagnostic action does not establish which cause is most likely. These fields are published since v0.1.23.
For an optional complete hypothesis order, add --rank-hypotheses when 2–4 locally plausible causes remain. Jev receives up to six pairwise choice questions in the same request as the separate next-step question. The CLI exposes hypothesis_order only when every pair has an allowed choice at the confidence threshold and the comparisons form an acyclic complete order. Otherwise the ranking status is incomplete or not_established, with no order inferred from caller order; the diagnostic step is validated independently. Local resolutions and unambiguous cases still skip remote calls. Opt-in ranking bypasses the choice-only typed cache. This feature does not claim improved diagnosis quality.
The built-in catalog cannot know when a private skill fits your work. Register a short, generic description explicitly, after reading its SKILL.md and checking the fields you are willing to share:
jevcompass skills add my-review-skill \
--capability 'Review local code changes' \
--use-when 'The task calls for source review' \
--avoid-when 'The task only needs implementation' \
--category review --domain softwareThe first run previews the fields and makes no change. If those fields are safe to send to OpenRouter when optional remote ranking is enabled, repeat with --approve-remote-metadata. Registration writes only your descriptions to CODEX_HOME/jevcompass/catalog.json (normally ~/.codex/jevcompass/catalog.json); it does not copy the skill body or path. Only an actually installed skill with matching name becomes a candidate. Keep descriptions generic: the skill ID, capability, use and avoid conditions may be sent to OpenRouter for ranking. Without an OpenRouter key, matching advice stays local. Remove an entry by editing that local JSON file; a changed catalog invalidates cached rankings. A suggestion never substitutes for reading the actual SKILL.md.
JevCompass works without an OpenRouter key: it can offer a local, unranked shortlist or stay silent. To let Jev rank eligible candidate choices, provide an OpenRouter API key. On a system with a supported keyring, install the optional dependency and enter the key through the hidden terminal prompt:
pipx inject jevcompass 'keyring>=25'
jevcompass auth status
jevcompass auth set
jevcompass doctorFor a new installation, pipx install 'jevcompass[secure-store]' includes the keyring dependency. If no supported keyring is available, supply OPENROUTER_API_KEY to the Codex process through your existing secret manager. Never paste a literal key into a hook command, repository file, or shell history. The environment variable takes precedence over the keyring; jevcompass auth status reports the source without printing the secret. jevcompass auth delete removes only the keyring entry after interactive confirmation.
jevcompass doctor checks model metadata without a paid decision request. Run jevcompass doctor --test-jev only when you explicitly want one synthetic, billed API test. A configured key permits remote ranking for eligible tasks; it does not guarantee a recommendation or prove that a hook delivered one to Codex.
After jevcompass install, inspect the two registrations in Codex /hooks, trust them if prompted, and start a new Desktop or CLI session. For a focused check, give Codex a substantive task such as planning a small Python package and ask it to report any JevCompass advice ID and suggested IDs from its initial context before its first tool call. An absent ID means that session did not receive advice; doctor and local hook metrics alone cannot prove delivery. The advisor may intentionally stay silent when only a generic shell command is available.
If your Codex host does not load either hook, run jevcompass doctor --json to distinguish registration from recent invocation, and use the explicit jevcompass recommend --category project-setup --domain python command while investigating. Never treat a manual result as proof that a hook delivered context to an agent.
A clean Codex profile may have only a generic shell tool. JevCompass stays silent in that case because repeating “use the shell” would add no value. To see a concrete suggestion, explicitly preview and install the two optional first-party coding skills:
jevcompass skills install --dry-run
jevcompass skills install
jevcompass recommend --category coding --domain pythonOn a fresh isolated Linux profile using the published 0.1.16 package with no OpenRouter key, the final command returned this excerpt:
Local unranked fallback; Jev did not select these candidates.
- tool `exec_command`: Codex built-in exec_command for bounded local shell commands
- skill `jevcompass-focused-tests`: Choose focused tests after a change and complete required repository validation
When verified timings show the complete required suite is already cheap and covers the change, the workflow can run it directly without ranking or repeating its subsets. Separately required focused checks and frozen benchmark gates remain mandatory.
The skill is available locally after installation; read its SKILL.md and use it only when its condition fits. The recommendation does not prove the agent read or followed it, and installing a skill does not guarantee faster or better work. If you want no added skills, leave the profile unchanged; doctor and recommend will explain or demonstrate when JevCompass has enough of your existing candidates to advise. Hook installation is separate: jevcompass install --dry-run, then jevcompass install, review /hooks, and start a fresh session. Optional OpenRouter ranking uses coarse metadata only and may incur charges.
- Prompt hook:
UserPromptSubmitclassifies an eligible prompt locally, then checks reviewed candidates discovered on the machine. Explicit Codex Desktop/CLI hook or skill setup can point to the installedopenai-docsskill; routine prompts and unavailable skills can remain silent. - Subagent hook:
SubagentStartcan suggest role-level candidates for Codex's built-inexplorerandworkerroles. That event does not include the subagent's task text. - Optional spawn advice:
jevcompass install --spawn-adviceadds a narrowly matchedPreToolUsehook for agent creation only (Agent,spawn_agent, orcollaborationspawn_agent). It locally classifies readable child task text or a descriptivetask_name, then sends only coarse metadata and reviewed candidate descriptions to Jev. It never gates a spawn, shell, Git, Helm, or network command. If neither field provides a clear category or the only candidates are generic shell and Git, it stays silent. A normal reinstall preserves the opt-in;jevcompass install --disable-spawn-adviceremoves it. Review and trust the updated hook in/hooks, then start a fresh session. This is experimental and disabled by default; verify the child sees an advice ID before its first tool before relying on it. - Manual mode:
jevcompass recommend --category CATEGORY --domain DOMAIN [--role ROLE]requests advice using explicit metadata. Usejevcompass recommend --helpfor accepted values. - Optional skill pack: from v0.1.16,
jevcompass skills installis an explicit, experimental opt-in that makes two reviewed coding workflows available locally. Matching-skill guidance is conditional on the task; installation is separate from hooks and never automatic. - Python test runner: for
testing/python, a clear unittest command in local GitHub Actions workflows suppresses a globally installedpytestcandidate and offersunittestinstead. If CI uses pytest or both runners, JevCompass does not infer a unittest requirement. Always follow the repository's exact test command. - Uncertainty: Local fallback advice is labeled unranked. If there is no useful candidate, JevCompass can stay silent. A configured MCP server is not assumed to be callable in the active session.
- Control: Advice does not run tools or skills, change permissions, block commands, or replace project instructions and required checks.
For setup diagnostics, run jevcompass doctor; use jevcompass --version to identify the installed command and jevcompass doctor --json for structured output. The optional --test-jev flag sends one synthetic, billed request.
Prompt classification and candidate discovery happen locally. When optional Jev ranking is used, JevCompass sends an allowlisted category, domain, role, an optional coarse planning focus (migration or implementation), criteria, and generic descriptions of reviewed candidates to OpenRouter's Decisions API. It does not send the raw prompt, source code, diffs, repository paths, memory contents, or unapproved SKILL.md descriptions. A skill added with --approve-remote-metadata is an exception you control: its ID and your generic capability, use and avoid descriptions can be sent to OpenRouter when remote ranking is enabled. Preview these fields first; omit names or details you consider private. The explicit coding workflows can additionally send fixed strategy/failure/change signals and opaque test candidates with allowlisted kind, coverage and runtime buckets. Raw diagnostics, command strings and measured durations stay local. OpenRouter receives the API key in the HTTPS authorization header.
Local metrics contain event/category/outcome/timing and a short advice ID; they do not record the prompt, model response, paths, or memory content. Local cache entries contain selection metadata. As with any external service, review OpenRouter's terms and data handling before enabling remote ranking.
Remote ranking is optional. Without a key or a reliable Jev response, JevCompass can use a small local shortlist when useful; otherwise it can stay silent. Network errors and timeouts do not block the Codex task.
- Python 3.11+
- Linux: local runtime and CLI checks have been performed.
- macOS: targeted, but runtime behavior has not been verified.
- Windows: not a verified target.
- Codex Desktop and CLI: the default installation registers
UserPromptSubmitandSubagentStart; optional agent-spawn advice adds only a narrowly matchedPreToolUsehook. Host delivery is version-, trust-, and session-dependent. Fresh Desktop prompt delivery and broad subagent coverage remain unverified.
JevCompass does not infer Codex Plan UI mode. Check /hooks and start a fresh session after installation or a Codex upgrade.
Evidence is deliberately limited to the environments tested:
-
Version 0.1.23 publishes the three explicit coding-decision commands (
strategy choose,tests discover/rank,triagewith observations, ranking and diagnostic costs), the opt-in typed-decision cache, and five opt-in bundled skills; per-change details and evidence boundaries are in CHANGELOG.md. The release revision passed 858 offline Python tests on Linux; PR #310 passed hosted CI and merged asd5b4772; tag-gated publish run 36792180928 passed the Python 3.11 suite, build, exact-tag distribution checks and Trusted Publishing. PyPI 0.1.23 wheel SHA-2565293e33c3879129756712132bae42aa8955884f3bc2c289984c71ffc2e87c87b, sdistbacbe3a3ac8d657802dc197155f41de561aea0ea45af26d2246e51d778135f6f; a fresh registry pipx install on Python 3.14 passeddoctorwith both advisory hooks and a keyless diagnostic-cost triage smoke resolved locally. No repeatable task benefit is claimed. -
The published 0.1.22 package places selected IDs at the top of each advisory. A fresh isolated CLI 0.155.1 session reported
exec_commandandunittestwith correlated advice ID0623c35bbefore its first tool. The matching local hook took 5.99 ms, exit was 0, and the fictional fixture was unchanged. This single smoke does not measure task benefit. -
PyPI v0.1.21 and its GitHub release passed 275 Python 3.11 tests, hosted CI, Trusted Publishing, exact-tag distribution checks and a clean pipx registry install with two advisory hooks and doctor PASS. A single isolated synthetic
testing/pythonrequest composed the shell executor andunittestlocally in 2.96 ms; it is not a host delivery or broad effectiveness measurement. -
A fresh isolated Codex CLI 0.155.1 smoke using the published v0.1.12 wheel received
openai-docsadvice before its first tool, with a matching local hook ID. Additional synthetic CLI pairs confirmed remote Jev advice IDs before the first tool. The published v0.1.13 wheel adds bounded skill use/skip conditions; these checks do not prove broad coverage or effectiveness. -
The initial four blinded synthetic CLI pairs tied on task quality. Later focused pairs had mixed results, including one baseline win and one treatment win; no repeatable speed or quality improvement has been demonstrated.
-
A native Desktop
explorerchild reported a correlated advice ID in an earlier smoke. With published 0.1.18 and two optional bundled skills explicitly installed, a freshworkerreported advice ID96b34a1dandexec_command/jevcompass-focused-testsbefore its first shell read; the matching local SubagentStart metric took 56.79 ms. A freshexplorerwith only generic candidates correctly stayed silent. These are isolated role-level delivery checks; fresh Desktop prompt delivery and outcome benefit remain unverified. -
PyPI v0.1.20 and the GitHub release suppress low-signal package-docs advice when only the generic shell is available. Hosted CI and Trusted Publishing passed; a fresh Python 3.11 pipx registry install registered two advisory hooks, passed doctor, and stayed silent for package-docs in a clean profile. PyPI v0.1.19 and its GitHub release include the local Python test runner relevance fix. Hosted CI and Trusted Publishing passed; a clean registry pipx install on Python 3.11 registered two hooks, passed doctor and suggested
unittestwithoutpytestin this repository. The firstuvinstall attempt briefly missed the just-published version; retry without its cache succeeded. PyPI v0.1.18 and its GitHub release contain the corrected approved-metadata disclosure. Trusted publish run 36142929270 passed; a clean PyPI pipx install on Python 3.11 registered the two default hooks anddoctorpassed. PyPI v0.1.17 and its GitHub release include the opt-in custom-skill catalog. Its PyPI README had an imprecise privacy sentence; v0.1.18 corrects it. The 0.1.17 wheel passed local content checks, 271 Python 3.11 tests, green hosted CI and isolated pipx installation. Earlier PyPI v0.1.16 and its GitHub release contain byte-identical Python modules, catalog and optional skill files. The exact-tag archive passed 254 Python 3.11 tests and an isolated pipx wheel install with two default hooks, noPreToolUsegate anddoctorPASS. This validates packaging and local setup, not task outcomes or Desktop/macOS runtime. -
The v0.1.14 source includes the changes from PR #40 (respect installed skill MCP prerequisites), PR #41 (C05 equal-environment comparison tied on blind scores), and PR #42 (opt-in agent-spawn advice). The spawn check was an isolated CLI source-checkout run: the child saw an advice ID before its first tool using the descriptive task title because the host message was encoded. The v0.1.14 wheel also passed an isolated pipx install and clean synthetic CLI delivery check: the child's context contained a spawn advice ID before its first tool, though it did not repeat the ID before that tool. At the v0.1.14 source check, Desktop and macOS were unverified; later named-role Desktop smokes are described above. Fresh Desktop prompt delivery, macOS runtime and outcome benefit remain unverified. The default install remains two advisory hooks; routine commands are not gated.
-
A fresh isolated Codex CLI 0.155.1 session with installed 0.1.19 reported
exec_commandandunittestwith correlated advice ID215744b8before its first tool (UserPromptSubmit/testing/local, 33.66 ms). The fictional fixture's required unittest suite failed an existing assertion; this is local advice delivery, not remote ranking or measured benefit. -
A 0.1.19 installed-package P01 coding pair with equal opt-in skills delivered local advice before first tool, but neither arm recorded a required unittest exit. A quality correction after unblinding cannot support a blind outcome claim. See PILOT.md.
-
One repeated 0.1.19 P01 pair froze an anonymous quality assessment before mapping: both helpers passed. Treatment had correlated local advice and a recognized successful unittest; baseline test execution was unobserved. This limited result does not prove that advice caused the difference or improve overall outcomes.
-
A separate installed 0.1.19 P05 documentation pair tied on blind README quality. Treatment suggested only
exec_command; a required fixture unittest failure was observed there, while baseline validation remained unobserved. Productive-action timing and advice benefit were not established. -
A fresh installed 0.1.20 Codex CLI 0.155.1 P05 smoke reported no advisory before its first tool; the matching
package-docs/low-signal-skiphook metric took 5.73 ms, and the fictional CLI help check exited 0. This validates abstention, not a speed or quality benefit. -
The maintainer upgraded pipx from 0.1.20 to 0.1.22 without changing
hooks.json;doctorstill found two advisory hooks and no Jev PreToolUse gate. A native Desktopworkerreported ID9ceb0fddandexec_command/jevcompass-focused-testsbefore its first tool, matching aSubagentStart/coding/localmetric of 54.64 ms. This is one role delivery smoke, not fresh Desktop prompt delivery or skill-use evidence. -
A corrected P07 CLI pair on published 0.1.22 used a writable fictional API-contract fixture. Blind authored-file review favored baseline; treatment returned
insufficient-candidatesand no Jev advice. This pair measures ordinary model variation under abstention, not an effect of JevCompass. See pilot evidence. -
macOS runtime behavior and broad usefulness remain unverified.
These checks do not establish that recommendations improve outcomes. If you test JevCompass, please share a reproducible example of advice that helped—or a case where silence was the right result.
More host validation, narrower coding recommendations, and a measured 20-case comparison are tracked in the roadmap. These are planned work, not shipped capabilities or proven productivity gains.
Issues and pull requests are welcome. Helpful contributions include reproducible compatibility reports, careful documentation fixes, and synthetic tests that preserve the privacy boundary. Please do not post credentials, raw prompts, private code, or unredacted logs.
JevCompass is licensed under the Apache License 2.0. Copyright 2026 Toni Nowak.
Included since v0.1.23: this cache and its readable provenance labels ship in the published package. On versions up to 0.1.22 the cache existed only in source builds, and a private wheel built from that source may still report version 0.1.22; that does not make it the published artifact.
Set JEVCOMPASS_TYPED_DECISION_CACHE=1 to reuse validated strategy, test-order and triage choices across CLI processes for up to 24 hours. It is disabled by default. The cache is bounded to 128 records and uses private files under ~/.cache/jevcompass/typed-decisions-v1; JEVCOMPASS_TYPED_CACHE_DIR can select a dedicated private directory. An existing directory with broader permissions is skipped, not modified.
Keys cover the exact sanitized request, model, caller policy and confidence threshold. Records contain only fixed choice tokens, confidence and creation time: no prompts, commands, source, memory, credentials or backend prose. Every hit is validated again and mapped to the current local candidates. Local resolutions still take precedence; injected clients bypass this cache.
CLI JSON exposes cache_hit. Readable strategy and test-order output labels cache reuse as a cached decision with no new API call; triage labels its preferred step cached_preferred_next_step. A hit reports no new provider usage. Original failures and mandatory validation remain unchanged. IO errors or invalid records silently fall through to ordinary advice. Cold and warm runs must be evaluated separately, including cache preparation costs. Cache reuse alone does not prove faster coding-task completion.
