Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions PILOT.md
Original file line number Diff line number Diff line change
Expand Up @@ -1732,8 +1732,16 @@ This integration is an offline measurement capability, not delivery or efficacy

The profile event parser now observes acknowledgment after both legacy and configured valid triage results. A real synthetic event stream verifies acknowledgment before the next tool and distinguishes missing or late acknowledgment. This fixes delivery measurement only; it does not establish native delivery or efficacy.

## VCR392 — Symmetric repair-arm Git validation

Repair-profile arms now receive isolated Git baselines before execution. Setup failure prevents launch. The runner records the agent’s exact diff-check invocation and exit codes; missing or unsuccessful checks prevent task acceptance in both arms. Protected fixture digests exclude only internal Git metadata. Four offline tests cover setup failure, metadata handling, whitespace errors and missing/nonzero checks; existing profile tests remain green. No live delivery or efficacy result is established.

## VCR389 — Prospective query migration repair fixture

Staged a new synthetic URL query migration task with versioned contracts, historical evidence, immutable tests, and an external hash-pinned oracle. Offline checks establish the expected initial failure, a source-only reference repair, and rejection of protected-file tampering or unexpected files. The oracle independently runs focused/full checks and checks the source patch.

This is preparation, not a launched pair or efficacy evidence. Launch remains gated on symmetric arm Git initialization and observed agent diff validation. The accepted diagnostic metadata includes contract confirmation, but the current runner does not yet request that candidate; causal ranking and diagnostic selection require separate integration.

## VCR394 — Independent diagnostic and causal ranking receipts

A valid diagnostic next step is now scored separately from complete causal ordering. Unestablished, partial or malformed rankings remain explicitly non-complete without erasing valid diagnostic delivery. Complete ranking requires a unique full causal permutation; contract confirmation remains a requested diagnostic candidate only. Three offline tests cover these boundaries. No live delivery or efficacy claim.
8 changes: 8 additions & 0 deletions TASKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -1136,8 +1136,16 @@ The runner now writes the bounded redacted measurement to its private output dir

The profile event parser now observes acknowledgment after both legacy and configured valid triage results. A real synthetic event stream verifies acknowledgment before the next tool and distinguishes missing or late acknowledgment. This fixes delivery measurement only; it does not establish native delivery or efficacy.

## VCR392 — Symmetric repair-arm Git validation

Repair-profile arms now receive isolated Git baselines before execution. Setup failure prevents launch. The runner records the agent’s exact diff-check invocation and exit codes; missing or unsuccessful checks prevent task acceptance in both arms. Protected fixture digests exclude only internal Git metadata. Four offline tests cover setup failure, metadata handling, whitespace errors and missing/nonzero checks; existing profile tests remain green. No live delivery or efficacy result is established.

## VCR389 — Prospective query migration repair fixture

Staged a new synthetic URL query migration task with versioned contracts, historical evidence, immutable tests, and an external hash-pinned oracle. Offline checks establish the expected initial failure, a source-only reference repair, and rejection of protected-file tampering or unexpected files. The oracle independently runs focused/full checks and checks the source patch.

This is preparation, not a launched pair or efficacy evidence. Launch remains gated on symmetric arm Git initialization and observed agent diff validation. The accepted diagnostic metadata includes contract confirmation, but the current runner does not yet request that candidate; causal ranking and diagnostic selection require separate integration.

## VCR394 — Independent diagnostic and causal ranking receipts

A valid diagnostic next step is now scored separately from complete causal ordering. Unestablished, partial or malformed rankings remain explicitly non-complete without erasing valid diagnostic delivery. Complete ranking requires a unique full causal permutation; contract confirmation remains a requested diagnostic candidate only. Three offline tests cover these boundaries. No live delivery or efficacy claim.
144 changes: 125 additions & 19 deletions scripts/pilot_contract_triage_pair.py
Original file line number Diff line number Diff line change
Expand Up @@ -361,6 +361,7 @@ def _case_treatment_prompt(profile: CaseProfile, advice_policy: str) -> str:
guidance += (
f"Complete the requested behavior change in {profile.source_file}. "
f"Do not modify {profile.focused_test_file} or configured evidence files. "
"Run git diff --check and require exit code zero after the change. "
"Run the focused test again and the required full test suite after the change. "
"The supervisor runs a separate immutable oracle."
)
Expand Down Expand Up @@ -413,7 +414,11 @@ def _triage_argv(exit_code: int, profile: CaseProfile | None = None) -> list[str
suffix = []
for kind in profile.triage_kinds:
suffix.extend(("--kind", kind))
for hypothesis in profile.triage_hypotheses:
candidates = list(profile.triage_hypotheses)
if ("confirm_behavior_contract" in profile.triage_accepted_ids
and "confirm_behavior_contract" not in candidates):
candidates.append("confirm_behavior_contract")
for hypothesis in candidates:
suffix.extend(("--hypothesis", hypothesis))
if profile.rank_hypotheses:
suffix.append("--rank-hypotheses")
Expand Down Expand Up @@ -498,20 +503,28 @@ def _validated_triage(payload: Any, focused_exit: int,
or value.get("test_failed") is not True
):
return {"status": "unscored", "candidate_ids": []}
parsed["hypothesis_order"] = []
parsed["hypothesis_ranking_status"] = "not_established"
if profile is not None and profile.rank_hypotheses:
order = value.get("hypothesis_order")
if (value.get("hypothesis_ranking_status") != "complete"
or not isinstance(order, list)
or len(order) != len(profile.triage_hypotheses)
or any(not isinstance(item, str) for item in order)
or set(order) != set(profile.triage_hypotheses)
or len(set(order)) != len(order)):
return {"status": "unscored", "candidate_ids": []}
parsed["hypothesis_ranking_status"] = "complete"
parsed["hypothesis_order"] = list(order)
else:
parsed["hypothesis_ranking_status"] = "not_established"
parsed["hypothesis_order"] = []
status = value.get("hypothesis_ranking_status")
causal_ids = {
item for item in profile.triage_hypotheses
if item != "confirm_behavior_contract"
}
if (status == "complete" and isinstance(order, list)
and all(isinstance(item, str) for item in order)
and len(order) == len(causal_ids)
and len(set(order)) == len(order)
and set(order) == causal_ids and len(causal_ids) >= 2):
parsed["hypothesis_ranking_status"] = "complete"
parsed["hypothesis_order"] = list(order)
elif status == "not_established" and order == []:
parsed["hypothesis_ranking_status"] = "not_established"
else:
# A valid next step remains valid independently of an incomplete
# or malformed causal ordering; never promote that order to success.
parsed["hypothesis_ranking_status"] = "incomplete"
return parsed


Expand All @@ -537,6 +550,8 @@ def _event_receipts(
evidence: set[str] = set()
focused_exits: list[int] = []
full_exits: list[int] = []
git_diff_check_exits: list[int] = []
git_diff_check_invoked = False
useful_failure_ms: float | None = None
useful_failure_observed = False
triage_seen = False
Expand Down Expand Up @@ -624,6 +639,10 @@ def _event_receipts(
if kind:
pending[event_id] = (kind, None)
argv = common._command_argv(item)
if (profile is not None and profile.outcome_mode == "repair"
and argv == ["git", "diff", "--check"]):
git_diff_check_invoked = True
pending[event_id] = ("git-diff-check", None)
if argv and "jevcompass" in argv:
triage_seen = True
exact = _is_triage(item, focused_exits[0] if focused_exits else None, profile)
Expand All @@ -643,7 +662,9 @@ def _event_receipts(
elif event_type == "item.completed" and event_id in pending:
kind, _ = pending.pop(event_id)
code = item.get("exit_code")
if kind == "focused" and isinstance(code, int) and not isinstance(code, bool):
if kind == "git-diff-check" and isinstance(code, int) and not isinstance(code, bool):
git_diff_check_exits.append(code)
elif kind == "focused" and isinstance(code, int) and not isinstance(code, bool):
focused_exits.append(code)
output = item.get("aggregated_output")
if not isinstance(output, str):
Expand Down Expand Up @@ -718,6 +739,16 @@ def _event_receipts(
"codex_billing_estimate": None,
"event_count": len(list(lines)) if isinstance(lines, list) else None,
}
if profile is not None and profile.outcome_mode == "repair":
result["agent_git_diff_check_invocation_observed"] = git_diff_check_invoked
result["agent_git_diff_check_exit_codes"] = git_diff_check_exits
result["agent_git_diff_check_exit_code"] = (
git_diff_check_exits[-1] if git_diff_check_exits else None
)
result["agent_git_diff_check_passed"] = bool(
git_diff_check_invoked and git_diff_check_exits
and all(code == 0 for code in git_diff_check_exits)
)
if advice_policy == "nonbinding":
result["workflow_acknowledgment"] = workflow_acknowledgment
result["triage_result_acknowledgment"] = triage_result_acknowledgment
Expand Down Expand Up @@ -807,8 +838,58 @@ def _run_arm(
return result, answer


def _fixture_digest(root: Path, *, exclude_relative: str | None = None) -> str | None:
"""Hash fixture files while ignoring generated Python bytecode."""
def _initialize_private_git_baseline(root: Path) -> bool:
"""Initialize an isolated Git baseline for one copied repair fixture."""
git = shutil.which("git")
if not git:
return False
env = {
"PATH": os.environ.get("PATH", ""),
"HOME": str(root),
"GIT_CONFIG_NOSYSTEM": "1",
"GIT_CONFIG_GLOBAL": os.devnull,
"GIT_TERMINAL_PROMPT": "0",
"GIT_AUTHOR_NAME": "JevCompass Fixture",
"GIT_AUTHOR_EMAIL": "fixture@example.invalid",
"GIT_COMMITTER_NAME": "JevCompass Fixture",
"GIT_COMMITTER_EMAIL": "fixture@example.invalid",
}
commands = (
[git, "init", "--quiet"],
[git, "add", "-A"],
[git, "-c", "commit.gpgsign=false", "-c", "core.hooksPath=/dev/null",
"commit", "--quiet", "-m", "Immutable synthetic fixture baseline"],
)
try:
for argv in commands:
completed = subprocess.run(
argv, cwd=root, env=env, stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL, timeout=10, check=False,
)
if completed.returncode != 0:
return False
except (OSError, subprocess.TimeoutExpired):
return False
return True


def _repair_git_diff_check_valid(arm: dict[str, Any]) -> bool:
"""Require an observed, successful diff check for repair-profile completion."""
exits = arm.get("agent_git_diff_check_exit_codes")
return bool(
arm.get("agent_git_diff_check_invocation_observed") is True
and isinstance(exits, list) and exits
and all(isinstance(code, int) and not isinstance(code, bool) and code == 0
for code in exits)
and arm.get("agent_git_diff_check_passed") is True
)


def _fixture_digest(
root: Path, *, exclude_relative: str | None = None,
exclude_internal_git: bool = False,
) -> str | None:
"""Hash fixture files while ignoring bytecode and optional private .git metadata."""
digest = hashlib.sha256()
try:
for path in sorted(root.rglob("*")):
Expand All @@ -817,6 +898,8 @@ def _fixture_digest(root: Path, *, exclude_relative: str | None = None) -> str |
relative_text = path.relative_to(root).as_posix()
if relative_text == exclude_relative:
continue
if exclude_internal_git and (relative_text == ".git" or relative_text.startswith(".git/")):
continue
relative = relative_text.encode("utf-8")
digest.update(len(relative).to_bytes(4, "big"))
digest.update(relative)
Expand Down Expand Up @@ -963,9 +1046,11 @@ def run_pair(
for name in (source_file_rel, test_file_rel)
}
initial_fixture_digest = _fixture_digest(fixture_source)
repair_profile = bool(case_profile and case_profile.outcome_mode == "repair")
initial_protected_digest = _fixture_digest(
fixture_source, exclude_relative=source_file_rel,
) if case_profile and case_profile.outcome_mode == "repair" else initial_fixture_digest
exclude_internal_git=repair_profile,
) if repair_profile else initial_fixture_digest
for true_arm in ("baseline", "treatment"):
fixtures[true_arm] = private_root / f"{true_arm}-fixture"
homes[true_arm] = private_root / f"{true_arm}-home"
Expand All @@ -982,6 +1067,19 @@ def run_pair(
return receipt
if digests["baseline"] != digests["treatment"]:
raise RuntimeError("fixture parity verification failed")
if repair_profile and not all(
_initialize_private_git_baseline(fixtures[name])
for name in ("baseline", "treatment")
):
receipt = {
"schema_version": 1, "run_id": run_id, "status": "failed",
"failure": "repair_git_setup_failed", "arms": {},
}
common._private_write(
output_dir / "receipt.json",
(json.dumps(receipt, sort_keys=True, indent=2) + "\n").encode("utf-8"),
)
return receipt

arms: dict[str, dict[str, Any]] = {}
answers: dict[str, str | None] = {}
Expand All @@ -1003,6 +1101,11 @@ def run_pair(
measurement_path=output_dir / f"{label}-agent-measurement.json",
profile=case_profile,
)
if repair_profile:
arms[label].setdefault("agent_git_diff_check_invocation_observed", False)
arms[label].setdefault("agent_git_diff_check_exit_codes", [])
arms[label].setdefault("agent_git_diff_check_exit_code", None)
arms[label].setdefault("agent_git_diff_check_passed", False)
arms[label]["fixture_sha256_before"] = digests[true_arm]

source_unchanged: dict[str, bool] = {}
Expand All @@ -1026,8 +1129,10 @@ def run_pair(
)
protected_fixture_unchanged[label] = (
initial_protected_digest is not None
and _fixture_digest(fixtures[true_arm], exclude_relative=source_file_rel)
== initial_protected_digest
and _fixture_digest(
fixtures[true_arm], exclude_relative=source_file_rel,
exclude_internal_git=repair_profile,
) == initial_protected_digest
)
source_bytes = _safe_artifact(fixtures[true_arm] / source_file_rel)
if source_bytes is not None:
Expand Down Expand Up @@ -1070,6 +1175,7 @@ def run_pair(
and source_changed.get(label) is True
and tests_unchanged.get(label) is True
and protected_fixture_unchanged.get(label) is True
and _repair_git_diff_check_valid(arm)
)
arm["validated_completion_ms"] = endpoint_ms if completion_valid else None
arm["task_outcome_status"] = (
Expand Down
4 changes: 4 additions & 0 deletions tests/test_pilot_contract_case_profile.py
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,10 @@ def fake_arm(**kwargs):
"first_useful_failure_observed": True,
"full_suite_invocation_observed": True,
"full_suite_exit": 0 if repair_succeeds else 1,
"agent_git_diff_check_invocation_observed": repair_succeeds,
"agent_git_diff_check_exit_codes": [0] if repair_succeeds else [],
"agent_git_diff_check_exit_code": 0 if repair_succeeds else None,
"agent_git_diff_check_passed": repair_succeeds,
}
if advice_policy == "nonbinding" and kwargs["treatment"]:
result.update({
Expand Down
Loading