Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 30 additions & 0 deletions docs/measured.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,36 @@ definition-level ranking removed 0.388 / 0.569 / 0.652; report hygiene removed 0
0.743; packet order only 0.366 / 0.536 / 0.558; full 0.486 / 0.710 / 0.750. Raw results: `swebench/results-final-test.jsonl` (outside the repository); harness
`scripts/swebench_localize.py`, scoring `scripts/research/eval_loc.py`.

### v1.3.0, run once on Lite and Verified (2026-10-09/10)

v1.3.0 adds files the report names (traceback frames, paths) as a vote and the function list
(`functions_to_look`). Same protocol: release binary, one run per benchmark, never inspected per
instance. 95% bootstrap intervals in brackets.

| file level | n | Acc@1 | Acc@3 | Acc@5 | Acc@10 | BM25 Acc@1/3/5/10 |
|---|---|---|---|---|---|---|
| Lite strict holdout | 276 | **0.507** [0.45,0.57] | **0.717** [0.66,0.77] | **0.750** [0.70,0.80] | **0.815** | 0.301 / 0.507 / 0.587 / 0.721 |
| Lite, LocAgent's subset | 274 | **0.500** [0.44,0.56] | **0.723** [0.67,0.77] | **0.755** [0.70,0.81] | **0.828** | 0.299 / 0.522 / 0.606 / 0.734 |
| Verified, all | 500 | **0.446** [0.40,0.49] | **0.680** [0.64,0.72] | **0.736** [0.69,0.77] | **0.814** | 0.216 / 0.392 / 0.490 / 0.642 |
| Verified, not in Lite | 403 | **0.437** [0.39,0.48] | **0.670** [0.63,0.72] | **0.732** [0.69,0.77] | **0.811** | 0.194 / 0.372 / 0.476 / 0.620 |

v1.2.0 → v1.3.0 on the same instances: Lite 276 Acc@1 0.486 → 0.507; Verified-not-Lite 0.392 →
0.437; Acc@5 unchanged within ±0.005. p50 2.9 s / p90 10.6 s per Lite issue (cold index included).

| function level (all edited functions in top k) | n | Acc@1 | Acc@5 | Acc@10 |
|---|---|---|---|---|
| Lite, LocAgent's subset | 274 | 0.168 [0.12,0.22] | **0.394** [0.34,0.46] | **0.482** [0.42,0.54] |
| Lite strict holdout, with function gold | 250 | 0.168 | 0.400 | 0.492 |
| Verified, all, with function gold | 459 | 0.163 | 0.346 | 0.416 |
| Verified not in Lite, with function gold | 375 | 0.157 | 0.312 | 0.381 |

Function gold = innermost function containing a removed line or insertion point of the reference
patch (Python `ast`); the 274 instances that have one are exactly LocAgent's subset. Published
(LocAgent Table 4, function Acc@5/@10): BM25 0.318/0.369, CodeRankEmbed 0.518/0.588,
Agentless+Claude-3.5 0.588, LocAgent+Claude-3.5 0.734/0.774. The engine lists functions only for
reports of 60+ words; shorter reports count as misses here (30 of 274). Scoring them out, as the
session-17 dev numbers did, gives 0.443/0.541 on 244 — the smaller denominator flatters.

## Speed (same machine, same hour, release binary without embeddings)

ultralytics, 931 files: index ≈3.6 s; one question end-to-end p50 ≈0.70 s (n=10, cold process each
Expand Down
43 changes: 43 additions & 0 deletions docs/research/contributions-log.md
Original file line number Diff line number Diff line change
Expand Up @@ -395,3 +395,46 @@ Engine on SWE-bench dev (189 issues with function gold), func Acc@1/5/10/20: 0.1
→ **0.175/0.323/0.402/0.460** (prototype predicted 0.174/0.321/0.400). Same run, file level with
everything since v1.2.0 (named files): 0.249/0.484/0.600/0.680 → **0.308/0.527/0.621/0.692**.
MCP: reports now also get `functions_to_look` (five, `path:Class.method Lx-Ly`).

### 8.13 v1.3.0 run once on Lite and Verified; function-level denominator fix (session 18, 2026-10-10)

Release binary v1.3.0 (tag `6e94b91`), one run per benchmark (`swebench/lite-v130-*.jsonl`,
`verified-v130-*.jsonl`; 300/300 and 500/500 scored, no checkout errors this time).

| file level | n | Acc@1 | Acc@3 | Acc@5 | Acc@10 | BM25 |
|---|---|---|---|---|---|---|
| Lite strict holdout | 276 | **0.507** [0.45,0.57] | **0.717** [0.66,0.77] | 0.750 [0.70,0.80] | 0.815 | 0.301/0.507/0.587/0.721 |
| LocAgent subset | 274 | **0.500** [0.44,0.56] | **0.723** [0.67,0.77] | 0.755 [0.70,0.81] | 0.828 | 0.299/0.522/0.606/0.734 |
| Verified, all | 500 | **0.446** [0.40,0.49] | 0.680 [0.64,0.72] | 0.736 [0.69,0.77] | 0.814 | 0.216/0.392/0.490/0.642 |
| Verified not in Lite | 403 | **0.437** [0.39,0.48] | 0.670 [0.63,0.72] | 0.732 [0.69,0.77] | 0.811 | 0.194/0.372/0.476/0.620 |

v1.2.0 → v1.3.0, same instances: Lite-276 Acc@1 0.486 → 0.507, Verified-not-Lite 0.392 → 0.437
(+0.045; the named-files vote, C13, dev gain was +0.049 — it transferred); Acc@5 within ±0.005.
Latency p50 2.9 s / p90 10.6 s (Lite), 2.5 s / 8.8 s (Verified), cold index included.

**Function level** (`func_eval.py`, now with `--ci`, `--gold-cache`):

| | n | Acc@1 | Acc@5 | Acc@10 | Acc@20 |
|---|---|---|---|---|---|
| LocAgent subset (= all Lite instances with function gold) | 274 | 0.168 [0.12,0.22] | **0.394** [0.34,0.46] | **0.482** [0.42,0.54] | 0.544 |
| Lite strict, with function gold | 250 | 0.168 | 0.400 | 0.492 | 0.548 |
| Verified, with function gold | 459 | 0.163 | 0.346 | 0.416 | 0.468 |
| Verified not in Lite | 375 | 0.157 | 0.312 | 0.381 | 0.435 |

Reference (LocAgent Table 4, function Acc@5/@10): BM25 0.318/0.369 · CodeRankEmbed 0.518/0.588 ·
Agentless+Claude-3.5 0.588 · LocAgent+Claude-3.5 0.734/0.774. Engine: above BM25 by +0.08/+0.11,
below the dense retriever.

**Evaluation fix (denominator).** `func_eval.py` used to score only instances whose result *has* a
`definitions` list; the engine emits it only for reports of ≥ 60 words, so short reports left the
denominator. Measured effect on LocAgent's 274: skipping them gives 244 instances and Acc@5/@10
0.443/0.541; counting them as misses (correct) gives 274 and 0.394/0.482. The session-17 dev
function numbers (189 issues, 0.323 Acc@5) used the old rule; phase 1 re-baselines dev with the
correct one. The 274 count now matches LocAgent's subset exactly, which also validates the
function-gold extraction. Engine follow-up: emit the function list for short reports too.
`git show` failures in gold extraction now retry and raise instead of reading as "no gold".

**Phase 2 blocker (2026-10-10).** Five Baseten keys (Kimi-K3, GLM-5.2-Fast, GLM-5.3,
DeepSeek-V4-Flash, DeepSeek-V4-Pro) authenticate (`/v1/models` 200) but every chat call returns
HTTP 402 "please check your current payment status" — the accounts have no credit. The strong-LLM
stage stays unmeasured until one is funded.
52 changes: 31 additions & 21 deletions docs/research/paper-draft.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,10 @@ either run an LLM agent over the repository (Agentless, LocAgent, SWE-agent) or
retriever (CodeRankEmbed). We ask how far a local engine gets with neither: a per-project code
graph, field-weighted lexical ranking, a definition-level ranking of long reports, and explicit
report hygiene, all on a laptop CPU. On SWE-bench Lite, run once after tuning only on the SWE-bench
dev split, the engine reaches file-level Acc@1/3/5 of 0.486/0.710/0.750 on 276 held-out issues
(plain BM25 0.301/0.507/0.587) in 2.8 s per issue including a cold index — above a code embedding
model at Acc@1 and level with Agentless+GPT-4o at Acc@5. We report every component's effect,
dev split, the engine reaches file-level Acc@1/3/5 of 0.507/0.717/0.750 on 276 held-out issues
(plain BM25 0.301/0.507/0.587) in 2.9 s per issue including a cold index — above a code embedding
model at Acc@1 and level with Agentless+GPT-4o at Acc@5; on SWE-bench Verified's 403 issues outside
Lite, 0.437/0.670/0.732 (BM25 0.194/0.372/0.476). We report every component's effect,
negative results (dense retrieval on CPU, cross-encoder reranking, blind graph expansion), and a
holdout protocol that caught two results that looked-at sets had suggested. TODO: LLM stage (D1),
Verified, ablations.
Expand Down Expand Up @@ -56,25 +57,34 @@ tokens, definition-level BM25 with title ×3, RRF with the packet order → `whe

## 5. Results

### 5.1 SWE-bench Lite (holdout, 276)

| method | LLM | Acc@1 | Acc@3 | Acc@5 |
|---|---|---|---|---|
| BM25 (ours) | – | 0.301 | 0.507 | 0.587 |
| **engine** | – | **0.486** | **0.710** | **0.750** |
| Jina-Code-v2 † | – | 0.434 | 0.712 | 0.803 |
| CodeRankEmbed † | – | 0.526 | 0.777 | 0.847 |
| Agentless + GPT-4o † | ✓ | 0.672 | 0.745 | 0.745 |
| LocAgent + Claude-3.5 † | ✓ | 0.777 | 0.920 | 0.942 |
| engine + local LLM (D1) | ✓ (3B, CPU) | TODO | TODO | TODO |

† LocAgent Table 4, their 274-instance subset (head-to-head on the same subset: TODO).
### 5.1 SWE-bench Lite — head-to-head on LocAgent's 274 instances (holdout, run once)

| method | LLM | file Acc@1 | Acc@3 | Acc@5 | func Acc@5 | func Acc@10 |
|---|---|---|---|---|---|---|
| BM25 (ours) | – | 0.299 | 0.522 | 0.606 | – | – |
| BM25 † | – | 0.387 | 0.518 | 0.617 | 0.318 | 0.369 |
| engine v1.2.0 | – | 0.474 | 0.708 | 0.752 | – | – |
| **engine v1.3.0** | – | **0.500** [0.44,0.56] | **0.723** | **0.755** | **0.394** | **0.482** |
| Jina-Code-v2 † | – | 0.434 | 0.712 | 0.803 | – | – |
| CodeRankEmbed † | – | 0.526 | 0.777 | 0.847 | 0.518 | 0.588 |
| Agentless + GPT-4o † | ✓ | 0.672 | 0.745 | 0.745 | – | – |
| Agentless + Claude-3.5 † | ✓ | 0.726 | 0.792 | 0.796 | 0.588 | – |
| LocAgent + Qwen2.5-7B (ft) † | ✓ | 0.708 | 0.847 | 0.883 | – | – |
| LocAgent + Claude-3.5 † | ✓ | 0.777 | 0.920 | 0.942 | 0.734 | 0.774 |
| engine + local 3B LLM (D1, dev pilot) | ✓ (CPU) | rejected on dev (§6) | | | | |
| engine + API LLM (Agentless-style) | ✓ | TODO (no funded key) | | | | |

† LocAgent Table 4. Strict holdout (276, excludes 24 instances looked at in development):
0.507/0.717/0.750/0.815 (BM25 0.301/0.507/0.587/0.721).

### 5.2 SWE-bench dev (225) and Verified (496 of 500)

dev: BM25 0.160/0.347/0.427/0.538 → engine 0.249/0.484/0.600/0.680 (Acc@1/3/5/10).
Verified (v1.2.0, run once): BM25 0.216/0.391/0.490/0.641 → engine 0.405/0.669/0.732/0.804; on the
403 Verified issues not in Lite: 0.194/0.372/0.476/0.620 → 0.392/0.655/0.727/0.809.
dev: BM25 0.160/0.347/0.427/0.538 → engine v1.3.0 0.308/0.527/0.621/0.692 (Acc@1/3/5/10).
Verified (run once per version): BM25 0.216/0.392/0.490/0.642 → v1.2.0 0.405/0.669/0.732/0.804 (496)
→ **v1.3.0 0.446/0.680/0.736/0.814** (500). On the 403 Verified issues not in Lite: BM25
0.194/0.372/0.476/0.620 → v1.2.0 0.392/0.655/0.727/0.809 → v1.3.0 0.437/0.670/0.732/0.811.
Function level (v1.3.0, all edited functions in top k): Verified 459 with function gold
0.163/0.346/0.416 (Acc@1/5/10); not in Lite (375) 0.157/0.312/0.381.

### 5.3 Ablation (Lite holdout, 276)

Expand All @@ -96,8 +106,8 @@ Verified (v1.2.0, run once): BM25 0.216/0.391/0.490/0.641 → engine 0.405/0.669

### 5.5 Cost

p50 2.8 s / p90 11.4 s per issue on a 6-core laptop CPU, cold index included; packet p50 11.1k
tokens; no network, no GPU.
p50 2.9 s / p90 10.6 s per Lite issue (Verified 2.5 s / 8.8 s) on a 6-core laptop CPU (Ryzen 5
7530U), cold index included; packet p50 11.1k tokens; no network, no GPU.

## 6. Negative results and what the holdouts caught

Expand Down
78 changes: 66 additions & 12 deletions scripts/research/func_eval.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,13 +8,22 @@
how LocAgent arrives at 274 of Lite's 300).

py -3 scripts/research/func_eval.py --results run.jsonl --data rows.json --repos <clones> [--subset ids.json]
[--ci] [--gold-cache gold.json] [--key definitions]

`--gold-cache` stores the function gold per instance (computing it runs `git show`
per patched file); `--key` scores another list field of `{path, name}` entries.
An instance whose result has no list (the engine emits `definitions` only for
reports of 60+ words) counts as a miss; `--skip-missing` drops it instead, which
is how numbers before session 18 were computed (a smaller, easier denominator).
"""
import argparse
import ast
import json
import os
import random
import subprocess
import sys
import time

sys.path.insert(0, os.path.dirname(__file__))
from locagent_subset import touched_old_lines # noqa: E402
Expand All @@ -41,11 +50,25 @@ def walk(node, stack):
return out


def git_show(clone, commit, path, tries=4):
"""File at `commit`, or "" when the patch creates it. Blobless clones fetch the
blob over the network here; a failed fetch must not read as "no function gold"
(the instance would silently leave the denominator), so failures retry, then raise."""
for attempt in range(tries):
p = subprocess.run(["git", "-C", clone, "show", f"{commit}:{path}"], capture_output=True,
text=True, encoding="utf-8", errors="ignore")
if p.returncode == 0:
return p.stdout
if "does not exist in" in p.stderr or "exists on disk, but not in" in p.stderr:
return ""
time.sleep(2 * (attempt + 1))
raise RuntimeError(f"git show {commit}:{path} failed in {clone}: {p.stderr.strip()[:200]}")


def gold_functions(clone, commit, patch):
gold = set()
for path, lines in touched_old_lines(patch).items():
src = subprocess.run(["git", "-C", clone, "show", f"{commit}:{path}"], capture_output=True,
text=True, encoding="utf-8", errors="ignore").stdout
src = git_show(clone, commit, path)
defs = defs_with_names(src)
for line in lines:
inside = [d for d in defs if d[0] <= line <= d[1]]
Expand All @@ -63,35 +86,66 @@ def same(pred, gold):
"." not in pn or "." not in gn or pn == gn)


def boot_ci(vals, n=2000, seed=0):
rnd = random.Random(seed)
means = sorted(sum(rnd.choice(vals) for _ in vals) / len(vals) for _ in range(n))
return means[int(0.025 * n)], means[int(0.975 * n) - 1]


def load_gold(rows, ids, repos, cache_path):
cache = {}
if cache_path and os.path.exists(cache_path):
cache = json.load(open(cache_path, encoding="utf-8"))
missing = [i for i in ids if i not in cache]
for iid in missing:
row = rows[iid]
clone = os.path.join(repos, row["repo"].split("/")[1])
cache[iid] = sorted(gold_functions(clone, row["base_commit"], row["patch"]))
if cache_path and missing:
json.dump(cache, open(cache_path, "w", encoding="utf-8"))
return {i: {tuple(g) for g in cache[i]} for i in ids}


def main():
ap = argparse.ArgumentParser()
ap.add_argument("--results", required=True)
ap.add_argument("--data", required=True)
ap.add_argument("--repos", required=True)
ap.add_argument("--subset", default="")
ap.add_argument("--ci", action="store_true")
ap.add_argument("--gold-cache", default="")
ap.add_argument("--key", default="definitions")
ap.add_argument("--skip-missing", action="store_true")
args = ap.parse_args()
rows = {r["instance_id"]: r for r in json.load(open(args.data, encoding="utf-8"))}
keep = None
if args.subset:
keep = {r["instance_id"] for r in json.load(open(args.subset, encoding="utf-8"))}
res = [json.loads(l) for l in open(args.results, encoding="utf-8") if l.strip()]
res = [r for r in res if "definitions" in r and (keep is None or r["instance_id"] in keep)]
res = [r for r in res if "error" not in r and (keep is None or r["instance_id"] in keep)]
if args.skip_missing:
res = [r for r in res if args.key in r]
golds = load_gold(rows, [r["instance_id"] for r in res], args.repos, args.gold_cache)
ks = (1, 5, 10, 20)
hits = {k: 0 for k in ks}
n = 0
hits = {k: [] for k in ks}
for r in res:
row = rows[r["instance_id"]]
clone = os.path.join(args.repos, row["repo"].split("/")[1])
gold = gold_functions(clone, row["base_commit"], row["patch"])
gold = golds[r["instance_id"]]
if not gold:
continue
n += 1
preds = [(d["path"], d["name"]) for d in r["definitions"]]
preds = [(d["path"], d["name"]) for d in r.get(args.key, [])]
for k in ks:
top = preds[:k]
hits[k] += all(any(same(p, g) for p in top) for g in gold)
hits[k].append(int(all(any(same(p, g) for p in top) for g in gold)))
n = len(hits[1])
print(f"instances with function gold: {n}")
print(" ".join(f"func Acc@{k} {hits[k] / max(n, 1):.3f}" for k in ks))
cells = []
for k in ks:
cell = f"func Acc@{k} {sum(hits[k]) / max(n, 1):.3f}"
if args.ci and n:
lo, hi = boot_ci(hits[k])
cell += f" [{lo:.2f},{hi:.2f}]"
cells.append(cell)
print(" ".join(cells))


if __name__ == "__main__":
Expand Down
Loading