diff --git a/docs/CHANGELOG.md b/docs/CHANGELOG.md index 4debc4a..f0907ac 100644 --- a/docs/CHANGELOG.md +++ b/docs/CHANGELOG.md @@ -18,6 +18,8 @@ All notable user-facing changes live here. The README stays a product guide, not ### Measured +- **v1.4.0 run once** — file level identical to 1.3.0; function level Acc@5 on LocAgent's 274 Lite + instances 0.394 → 0.460 (Acc@10 0.482 → 0.540), on Verified 0.346 → 0.429. - v1.3.0 run once on SWE-bench Lite (276 strict holdout: Acc@1/3/5 0.507/0.717/0.750) and Verified (500: 0.446/0.680/0.736); function level on LocAgent's 274: Acc@5/@10 0.394/0.482. - Fresh plain-language holdout (jsoup, Java, gold locked before the run): list R@3 0.750 vs BM25 diff --git a/docs/measured.md b/docs/measured.md index cfc0994..d15b954 100644 --- a/docs/measured.md +++ b/docs/measured.md @@ -119,6 +119,23 @@ Agentless+Claude-3.5 0.588, LocAgent+Claude-3.5 0.734/0.774. The engine lists fu reports of 60+ words; shorter reports count as misses here (30 of 274). Scoring them out, as the session-17 dev numbers did, gives 0.443/0.541 on 244 — the smaller denominator flatters. +### v1.4.0, run once on Lite and Verified (2026-10-10) + +v1.4.0 adds the functions a traceback runs through to the function list and lists functions for +every prompt. File level is identical to v1.3.0 on every instance (it was not touched). Function +level, every edited function in the top k, 95% bootstrap CI for Acc@5: + +| function level | n | v1.3.0 Acc@1 / @5 / @10 | **v1.4.0 Acc@1 / @5 / @10** | +|---|---|---|---| +| Lite, LocAgent's subset | 274 | 0.168 / 0.394 / 0.482 | **0.223 / 0.460** [0.40,0.52] **/ 0.540** | +| Lite strict holdout, with function gold | 250 | 0.168 / 0.400 / 0.492 | **0.228 / 0.472 / 0.556** | +| Verified, with function gold | 459 | 0.163 / 0.346 / 0.416 | **0.198 / 0.429** [0.38,0.47] **/ 0.497** | +| Verified not in Lite | 375 | 0.157 / 0.312 / 0.381 | **0.189 / 0.397 / 0.464** | + +Published (LocAgent Table 4, function Acc@5/@10): BM25 0.318/0.369, CodeRankEmbed 0.518/0.588, +Agentless+Claude-3.5 0.588, LocAgent+Claude-3.5 0.734/0.774. p50 2.8 s / p90 11.2 s per Lite issue. +Tables regenerate with `scripts/research/paper_tables.py --swe `. + ## Speed (same machine, same hour, release binary without embeddings) ultralytics, 931 files: index ≈3.6 s; one question end-to-end p50 ≈0.70 s (n=10, cold process each diff --git a/docs/planning/15-next-session-brief.fa.md b/docs/planning/15-next-session-brief.fa.md new file mode 100644 index 0000000..fad553f --- /dev/null +++ b/docs/planning/15-next-session-brief.fa.md @@ -0,0 +1,78 @@ +# راهنمای session 19 — بعد از v1.4.0 + +> اول این را بخوان، بعد `docs/research/contributions-log.md` §8.13–8.17 و `docs/research/paper-draft.md`. +> همه‌ی اعداد اندازه‌گیری‌شده‌اند؛ منبع هر کدام در contributions-log است. + +## ۱. وضعیت در پایان session 18 (۲۰۲۶-۱۰-۱۰) + +**نسخه‌ها:** v1.4.0 منتشر و نصب شد (PR #144، tag روی `ee6cc27`، پشتیبان `neuromesh-v1.3.0-backup.exe`). تغییرات engine: رأی traceback→تابع، لیست تابع برای همه‌ی پرسش‌ها در +`packet --json`، اصلاح سقف واژه‌های query در رتبه‌بندی definition. سوییچ پژوهشی `NM_DIAG=1`. + +**اعداد (فایل = همه‌ی فایل‌های gold در top-k؛ تابع = همه‌ی توابع ویرایش‌شده در top-k):** + +| بنچمارک | نسخه | Acc@1 | Acc@3 | Acc@5 | Acc@10 | +|---|---|---|---|---|---| +| Lite سخت (۲۷۶) | v1.3.0 | 0.507 | 0.717 | 0.750 | 0.815 | +| LocAgent (۲۷۴) فایل | v1.3.0 | 0.500 | 0.723 | 0.755 | 0.828 | +| LocAgent (۲۷۴) تابع | v1.3.0 | 0.168 | – | 0.394 | 0.482 | +| Verified (۵۰۰) | v1.3.0 | 0.446 | 0.680 | 0.736 | 0.814 | +| dev تابع (۲۱۰) | v1.3.0 → v1.4.0 | 0.157 → 0.190 | – | 0.290 → 0.324 | 0.362 → 0.390 | +| LocAgent (۲۷۴) تابع | **v1.4.0** | **0.223** | – | **0.460** | **0.540** | +| Verified تابع (۴۵۹) | v1.3.0 → **v1.4.0** | 0.163 → **0.198** | – | 0.346 → **0.429** | 0.416 → **0.497** | +| فایل Lite/Verified | v1.4.0 | = v1.3.0 (دست نخورد) | | | | + +**زبان ساده:** jsoup (holdout تازه‌ی پنجم، جاوا) R@3 فهرست 0.750 در برابر BM25 0.667 (پیش از S1: 0.417). +همه‌ی پنج holdout زبان ساده اکنون دیده شده‌اند — ادعای جدید زبان ساده holdout ششم لازم دارد. + +## ۲. قواعد ثابت (همان session 18) + +1. تنظیم فقط روی SWE-bench dev (۲۲۵؛ dev-fast ۵۹ / dev-rest ۱۶۶). Lite/Verified یک بار برای هر نسخه، بدون نگاه per-instance. +2. اول prototype پایتونی آفلاین (`scripts/research/priors.py`، `--prior ext` برای هر لیست ذخیره‌شده)، آستانه از قبل: + file Acc@1 یا func Acc@5 ≥ +0.02 در **هر دو** نیمه. آستانه را بعد از دیدن عدد عوض نکن. +3. **baseline درست:** v1.4.0 روی dev = `swebench/dev-s18a-all.jsonl` (NM_DEFINITIONS=300). فایل + `dev-func300-all.jsonl` قدیمی است (پیش از #140) — استفاده نکن. +4. مخرج سطح تابع: instance بدون لیست = miss (`func_eval.py` پیش‌فرض). +5. یک build در هر لحظه؛ `df -h /c` (session 18: ۳۸ → ۲۷ GB، بخشی از پروژه‌ی دیگر). +6. تست MCP با `NEUROMESH_NO_BROWSER=1` و Python subprocess. README.md و nm.config.json را commit نکن. +7. harness با `--no-bm25` (BM25 در اجراهای v1.3.0 ذخیره است) — چند برابر سریع‌تر. +8. heredoc حاوی سه backtick ابزار Bash را می‌شکند؛ اسکریپت patch را با Write بنویس. + +## ۳. ابزارهای جدید session 18 + +| ابزار | کار | +|---|---| +| `scripts/research/priors.py` | prototype هر prior (traceback, repro, history, testlink, codeonly, ext) روی دو نیمه | +| `scripts/research/chunk_prf.py` | BM25 سطح definition پایتونی + RM3 + dense روی short-list (با checkout) | +| `scripts/research/dense_shortlist.py` | jina-code روی top-k definition engine (بدون checkout، با cache) | +| `scripts/research/diag_eval.py` | ارزیابی واریانت‌های `NM_DIAG` (no_stem, owner, raw, no_trace) و RRF آن‌ها | +| `scripts/research/doc_bridge.py` | لیست engine برای پرسش‌های ساده (cache) + آزمایش پل doc→code | +| `scripts/research/paper_tables.py` | همه‌ی جدول‌های مقاله از فایل‌های نتیجه | +| `scripts/research/llm_localize.py --model --functions` | مرحله‌ی LLM سبک Agentless (فایل → تابع)، آماده | + +## ۴. فازهای پیشنهادی session 19 + +### فاز A — مرحله‌ی LLM (مسدود: اعتبار API) +هر پنج کلید Baseten که پارسا داد احراز هویت می‌شوند ولی chat → HTTP 402 (بدون اعتبار). با شارژ یکی: +```bash +export LLM_API_KEY= +py -3 scripts/research/llm_localize.py --results ../swebench/dev-s18a-all.jsonl --data ../swebench/devset/dev-fast.json \ + --repos ../swebench/repos --work C:/w/llm --out ../swebench/llm-dev-fast.jsonl \ + --endpoint https://inference.baseten.co/v1 --model deepseek-ai/DeepSeek-V4-Pro-0813 --max-tokens 4000 --functions +``` +معیار: Acc@1 فایل و func Acc@5 در برابر engine تنها، با توکن به ازای issue (Agentless ~ده‌ها هزار توکن). + +### فاز B — نزدیک‌خطاها روی split بزرگ‌تر +dense short-list (file@1 +0.017/+0.024، +10 ثانیه) و ادغام دو شاخص stem/unstemmed (func@1 +0.039/+0.026) +هر دو مثبت ولی زیر آستانه. پیشنهاد: dev بزرگ‌تر (SWE-Gym یا SWE-bench train، مخازن غیر Lite/Verified) +تا قدرت آماری کافی باشد؛ تصمیم یک‌باره، با آستانه‌ی ثابت. + +### فاز C — G (مقاله‌ی main track) +holdout ≥ ۵۰ پرسش زبان ساده با annotator دوم (پارسا/تیم). بدون آن فقط workshop/industry. + +### فاز D — مقاله +`paper-draft.md` با `paper_tables.py` هم‌گام؛ بخش‌های ۱، ۲ و ۸ (مقدمه، کار مرتبط، جمع‌بندی) نوشته شوند. + +## ۵. موانعی که فقط پارسا باز می‌کند +- اعتبار برای یکی از کلیدهای Baseten (یا هر endpoint سازگار OpenAI). +- annotator دوم برای gold مستقل. +- ری‌استارت اپ Claude: MCP دسکتاپ هنوز پروسه‌های ۸ اکتبر (v1.1.0) را اجرا می‌کند. diff --git a/docs/research/contributions-log.md b/docs/research/contributions-log.md index fd59cf3..b379970 100644 --- a/docs/research/contributions-log.md +++ b/docs/research/contributions-log.md @@ -561,3 +561,28 @@ dev-fast (file Acc@1 +0.017 = one issue of 59; func Acc@5 +0.019/+0.013). Noted: alone puts the right file first more often than the engine (0.356 vs 0.307) while losing depth — a cheap re-ranker of the top few files, not a retriever. Candidate for an opt-in deep mode, to be decided on a larger split. + +### 8.17 v1.4.0 run once on Lite and Verified (2026-10-10) + +Release binary v1.4.0 (tag on `ee6cc27`, #144), one run each (`swebench/lite-v140-*.jsonl`, +`verified-v140-*.jsonl`; 300/300 and 500/500 scored, `--no-bm25`). File level identical to v1.3.0 +on every instance. Function level (all edited functions in top k): + +| | n | v1.3.0 Acc@1/5/10 | v1.4.0 Acc@1/5/10 | Δ@5 | +|---|---|---|---|---| +| LocAgent subset | 274 | 0.168/0.394/0.482 | **0.223/0.460/0.540** (@5 CI [0.40,0.52]) | +0.066 | +| Lite strict, with function gold | 250 | 0.168/0.400/0.492 | **0.228/0.472/0.556** | +0.072 | +| Verified, with function gold | 459 | 0.163/0.346/0.416 | **0.198/0.429/0.497** (@5 CI [0.38,0.47]) | +0.083 | +| Verified not in Lite | 375 | 0.157/0.312/0.381 | **0.189/0.397/0.464** | +0.085 | + +Dev predicted +0.034 at func Acc@5 (0.290 → 0.324); the holdouts gained about twice that. Not +because the holdouts have more tracebacks or short reports (aggregate input statistics: frames in +20% / 20% / 14% of dev / Lite / Verified issues, under 60 words 10% / 12% / 14%), but because their +function gold is smaller: one edited function in 45% of dev issues (mean 3.7) vs 86% of Lite +(mean 1.17) and 72% of Verified (1.87) — when every edited function must be in the top k, a +correctly placed function pays off far more often on single-function gold. Dev understates +function-level effects on Lite-like data. Against published function-level numbers (LocAgent Table 4, Lite): above BM25 +(0.318/0.369) by +0.14/+0.17, below CodeRankEmbed (0.518/0.588) by −0.06/−0.05 — the gap to the +GPU-scale dense retriever halved (it was −0.12/−0.11 with v1.3.0). +Installed as the dogfood binary (v1.3.0 kept as `neuromesh-v1.3.0-backup.exe`); MCP banner 1.4.0, +exits on stdin EOF (probe with `NEUROMESH_NO_BROWSER=1`). diff --git a/docs/research/paper-draft.md b/docs/research/paper-draft.md index 4076fd5..c7b6a7f 100644 --- a/docs/research/paper-draft.md +++ b/docs/research/paper-draft.md @@ -1,7 +1,8 @@ # Paper draft (living) — *Where to look: LLM-free, CPU-only code localisation for coding agents, measured on holdouts* -Status: skeleton with measured numbers (session 17). Every number traces to -`contributions-log.md`; TODO marks what is not measured yet. +Status: full draft of every section except the LLM-stage row (blocked on API credit), numbers +through session 18. Every number traces to `contributions-log.md`; the SWE-bench tables are rebuilt +from the result files by `scripts/research/paper_tables.py`. TODO marks what is not measured yet. ## Abstract (draft) @@ -14,29 +15,66 @@ dev split, the engine reaches file-level Acc@1/3/5 of 0.507/0.717/0.750 on 276 h (plain BM25 0.301/0.507/0.587) in 2.9 s per issue including a cold index — above a code embedding model at Acc@1 and level with Agentless+GPT-4o at Acc@5; on SWE-bench Verified's 407 issues outside Lite, 0.435/0.671/0.732 (BM25 0.194/0.373/0.477). We report every component's effect, -negative results (dense retrieval on CPU, cross-encoder reranking, blind graph expansion), and a -holdout protocol that caught two results that looked-at sets had suggested. TODO: LLM stage (D1), -Verified, ablations. +negative results (dense retrieval on CPU, cross-encoder reranking, blind graph expansion, commit +history, pseudo-relevance feedback, documentation bridges), and a holdout protocol that caught two +results that looked-at sets had suggested and confirmed a third on a fresh repository. At function +level the engine lists the edited function among its first five on 0.394 of LocAgent's 274 Lite +instances (BM25 0.318; CodeRankEmbed 0.518) with no model at all, 0.429 on Verified. ## 1. Introduction -- Problem: context for coding agents; cost of reading; localisation as the first step. -- Gap: LLM-heavy or GPU-heavy localisers; local-first tools (Aider RepoMap, Continue) are not - measured on standard localisation benchmarks; tuning on the evaluation set is common. -- Contributions: - 1. A holdout protocol for retrieval engines (gold locked before runs; dev vs holdout labelled per - set; one run per holdout) and evidence it matters (§6). - 2. An LLM-free engine: prompt-evidence seeding, BM25F with a comment field, definition-level - ranking for reports, report hygiene, localisation list (C2–C12 in the log). - 3. Results on SWE-bench Lite/dev/Verified and on four plain-language holdouts in four languages. - 4. Negative results with numbers. +A coding agent asked to fix an issue first has to find where the fix goes. On SWE-bench that step — +localisation — is where agent pipelines spend much of their budget: Agentless prompts a model with +the repository tree, then with file skeletons, then with code; LocAgent and SWE-agent run multi-step +LLM searches. The strongest model-free alternative published, CodeRankEmbed, embeds every function +of the repository with a 137M-parameter encoder — hours of compute per large repository on a +laptop CPU (we measured 2.6 chunks/s for a comparable model, §6). Local-first assistants (Aider's +RepoMap, Continue, Cody) ship structural or lexical retrieval but are not measured on standard +localisation benchmarks. + +We ask a narrow question: **how far does a local engine get with no LLM and no GPU**, if every +design decision is measured on a tuning split and confirmed once on held-out data? The engine +(open source) builds a per-project code graph in seconds, ranks files with field-weighted BM25, +ranks *definitions* for long reports, cleans issue-template scaffolding out of the query, and +reads the report's own pointers — the files and functions its traceback runs through. + +Contributions: +1. **A holdout protocol for retrieval engines** — gold locked by commit before any run, a tuning + split kept apart from every evaluation set, one run per released version on Lite and Verified, + and per-set labels for what has been looked at (§4). It caught two results that looked-at sets + suggested (a cross-encoder reranker, a neighbour vote) and confirmed one on a fresh repository. +2. **An LLM-free localiser** that reaches file-level Acc@1 0.507 on 276 held-out SWE-bench Lite + issues (BM25 0.301) at 2.9 s per issue on a laptop CPU, above a code-embedding retriever at + Acc@1, and lists the edited function in its top five on 0.460 of LocAgent's subset (§5). +3. **Ablations and costs** for every component, on the holdout (§5.3, §5.5). +4. **Negative results with numbers** — a dozen ideas from the bug-localisation and retrieval + literature that did not survive the protocol, and why (§6). ## 2. Related work -Agentless (hierarchical LLM localisation), LocAgent (graph + LLM agent), SWE-agent / OpenHands / -MoatlessTools (agentic search), CodeRankEmbed / Jina code embeddings (dense), BM25, Aider RepoMap -(PageRank over tags), Continue/Cody (retrieve + rerank), BM25F (Robertson et al.), hybrid fusion -(Bruch et al. 2023), RRF (Cormack et al. 2009). +**LLM localisers.** Agentless narrows hierarchically (files from the repository tree, then +classes/functions from skeletons, then lines) with one model call per level and sampling + voting. +LocAgent indexes a heterogeneous code graph (files, classes, functions; contain/import/invoke/ +inherit edges) and lets an LLM agent search it with entity-content BM25 and graph traversal tools; +it is the strongest published localiser on SWE-bench Lite (file Acc@5 0.942 with Claude-3.5). +SWE-agent, OpenHands and MoatlessTools search through shell or retrieval tools inside an agent loop. +Our engine borrows LocAgent's entity-level content ranking and Agentless's file→function narrowing, +without the model. + +**Dense retrieval.** CodeRankEmbed and Jina code embeddings rank function chunks by cosine with the +issue; LocAgent reports them as the strongest model-free baselines. We measure the CPU cost of +whole-repository chunk embedding and a short-list variant that embeds only the engine's top +definitions (§6). + +**Classic bug localisation.** IR-based bug localisation (BugLocator: rVSM plus similar fixed bugs; +Locus: change hunks; Rahman & Roy: query reformulation) uses report text, history and structure. We +test the history and reformulation (RM3) ideas under the same protocol and report them negative on +SWE-bench dev. + +**Ranking machinery.** BM25F (Robertson et al.), reciprocal rank fusion (Cormack et al. 2009) and +convex lexical/dense fusion (Bruch et al. 2023). Local-first assistants: Aider's RepoMap (PageRank +over tags; we measure it as a retriever, R@5 0.04–0.48 on our sets), Continue and Cody (retrieve + +rerank; a code-trained cross-encoder did not survive our holdout). ## 3. System @@ -54,11 +92,20 @@ skipped, voting at double weight deepest-first. ## 4. Evaluation protocol -- SWE-bench dev (225) = tuning; dev-fast (59) for iteration; Lite test (300) = holdout run once; - the 24 instances looked at in an earlier session excluded from the strict row (276). -- Metric: file-level Acc@k (all gold files in top k), bootstrap 95% CIs. -- Plain-language: four holdouts (ripgrep/Rust, click/Python, cobra/Go, axios/JS), 12 q each, gold - locked by commit before any run. +- **Tuning split:** SWE-bench dev (225 issues; pvlib, pydicom, sqlfluff, astroid, pyvista, + marshmallow — no repository shared with Lite or Verified), halves dev-fast (59, ten per repository) + and dev-rest (166). A change is kept only if file Acc@1 or function Acc@5 rises by at least 0.02 + in *both* halves — a bar fixed before each measurement. Every idea is first a Python prototype over + stored engine outputs, then an engine port re-measured on the full split. +- **Holdouts:** SWE-bench Lite test (300) and Verified (500), each run once per released version with + the release binary, never inspected per instance. The 24 Lite instances looked at in an early + session are excluded from the strict row (276); LocAgent's 274-instance subset (patches that edit an + existing function, rebuilt exactly) gives the head-to-head; 407 Verified issues are outside Lite. +- **Metrics:** file-level Acc@k (all gold files in the top k) and function-level Acc@k (all edited + functions in the top k; an instance without a function list counts as a miss), 95% bootstrap CIs. +- **Plain-language questions:** five holdouts (ripgrep/Rust, click/Python, cobra/Go, axios/JS, + jsoup/Java), 12 questions each, gold written from source and locked by commit before the first run; + metric = R@3 of the localisation list and packet recall/precision. ## 5. Results @@ -69,7 +116,8 @@ skipped, voting at double weight deepest-first. | BM25 (ours) | – | 0.299 | 0.522 | 0.606 | – | – | | BM25 † | – | 0.387 | 0.518 | 0.617 | 0.318 | 0.369 | | engine v1.2.0 | – | 0.474 | 0.708 | 0.752 | – | – | -| **engine v1.3.0** | – | **0.500** [0.44,0.56] | **0.723** | **0.755** | **0.394** | **0.482** | +| engine v1.3.0 | – | 0.500 [0.44,0.56] | 0.723 | 0.755 | 0.394 | 0.482 | +| **engine v1.4.0** | – | **0.500** [0.44,0.56] | **0.723** | **0.755** | **0.460** [0.40,0.52] | **0.540** | | Jina-Code-v2 † | – | 0.434 | 0.712 | 0.803 | – | – | | CodeRankEmbed † | – | 0.526 | 0.777 | 0.847 | 0.518 | 0.588 | | Agentless + GPT-4o † | ✓ | 0.672 | 0.745 | 0.745 | – | – | @@ -82,15 +130,21 @@ skipped, voting at double weight deepest-first. † LocAgent Table 4. Strict holdout (276, excludes 24 instances looked at in development): 0.507/0.717/0.750/0.815 (BM25 0.301/0.507/0.587/0.721). -### 5.2 SWE-bench dev (225) and Verified (496 of 500) +### 5.2 SWE-bench dev (225) and Verified (500) -dev: BM25 0.160/0.347/0.427/0.538 → engine v1.3.0 0.308/0.527/0.621/0.692 (Acc@1/3/5/10). +dev, file level: BM25 0.160/0.347/0.427/0.538 → engine v1.3.0 0.308/0.527/0.621/0.692 (Acc@1/3/5/10; +v1.4.0 identical). Function level (210 issues with function gold), v1.3.0 → v1.4.0: Acc@1 +0.157 → 0.190, @5 0.290 → 0.324, @10 0.362 → 0.390; per half fast 0.135/0.269 → 0.173/0.308, rest +0.165/0.297 → 0.196/0.329 (Acc@1/@5). Verified (run once per version): BM25 0.216/0.392/0.490/0.642 → v1.2.0 0.405/0.669/0.732/0.804 (496) → **v1.3.0 0.446/0.680/0.736/0.814** (500). On the Verified issues not in Lite (407; v1.2.0 lost 4 to checkout errors): BM25 0.194/0.372/0.476/0.620 → v1.2.0 0.392/0.655/0.727/0.809 (403 scored) → v1.3.0 0.435/0.671/0.732/0.811 (all 407; BM25 on 407: 0.194/0.373/0.477/0.622). -Function level (v1.3.0, all edited functions in top k): Verified 459 with function gold -0.163/0.346/0.416 (Acc@1/5/10); not in Lite (375) 0.157/0.312/0.381. +Function level (all edited functions in top k), v1.3.0 → v1.4.0: Verified 459 with function gold +0.163/0.346/0.416 → 0.198/0.429/0.497 (Acc@1/5/10); not in Lite (375) 0.157/0.312/0.381 → +0.189/0.397/0.464. On Lite the gain (+0.066 @5) was twice the dev prediction (+0.034): dev's function gold is +larger (one edited function in 45% of dev issues vs 86% of Lite), so dev understates +function-level effects on Lite-like data. ### 5.3 Ablation (Lite holdout, 276) @@ -103,12 +157,17 @@ Function level (v1.3.0, all edited functions in top k): Verified 459 with functi ### 5.4 Plain-language holdouts (R@3) -| set | BM25 | engine | -|---|---|---| -| ripgrep (Rust) | 0.500 | 0.500 | -| click (Python) | 0.750 | 0.917 | -| cobra (Go) | 0.875 | 0.875 | -| axios (JS, fresh) | 0.583 | 0.833 | +| set | BM25 | engine before S1 (v1.2.0) | engine (S1) | +|---|---|---|---| +| ripgrep (Rust) | 0.500 | 0.208 | 0.500 | +| click (Python) | 0.750 | 0.750 | 0.917 | +| cobra (Go) | 0.875 | 0.792 | 0.875 | +| axios (JS, fresh at its run) | 0.583 | – | 0.833 | +| **jsoup (Java, fresh; gold locked; run once)** | **0.667** | **0.417** | **0.750** | + +S1 (fusing the packet order with the whole-question ranking for short questions) was decided on +the first three (looked at), was neutral on axios, and is confirmed on jsoup (+0.333 over the +packet order, +0.083 over BM25). ### 5.5 Cost @@ -134,6 +193,10 @@ p50 2.9 s / p90 10.6 s per Lite issue (Verified 2.5 s / 8.8 s) on a 6-core lapto ## 7. Threats to validity +The tuning split differs from the holdouts in gold size (dev: 1.9 gold files and 3.7 gold functions +per issue on average; Lite is single-file with 1.17 gold functions), so dev numbers are lower and +function-level gains transfer larger than predicted. + Bookkeeping is itself a threat: two of our own evaluation slips were caught only by re-deriving counts against an external reference (LocAgent's 274) — a denominator that silently dropped short reports, and a prototype baseline taken from an older run. Both are reported with their effect. @@ -141,4 +204,14 @@ reports, and a prototype baseline taken from an older run. Both are reported wit Different instance subsets vs published numbers; single annotator for plain-language gold; 12 questions per plain-language holdout; Windows-only timing; blobless clones / network. -## 8. Conclusion — TODO +## 8. Conclusion + +A model-free engine on a laptop CPU localises SWE-bench Lite issues at file Acc@1 0.507 (276 +held-out issues), above a code-embedding retriever and below LLM agents, at a few seconds and no +tokens per issue; at function level it is above BM25 and below the dense retriever. The holdout +protocol mattered: two of the ideas that looked best on looked-at sets lost on fresh data, and most +literature priors we tried (history, query expansion, documentation bridges, dense re-ranking of a +short list) did not clear a bar fixed before measuring. What remains for a main-track paper: an LLM +stage on top of `where_to_look` measured in tokens and Acc@1 against Agentless and LocAgent (blocked +on API credit at the time of writing), and plain-language holdouts of at least fifty questions with +a second annotator. diff --git a/scripts/research/paper_tables.py b/scripts/research/paper_tables.py index 67601d8..a159fb7 100644 --- a/scripts/research/paper_tables.py +++ b/scripts/research/paper_tables.py @@ -37,6 +37,20 @@ "Verified, all": ("verified", None), "Verified, not in Lite": ("verified", "devset/verified-not-lite.json"), } +DEV = [ # SWE-bench dev (tuning split): engine runs per version + ("engine v1.3.0", "dev-func2-all.jsonl"), + ("engine v1.4.0", "dev-s18a-all.jsonl"), +] +# Plain-language list R@3: (set, gold toml under the repo, {label: cached lists}) +PLAIN = [ + ("ripgrep (Rust)", "tests/third_party/concept-holdout/ripgrep/gold_tasks.toml", {"v1.3.0": "plain-lists-v130.jsonl"}), + ("click (Python)", "tests/third_party/concept-holdout2/click/gold_tasks.toml", {"v1.3.0": "plain-lists-v130.jsonl"}), + ("cobra (Go)", "tests/third_party/concept-holdout3/cobra/gold_tasks.toml", {"v1.3.0": "plain-lists-v130.jsonl"}), + ("axios (JavaScript)", "tests/third_party/concept-holdout4/axios/gold_tasks.toml", {"v1.3.0": "plain-lists-v130.jsonl"}), + ("jsoup (Java)", "tests/third_party/concept-holdout5/jsoup/gold_tasks.toml", + {"v1.2.0": "plain-lists-jsoup-v120.jsonl", "v1.3.0": "plain-lists-jsoup-v130.jsonl", + "v1.4.0": "plain-lists-jsoup-s18b.jsonl"}), +] ABLATION = [ # Lite strict holdout, v1.2.0 engine with one component switched off (NM_ABLATE build) ("nothing removed", "results-final-test.jsonl", "loc_files"), ("report hygiene + code tokens (C11)", "abl-hyg-all.jsonl", "loc_files"), @@ -174,6 +188,48 @@ def main(): lines.append(f"| {label} | " + " | ".join(file_row(load(p), ids, key, ci=False)) + " |") lines.append("") + lines += ["## SWE-bench dev (tuning split, 225; halves dev-fast 59 / dev-rest 166)", "", + "| run | half | file Acc@1 | @3 | @5 | @10 | func Acc@1 | @5 | @10 |", "|---|---|---|---|---|---|---|---|---|"] + dev_rows = {r["instance_id"]: r for r in json.load(open(os.path.join(swe, "devset/swe-dev.json"), encoding="utf-8"))} + dev_gold = load_gold(dev_rows, list(dev_rows), repos, os.path.join(swe, "func-gold-dev.json")) + fast = {r["instance_id"] for r in json.load(open(os.path.join(swe, "devset/dev-fast.json"), encoding="utf-8"))} + for label, f in DEV: + p = os.path.join(swe, f) + if not os.path.exists(p): + continue + res = load(p) + for half, ids in (("fast", [i for i in dev_rows if i in fast]), ("rest", [i for i in dev_rows if i not in fast]), + ("all", list(dev_rows))): + fcells = file_row(res, ids, "loc_files", ci=False) + _, ucells = func_row(res, ids, dev_gold) + lines.append(f"| {label} | {half} | " + " | ".join(fcells + [c.split(" [")[0] for c in ucells]) + " |") + lines.append("") + + lines += ["## Plain-language questions, localisation list R@3", "", + "| set | run | R@3 |", "|---|---|---|"] + try: + import tomllib + except ImportError: # Python < 3.11 + import tomli as tomllib + root = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..")) + for name, gold_toml, runs_ in PLAIN: + gold = {t["id"]: t["gold_files"] for t in tomllib.load(open(os.path.join(root, gold_toml), "rb"))["task"]} + for label, f in runs_.items(): + p = os.path.join(swe, f) + if not os.path.exists(p): + continue + lists = {} + for line in open(p, encoding="utf-8"): + if line.strip(): + r = json.loads(line) + if r["id"] in gold: + lists[r["id"]] = r["localization"] + if len(lists) < len(gold): + continue + r3 = sum(sum(1 for g in gold[i] if g in lists[i][:3]) / len(gold[i]) for i in gold) / len(gold) + lines.append(f"| {name} | engine {label} | {r3:.3f} |") + lines.append("") + lines += ["## Cost per issue (cold index included, laptop CPU)", "", "| run | n | latency p50 | p90 | packet tokens p50 |", "|---|---|---|---|---|"] for (bench, label), (res, _) in runs.items():