Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,8 @@ All notable user-facing changes live here. The README stays a product guide, not

### Measured

- **v1.4.0 run once** — file level identical to 1.3.0; function level Acc@5 on LocAgent's 274 Lite
instances 0.394 → 0.460 (Acc@10 0.482 → 0.540), on Verified 0.346 → 0.429.
- v1.3.0 run once on SWE-bench Lite (276 strict holdout: Acc@1/3/5 0.507/0.717/0.750) and Verified
(500: 0.446/0.680/0.736); function level on LocAgent's 274: Acc@5/@10 0.394/0.482.
- Fresh plain-language holdout (jsoup, Java, gold locked before the run): list R@3 0.750 vs BM25
Expand Down
17 changes: 17 additions & 0 deletions docs/measured.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,6 +119,23 @@ Agentless+Claude-3.5 0.588, LocAgent+Claude-3.5 0.734/0.774. The engine lists fu
reports of 60+ words; shorter reports count as misses here (30 of 274). Scoring them out, as the
session-17 dev numbers did, gives 0.443/0.541 on 244 — the smaller denominator flatters.

### v1.4.0, run once on Lite and Verified (2026-10-10)

v1.4.0 adds the functions a traceback runs through to the function list and lists functions for
every prompt. File level is identical to v1.3.0 on every instance (it was not touched). Function
level, every edited function in the top k, 95% bootstrap CI for Acc@5:

| function level | n | v1.3.0 Acc@1 / @5 / @10 | **v1.4.0 Acc@1 / @5 / @10** |
|---|---|---|---|
| Lite, LocAgent's subset | 274 | 0.168 / 0.394 / 0.482 | **0.223 / 0.460** [0.40,0.52] **/ 0.540** |
| Lite strict holdout, with function gold | 250 | 0.168 / 0.400 / 0.492 | **0.228 / 0.472 / 0.556** |
| Verified, with function gold | 459 | 0.163 / 0.346 / 0.416 | **0.198 / 0.429** [0.38,0.47] **/ 0.497** |
| Verified not in Lite | 375 | 0.157 / 0.312 / 0.381 | **0.189 / 0.397 / 0.464** |

Published (LocAgent Table 4, function Acc@5/@10): BM25 0.318/0.369, CodeRankEmbed 0.518/0.588,
Agentless+Claude-3.5 0.588, LocAgent+Claude-3.5 0.734/0.774. p50 2.8 s / p90 11.2 s per Lite issue.
Tables regenerate with `scripts/research/paper_tables.py --swe <swebench dir>`.

## Speed (same machine, same hour, release binary without embeddings)

ultralytics, 931 files: index ≈3.6 s; one question end-to-end p50 ≈0.70 s (n=10, cold process each
Expand Down
78 changes: 78 additions & 0 deletions docs/planning/15-next-session-brief.fa.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
# راهنمای session 19 — بعد از v1.4.0

> اول این را بخوان، بعد `docs/research/contributions-log.md` §8.13–8.17 و `docs/research/paper-draft.md`.
> همه‌ی اعداد اندازه‌گیری‌شده‌اند؛ منبع هر کدام در contributions-log است.

## ۱. وضعیت در پایان session 18 (۲۰۲۶-۱۰-۱۰)

**نسخه‌ها:** v1.4.0 منتشر و نصب شد (PR #144، tag روی `ee6cc27`، پشتیبان `neuromesh-v1.3.0-backup.exe`). تغییرات engine: رأی traceback→تابع، لیست تابع برای همه‌ی پرسش‌ها در
`packet --json`، اصلاح سقف واژه‌های query در رتبه‌بندی definition. سوییچ پژوهشی `NM_DIAG=1`.

**اعداد (فایل = همه‌ی فایل‌های gold در top-k؛ تابع = همه‌ی توابع ویرایش‌شده در top-k):**

| بنچمارک | نسخه | Acc@1 | Acc@3 | Acc@5 | Acc@10 |
|---|---|---|---|---|---|
| Lite سخت (۲۷۶) | v1.3.0 | 0.507 | 0.717 | 0.750 | 0.815 |
| LocAgent (۲۷۴) فایل | v1.3.0 | 0.500 | 0.723 | 0.755 | 0.828 |
| LocAgent (۲۷۴) تابع | v1.3.0 | 0.168 | – | 0.394 | 0.482 |
| Verified (۵۰۰) | v1.3.0 | 0.446 | 0.680 | 0.736 | 0.814 |
| dev تابع (۲۱۰) | v1.3.0 → v1.4.0 | 0.157 → 0.190 | – | 0.290 → 0.324 | 0.362 → 0.390 |
| LocAgent (۲۷۴) تابع | **v1.4.0** | **0.223** | – | **0.460** | **0.540** |
| Verified تابع (۴۵۹) | v1.3.0 → **v1.4.0** | 0.163 → **0.198** | – | 0.346 → **0.429** | 0.416 → **0.497** |
| فایل Lite/Verified | v1.4.0 | = v1.3.0 (دست نخورد) | | | |

**زبان ساده:** jsoup (holdout تازه‌ی پنجم، جاوا) R@3 فهرست 0.750 در برابر BM25 0.667 (پیش از S1: 0.417).
همه‌ی پنج holdout زبان ساده اکنون دیده شده‌اند — ادعای جدید زبان ساده holdout ششم لازم دارد.

## ۲. قواعد ثابت (همان session 18)

1. تنظیم فقط روی SWE-bench dev (۲۲۵؛ dev-fast ۵۹ / dev-rest ۱۶۶). Lite/Verified یک بار برای هر نسخه، بدون نگاه per-instance.
2. اول prototype پایتونی آفلاین (`scripts/research/priors.py`، `--prior ext` برای هر لیست ذخیره‌شده)، آستانه از قبل:
file Acc@1 یا func Acc@5 ≥ +0.02 در **هر دو** نیمه. آستانه را بعد از دیدن عدد عوض نکن.
3. **baseline درست:** v1.4.0 روی dev = `swebench/dev-s18a-all.jsonl` (NM_DEFINITIONS=300). فایل
`dev-func300-all.jsonl` قدیمی است (پیش از #140) — استفاده نکن.
4. مخرج سطح تابع: instance بدون لیست = miss (`func_eval.py` پیش‌فرض).
5. یک build در هر لحظه؛ `df -h /c` (session 18: ۳۸ → ۲۷ GB، بخشی از پروژه‌ی دیگر).
6. تست MCP با `NEUROMESH_NO_BROWSER=1` و Python subprocess. README.md و nm.config.json را commit نکن.
7. harness با `--no-bm25` (BM25 در اجراهای v1.3.0 ذخیره است) — چند برابر سریع‌تر.
8. heredoc حاوی سه backtick ابزار Bash را می‌شکند؛ اسکریپت patch را با Write بنویس.

## ۳. ابزارهای جدید session 18

| ابزار | کار |
|---|---|
| `scripts/research/priors.py` | prototype هر prior (traceback, repro, history, testlink, codeonly, ext) روی دو نیمه |
| `scripts/research/chunk_prf.py` | BM25 سطح definition پایتونی + RM3 + dense روی short-list (با checkout) |
| `scripts/research/dense_shortlist.py` | jina-code روی top-k definition engine (بدون checkout، با cache) |
| `scripts/research/diag_eval.py` | ارزیابی واریانت‌های `NM_DIAG` (no_stem, owner, raw, no_trace) و RRF آن‌ها |
| `scripts/research/doc_bridge.py` | لیست engine برای پرسش‌های ساده (cache) + آزمایش پل doc→code |
| `scripts/research/paper_tables.py` | همه‌ی جدول‌های مقاله از فایل‌های نتیجه |
| `scripts/research/llm_localize.py --model --functions` | مرحله‌ی LLM سبک Agentless (فایل → تابع)، آماده |

## ۴. فازهای پیشنهادی session 19

### فاز A — مرحله‌ی LLM (مسدود: اعتبار API)
هر پنج کلید Baseten که پارسا داد احراز هویت می‌شوند ولی chat → HTTP 402 (بدون اعتبار). با شارژ یکی:
```bash
export LLM_API_KEY=<key>
py -3 scripts/research/llm_localize.py --results ../swebench/dev-s18a-all.jsonl --data ../swebench/devset/dev-fast.json \
--repos ../swebench/repos --work C:/w/llm --out ../swebench/llm-dev-fast.jsonl \
--endpoint https://inference.baseten.co/v1 --model deepseek-ai/DeepSeek-V4-Pro-0813 --max-tokens 4000 --functions
```
معیار: Acc@1 فایل و func Acc@5 در برابر engine تنها، با توکن به ازای issue (Agentless ~ده‌ها هزار توکن).

### فاز B — نزدیک‌خطاها روی split بزرگ‌تر
dense short-list (file@1 +0.017/+0.024، +10 ثانیه) و ادغام دو شاخص stem/unstemmed (func@1 +0.039/+0.026)
هر دو مثبت ولی زیر آستانه. پیشنهاد: dev بزرگ‌تر (SWE-Gym یا SWE-bench train، مخازن غیر Lite/Verified)
تا قدرت آماری کافی باشد؛ تصمیم یک‌باره، با آستانه‌ی ثابت.

### فاز C — G (مقاله‌ی main track)
holdout ≥ ۵۰ پرسش زبان ساده با annotator دوم (پارسا/تیم). بدون آن فقط workshop/industry.

### فاز D — مقاله
`paper-draft.md` با `paper_tables.py` هم‌گام؛ بخش‌های ۱، ۲ و ۸ (مقدمه، کار مرتبط، جمع‌بندی) نوشته شوند.

## ۵. موانعی که فقط پارسا باز می‌کند
- اعتبار برای یکی از کلیدهای Baseten (یا هر endpoint سازگار OpenAI).
- annotator دوم برای gold مستقل.
- ری‌استارت اپ Claude: MCP دسکتاپ هنوز پروسه‌های ۸ اکتبر (v1.1.0) را اجرا می‌کند.
25 changes: 25 additions & 0 deletions docs/research/contributions-log.md
Original file line number Diff line number Diff line change
Expand Up @@ -561,3 +561,28 @@ dev-fast (file Acc@1 +0.017 = one issue of 59; func Acc@5 +0.019/+0.013). Noted:
alone puts the right file first more often than the engine (0.356 vs 0.307) while losing depth —
a cheap re-ranker of the top few files, not a retriever. Candidate for an opt-in deep mode, to be
decided on a larger split.

### 8.17 v1.4.0 run once on Lite and Verified (2026-10-10)

Release binary v1.4.0 (tag on `ee6cc27`, #144), one run each (`swebench/lite-v140-*.jsonl`,
`verified-v140-*.jsonl`; 300/300 and 500/500 scored, `--no-bm25`). File level identical to v1.3.0
on every instance. Function level (all edited functions in top k):

| | n | v1.3.0 Acc@1/5/10 | v1.4.0 Acc@1/5/10 | Δ@5 |
|---|---|---|---|---|
| LocAgent subset | 274 | 0.168/0.394/0.482 | **0.223/0.460/0.540** (@5 CI [0.40,0.52]) | +0.066 |
| Lite strict, with function gold | 250 | 0.168/0.400/0.492 | **0.228/0.472/0.556** | +0.072 |
| Verified, with function gold | 459 | 0.163/0.346/0.416 | **0.198/0.429/0.497** (@5 CI [0.38,0.47]) | +0.083 |
| Verified not in Lite | 375 | 0.157/0.312/0.381 | **0.189/0.397/0.464** | +0.085 |

Dev predicted +0.034 at func Acc@5 (0.290 → 0.324); the holdouts gained about twice that. Not
because the holdouts have more tracebacks or short reports (aggregate input statistics: frames in
20% / 20% / 14% of dev / Lite / Verified issues, under 60 words 10% / 12% / 14%), but because their
function gold is smaller: one edited function in 45% of dev issues (mean 3.7) vs 86% of Lite
(mean 1.17) and 72% of Verified (1.87) — when every edited function must be in the top k, a
correctly placed function pays off far more often on single-function gold. Dev understates
function-level effects on Lite-like data. Against published function-level numbers (LocAgent Table 4, Lite): above BM25
(0.318/0.369) by +0.14/+0.17, below CodeRankEmbed (0.518/0.588) by −0.06/−0.05 — the gap to the
GPU-scale dense retriever halved (it was −0.12/−0.11 with v1.3.0).
Installed as the dogfood binary (v1.3.0 kept as `neuromesh-v1.3.0-backup.exe`); MCP banner 1.4.0,
exits on stdin EOF (probe with `NEUROMESH_NO_BROWSER=1`).
Loading
Loading