Repository navigation
Issue mode: definition-level ranking + report hygiene (SWE dev Acc@5 0.456 → 0.596) - #134
Merged
Merged
Conversation
…fast Acc@5 0.456 -> 0.596
Measured on SWE-bench *dev* (tuning set; no repository shared with Lite test),
dev-fast = 10 issues per dev repository, paired runs, file-level Acc@k:
Acc@1 Acc@3 Acc@5 Acc@10
plain BM25 0.175 0.386 0.474 0.526
before 0.246 0.421 0.456 -
this 0.263 0.509 0.596 0.719
- Issue-template scaffolding is stripped from reports (>= 60 words): HTML
comments, links, checklists, section headings (not `#` lines inside code
fences). It ranked CODE_OF_CONDUCT.md and the CLI module on sqlfluff.
- Code-like tokens (`L031`, `LT02`, `E501`, `utf8`) are kept whole: every word
tokenizer dropped the one term a rule report shares with `rules/L031.py`.
- Definition-level BM25 (`chunk_rank`): file head + every definition span the
graph recorded, path words on each, a file scores its best definition, the
report's title counts 3x. Built lazily from the sources, cached like
file_rank. Prototype in scripts/research/lex_variants.py first.
- `localization` in `packet --json`: the packet's files, then the ranking's
runners-up, tests/docs/fixtures last, and for a long report RRF with the
definition ranking. A packet holds 1-5 files; Acc@5 was capped by it.
Rejected with numbers (contributions-log 8.3/8.4): real body tf in BM25F
(0.544 -> 0.509 @5), the same fusion choosing packet seeds (no change), a
cross-encoder (jina-reranker-v2) on issues (0.492 -> 0.373 @3; it helps short
plain-language questions: ripgrep 0.208 -> 0.583, click 0.750 -> 0.958).
Twelve sets unchanged (dev 0.938 ... holdout-web 0.917/0.643, concept 0.679/0.398,
concept-holdout 0.500/0.156, concept-holdout2 0.958/0.342).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
SWE-bench dev split (tuning set, disjoint from Lite test), dev-fast (59 issues), paired, file-level Acc@k:
L031,E501,utf8)chunk_rank, title ×3), RRF-fused into a newlocalizationlist inpacket --jsonscripts/research/(prototype, ablations, reranker tests)Rejected with numbers: body tf in BM25F, fusion for packet seeds, cross-encoder on issues (details in contributions-log §8.3–8.4).
Twelve benchmark sets unchanged.
🤖 Generated with Claude Code