Close out spec 004: implemented, #140 and #125 closed with credit - #206
Merged
Conversation
T012-T014. All 15 tasks complete. Each contributor was told specifically what of theirs was used. #140's instruction wording is in language_instruction.py nearly verbatim, because it got the part that is easy to miss: nomenclature must survive untranslated, and SET, MAX and CAT are gene symbols as well as English words. #125's mechanism -- a prompt variable rather than a concatenated query -- is the one #205 uses. #140 was also told where the 0-of-10 figure came from and that it was wrong, since a dramatic number in a rejection deserves an explanation. #125 was told plainly that closing it is not a decision about its hallucination grader, which belongs to #123 and is worth having. Records both wrong claims and the lesson behind the second: the first review checked whether each claim was true without checking whether the test measured the product, which is the same failure evaluator.py had four days earlier. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
All 15 tasks complete. Both contributed PRs closed, each told specifically what of theirs was used.
@bleedblack1 (#140) — the instruction wording is in
language_instruction.pynearly verbatim. It got the part that is easy to miss: nomenclature must survive untranslated, andSET,MAXandCATare gene symbols as well as ordinary English words. They also picked the right profile.They were also told where the
0 of 10figure came from and that it was wrong — a dramatic number in a rejection deserves an explanation, and the honest figure is half the context, not all of it.@bhavyakeerthi3 (#125) — the mechanism is theirs: a prompt variable rather than a concatenated query. They also kept the translate-to-English step while rewriting the rephrase prompt, which was the obvious thing to get wrong. And they were told plainly that closing it is not a decision about their hallucination grader, which belongs to #123 and is worth having.
Two wrong claims recorded, and the lesson
The second is the useful one. The first review checked whether each claim was true without checking whether the test measured the product — the same failure
evaluator.pyhad four days earlier.Worth watching: the French answer carried 2
R-HSAcitations against English's 9. One sample, so not a finding — but if non-English answers systematically cite less, this feature closes a language gap while opening a quality one.🤖 Generated with Claude Code