HuggingFaceNmtEngine builds beam search alignments from the wrong beam at each step. It happens when num_beams is greater than 1 and output_attentions is enabled, on models with a decoder start token, such as NLLB and M2M100. Alignments are wrong wherever the selected hypothesis came from a different beam.
Affected: _forward in the translation pipeline in machine/translation/huggingface/hugging_face_nmt_engine.py.
Versions: main (transformers 4.47.1) and #363 (transformers 5.14.1).
Cause: beam_indices has one entry per generated token, starting with the first token after the decoder start token. The engine slices beam_indices[:, start_index:], which drops the first generated token's entry and shifts the rest by one step. The last step gets -1 % num_beams on main and beam 0 on #363.
Evidence: with stas/tiny-m2m_100, num_beams=2 and max_length=10 on transformers 5.14.1, sequences has shape [1, 10] and beam_indices has shape [1, 9], with values [[0, 0, 0, 0, 0, 0, 1, 0, 0]]. transformers 4.47.1 fills beam_indices the same way (indices[i, : len(best_idx)] in generation/beam_search.py), with a trailing -1.
Expected: step i uses the beam in beam_indices[:, i], with no shift.
Test: the expected alignments in test_translate_n_batch_beam come from the current output, so they do not catch this. A test with a hypothesis that switches beams would.
HuggingFaceNmtEnginebuilds beam search alignments from the wrong beam at each step. It happens whennum_beamsis greater than 1 andoutput_attentionsis enabled, on models with a decoder start token, such as NLLB and M2M100. Alignments are wrong wherever the selected hypothesis came from a different beam.Affected:
_forwardin the translation pipeline inmachine/translation/huggingface/hugging_face_nmt_engine.py.Versions:
main(transformers 4.47.1) and #363 (transformers 5.14.1).Cause:
beam_indiceshas one entry per generated token, starting with the first token after the decoder start token. The engine slicesbeam_indices[:, start_index:], which drops the first generated token's entry and shifts the rest by one step. The last step gets-1 % num_beamsonmainand beam 0 on #363.Evidence: with
stas/tiny-m2m_100,num_beams=2andmax_length=10on transformers 5.14.1,sequenceshas shape[1, 10]andbeam_indiceshas shape[1, 9], with values[[0, 0, 0, 0, 0, 0, 1, 0, 0]]. transformers 4.47.1 fillsbeam_indicesthe same way (indices[i, : len(best_idx)]ingeneration/beam_search.py), with a trailing-1.Expected: step i uses the beam in
beam_indices[:, i], with no shift.Test: the expected alignments in
test_translate_n_batch_beamcome from the current output, so they do not catch this. A test with a hypothesis that switches beams would.