Official code for the paper "On Mitigation of Subliminal Learning in Large Language Models"
-
Updated
Sep 1, 2026 - Python
Official code for the paper "On Mitigation of Subliminal Learning in Large Language Models"
From MNIST noise distillation to Llama-3.1-70B: subliminal learning, token entanglement, causal state transfer, and multi-token confounds.
Code to reproduce the results of the paper "Learning Through Noise: Why Subliminal Learning Works and When It Fails"
From-scratch NumPy MLPs for MNIST, plus a larger teacher–student setup that demonstrates subliminal learning (implementing Cloud et al. (2025)).
Replicating latent-space backdoor leakage and behavioral transfer in LLMs using Pythia-70M.
Blind attribution of the principal behind a covertly poisoned training corpus, against 47 candidates. Secret Loyalties hackathon, Jul 2026; corrected Sep 2026.
To associate your repository with the subliminal-learning topic, visit your repo's landing page and select "manage topics."