flâneur

Student Projects - CS 2881R AI Safety | CS 2881 AI Safety

boazbk.github.io · 1,384 words · saved by 1 readers

Harvard CS 2881R

Harvard CS 2881R Student Final Projects - Fall 2025 This page showcases the final research projects from CS 2881R: AI Safety. Students conducted original research on topics spanning interpretability, alignment, adversarial robustness, and AI governance. Video of oral presentations Mechanisms of Subliminal Learning Subliminal learning is a recently discovered failure mode of distillation and post-training where a student model inherits a teacher's hidden traits (e.g., "liking owls") from data that appears semantically unrelated (e.g., number lists). We study the mechanisms of subliminal…

saved by

related reading