Discovering Concept-Editing Algorithms With LLM Agents
Scherlis, Bishka, et al., 2026
We tasked LLM agents with inventing concept-erasure algorithms. They discovered six algorithm families across 50 concepts, all of which out-perform LEACE on a non-linear probe.
How Language Models Understand Nullability
Sanchez-Stern and Tondwalkar, 2025
We study how models represent the nullability of program values. We measure how well models of
various sizes, at various training checkpoints, complete programs that use nullable values, and then
extract an internal representation of nullability.