AI models now match or exceed human experts across many domains, from protein structure prediction and clinical forecasting to astronomy, climate science, animal communication, and strategic games.
To reach that performance, they appear to have learned patterns and regularities that humans have not yet articulated. Interpretability gives us a way to read those representations back out.
This workshop reframes interpretability as a tool for discovery, not just explanation. We bring together interpretability researchers and scientists to turn what models encode into knowledge that domain experts can test and validate.