Workshop focus

From prediction to discovery

Interpretability is more than a debugging tool. It can bridge black-box models and scientific breakthroughs across unfamiliar domains, architectures, and modalities.

Premise

Explanation is only the beginning.

Explore the scope ↓

AI models now match or exceed human experts across many domains, from protein structure prediction and clinical forecasting to astronomy, climate science, animal communication, and strategic games.

To reach that performance, they appear to have learned patterns and regularities that humans have not yet articulated. Interpretability gives us a way to read those representations back out.

This workshop reframes interpretability as a tool for discovery, not just explanation. We bring together interpretability researchers and scientists to turn what models encode into knowledge that domain experts can test and validate.

Core questions

What must we solve?

Discovery requires more than a compelling visualization. It requires methods, translation, and validation.

Question 01

Adapting modalities

How can interpretability methods adapt to architectures, inductive biases, and data modalities beyond standard language and vision models?

Question 02

Translating knowledge

How can interpretable internal representations become actionable, testable new knowledge?

Question 03

Open scientific problems

What questions can interpretability help answer that are currently unknown to humans?

Intellectual scope

Workshop topics

We invite work at the boundary between understanding models and producing knowledge about the world.

01

Methods and models for knowledge discovery

Interpretability methods, evaluation frameworks, and model designs that support discovery across unfamiliar architectures and modalities.

02

Interpretability-driven discovery

Empirical case studies where examining model internals surfaces non-obvious, verifiable knowledge in science and expert domains.

03

Perspectives on novel knowledge

Position papers and theory on scope, limitations, epistemology, validation, and open problems in interpretability for discovery.

04

Failure cases and negative results

Careful analyses of unsuccessful approaches, misleading interpretations, failed validation, and practical limits that clarify when interpretability does or does not support reliable discovery.