Welcome to Belief Updates
The Simplex Research Blog
At Simplex, we are working to build a fundamental science of intelligence. We want a concrete understanding of what it means for a system to represent, understand, and act in the world. Ultimately, we want a framework that applies as well to LLMs as it does to brains, and to intelligences that do not yet exist.
The deeper understanding that a fundamental science provides will require new ways of thinking about and orienting towards the nature of internal structure in these networks, and how that internal structure relates to the behaviors these systems carry out. For the last two years, our concrete entrypoint to this massive problem has been the activation geometry of neural networks.
As the field of interpretability has grown, an increasing number of scientific findings have found geometric structures within LLM activations [1, 2, 3, 4, 5, 6]. Much of this work remains empirically guided: finding interesting shapes inside models and trying to work backward to their meaning. In the face of this growing empirical support, the field still lacks the fundamental conceptual and formal language for answering why these structures appear, what dictates their form, and how they relate to a network’s understanding of the world [see also 7, 8].
That is why we are launching Belief Updates. We chose that name for two reasons.
The first is that we’ve developed a theory that understands the shape of a neural network’s internal activations as the direct mathematical consequence of predicting over a structured world [9, 10]. When a network successfully predicts data generated by a structured environment, its internal activations naturally arrange into geometric structures corresponding to beliefs over the hidden states of the world, beliefs that update token by token as observations arrive in context [11].
The second reason is about the scientific approach we take at Simplex. We want to be precise enough to be possibly wrong. We are building a theory that makes concrete, testable, falsifiable predictions. We don’t just want to observe that geometric structure appears; we want a framework that predicts exactly why it appears, when it will form, and what precise shapes we should expect. This allows us to design experiments that can test whether our theoretical framework is correct or needs to be updated. It is hard to overstate how deeply we believe in this approach. Having theories amenable to tight empirical feedback is an important setup for rapid scientific progress.
In the process of developing our theoretical assumptions and empirically testing their implications, we have found places where our assumptions broke and needed to be revised [12, 13]. Our thinking about what these geometries are, and how they relate to computation, features, and circuits, has continually evolved. We have built up a way of thinking about these systems that we think is unique, useful1, and beautiful.
Interpretability work is often motivated by pragmatic goals, like building methods to monitor and control AI systems [e.g., 16]. But it is not obvious what the right relationship to systems potentially more intelligent and capable than us will be. If we are going to face that question well, we will need to understand the nature of cognition in both the artificial systems humanity is building, and also in ourselves.
Our ambitions as an organization are high. We intend to fundamentally change not only how interpretability researchers do their work, but also how the larger community thinks about what exactly intelligence is made of. In order to carry out this important work, we intend to grow our team to at least 50 people in the next two years2.
Simplex is made up of people from different backgrounds, all oriented toward this scientific mission. Here, we hope to share exactly what we mean by building a fundamental science of intelligence. This will take the form of blog posts and interactive demos. Some will accompany our papers. Some will explain the math and concepts we find ourselves excited by. Some will be hot takes. And some will simply be us thinking out loud about something we don’t yet understand.
We want each post to be either surprising or to say something we believe that hasn’t quite been said elsewhere3. Across all of them, we aim to be serious about the science, playful with the ideas, and honest about what we don’t know. We hope that this presentation of our thinking and work will be more true to the science, and to the standard of rigor we hold ourselves to, than papers alone can be.
Acknowledgments
Paul Riechers co-founded Simplex, and his thinking runs through every part of this post. The way of thinking described here belongs to the whole Simplex team, past and present. Thanks in particular to Eric Alt, Loren Amdahl-Culleton, Casper Christensen, Asvin Gothandaraman, Selma Mazioud, Eric Michaud, Xavier Poncini, Kyle Ray, McNair Shah, and Jasmina Urdshals for comments on drafts. Loren Amdahl-Culleton and Eric Michaud built the site, and Eric Michaud, Kyle Ray and Casper Christensen shaped and ran the pipeline behind it.