Mechanistic interpretability, done from outside the field and published with its null results
I study what a language model does on the inside. There is an original paper with its PDF and interactive demo, the first consolidated Spanish-language course on the subject, and a new project published before it has results. I come from eighteen years in energy, and that is where this points.
Frontier-model use outside sanctioned controls is already measured inside critical-infrastructure organizations. Once agents move from advising to acting on a pipeline or a grid, someone will have to show when that is safe. I spent eighteen years inside that industry, and it strikes me as where this ends up.
What I cannot claim: nothing I research today applies to an operating decision. My experiments separate "Italy" from "capital-of" in the internal state of a small model, in English, over toy factual relations. Between that and auditing an agent with authority over a pipeline there is a gap I would rather state than paper over.
Agenda
What comes next, in order, and where it stalls
Each rung rests on the one below it and none of them is climbed yet. The first two follow from lines the paper leaves open; the third is a different size of problem.
Causal diagnosis of an error
When a model gets something wrong, tell apart whether it confused the thing, picked the wrong operation, or failed to bind one to the other. The paper shows those parts can be manipulated separately; what is missing is a diagnosis that holds when you don't already know the right answer.
Reading intent, not the act
Detect which operation a model is about to run, rather than auditing the one it already ran. This travels with its counterpoint, which does not get published separately: a direction you can monitor is a direction someone else can push on.
Leaving the toy domain
A country and its capital have one unambiguous answer; an engineering task does not. That step is not a corollary of the previous two but a project of its own, and it is the point where this needs funded time or a group already working inside.
The six pillars of AI safety, where interpretability sits among them, and the long argument behind these rungs: the full map (in Spanish).
A model's thought about a fact splits into parts (the thing, and the question asked of it), and those parts suffice: average them, recombine them, write them back inside, and the model speaks the answer. Three models, two domains where it works and two where it provably doesn't. Interactive demo included.
The first consolidated course on the subject in Spanish: six pages from "what is a token" to attribution graphs. Alongside it, a map of the six pillars of AI safety, ten open research lines, and resources curated link by link.
The lab notebook: 989 lines in chronological order with the experiments, the measurement bugs I hit while building, and the replication that died under matched controls and turned into the paper. Every number recomputes from the artifacts.
The newest one, still without results: residual-gauge, importing a basis for the internal state from a symmetry the task already has, with thresholds pre-registered before anything ran.
Who's behind this
Matías Podeley. Eighteen years in energy in Argentina: reservoir engineering, then business development. Now independent interpretability research. Buenos Aires.
The eighteen years of energy sitting behind this: profile.
If you work in this
Funders and collaborators
If you fund research in this area, run a fellowship or incubator, or work in AI safety and want a collaborator who knows energy operations from the inside: email me.