About Me

In pursuit of a world where everyone wants to cooperate in prisoner's dilemmas.

I work full-time on making AI go well, having paused my MSc in Artificial Intelligence at the University of Groningen to do so. I believe that within a worryingly short time frame AI will become more powerful, and currently, by default, more uninterpretable. Combined with the enormous incentives to prioritize capabilities over safety, this creates a serious risk of developing systems we do not fully understand and cannot reliably control.

I work on this by building organizations; previously, I also did technical research. I direct Safe AI Netherlands (SAIN), the national AI safety field-building organization with chapters in Groningen, Utrecht, and Amsterdam, and new chapters coming in Eindhoven, Delft, and The Hague. In 2026, we received a $2M two-year grant, and we now run on a team of 4 full-time staff (including me), 10 part-timers, and 30 volunteers. Through SAIN we educate hundreds of people, run a Project Hub whose research has appeared at NeurIPS and ICLR workshops, and collaborate with Dutch municipalities. I am also co-founding and co-directing the AI Safety Hub Amsterdam at The Stack, Amsterdam's new AI hub. On the technical side, I have worked on multi-agent game theory, mechanistic interpretability, and agentic LLM behavior under EU legislation.

Education

MSc Artificial Intelligence

University of Groningen · 2025 – 2027 (paused)

GPA: 8.8/10

Focus on AI alignment, mechanistic interpretability, and game theory. Paused in 2026 to lead SAIN full-time.

BSc Artificial Intelligence

University of Groningen · 2022 – 2025

GPA: 8.9/10, Cum Laude

BSc thesis on transformer-based chemical foundation models, which led to a follow-up publication.

Research Interests

Mechanistic Interpretability

Understanding how neural networks work internally. If we can't look inside these systems and understand what they're doing, we can't trust them. See this paper.

Multi-Agent Game Theory

How game-theoretic frameworks can help us think about multi-agent alignment and cooperation, i.e., how to defeat Moloch. See this paper.

LLM Agent Evaluations

Measuring how LLM agents behave when deployed in realistic settings, including whether they break the law. See this paper.

AI Governance & Field-Building

Governance frameworks, EU law compliance, and growing organizations that take AI safety seriously.

AI Alignment

How do we make sure AI systems actually do what we want? The problem is harder than it sounds. See this post.

What I'm Working On

  • ● Directing Safe AI Netherlands (SAIN), funded with a $2M two-year grant: growing the team to 4 full-time staff (including me), 10 part-timers, and 30 volunteers, and expanding to Eindhoven, Delft, and The Hague.
  • ● Co-founding the AI Safety Hub Amsterdam at The Stack, Amsterdam's new AI hub, which I co-direct.
  • ● Writing on my Substack and SAIN's Substack about AI safety, philosophy, books, and more.

CV

For my complete educational background, work experience, and research projects:

Alexander Müller

Location

Amsterdam, Netherlands

Education

MSc Artificial Intelligence (paused)

University of Groningen

Role

Director, Safe AI Netherlands (SAIN)

Co-founder & Co-Director, AI Safety Hub Amsterdam