About
Hi, I am Ayodele. I am currently transitioning into AI Safety, and I am a part-time AI Safety Research Intern at Hux AI, where I am currently working on LLM Guardrails Evaluation Toolkit.
I see AI safety as one of the most important problems of this era. My focus within the space is mechanistic interpretability, taking a trained model and trying to reverse-engineer the algorithms and circuits it has learned. I am also interested in how interpretability findings can be operationalised into practical auditing protocols, ways of systematically testing whether a model is doing what it is claimed to be doing.
Before transitioning into AI safety research, I spent several years building software products. I completed the BlueDot Impact Technical AI Safety course, which introduced me to the alignment research landscape and led me toward mechanistic interpretability. I have been working through the ARENA curriculum independently, running experiments on transformer circuits and feature geometry. I am also a course navigator at Lens Academy, helping participants navigate structured AI safety learning pathways. I have two peer-reviewed publications in applied machine learning and manufacturing, and recently submitted a project to the Apart Research AI Safety Hackathon.
If you want to get in touch, you can email me at ayo@ayabraham.com or find me on LinkedIn. I am particularly interested in conversations about mechanistic interpretability, AI auditing, and alignment research.
About My WritingI write about AI safety, mechanistic interpretability, alignment, and the technical and philosophical questions at the frontier of AI development. Most of my writing is the product of trying to understand something precisely enough to explain it clearly. If I have written about a topic, it usually means I have spent time confused about it and found a way through.
I also write tutorial-style posts for concepts I had to reconstruct from scattered sources. If you want to know where to start, the AI Safety essays are the most relevant to my current research focus. The Tutorials section covers technical mechanics I have worked through carefully.
Research and ProjectsMy primary active project is Secret Loyalty Auditing, submitted to the Apart Research AI Safety Hackathon and currently under review. The project investigates whether hidden objectives in synthetic LLM organisms can be detected through mechanistic interpretability techniques. Specifically, it evaluates how auditor affordance levels influence the linear decodability of hidden objectives from transformer activations, and explores lightweight auditing protocols based on activation probing.
I maintain a set of independent mechanistic interpretability experiments studying superposition, feature geometry, and transformer circuits. You can find these on my GitHub. I have also built Lerna, an open-source personal AI curriculum architect that generates structured learning paths for technical topics, and meBible, an AI-powered Bible study application.
My two peer-reviewed publications are in applied machine learning and advanced manufacturing. You can read them on the Publications page.
Questions I Am Thinking AboutHow do we distinguish a model that has genuinely internalised a value from one that has learned to mimic the outputs of an agent with that value? Can sparse autoencoders identify features that are causally responsible for harmful behaviours, or only features that correlate with them? What is the right unit of analysis for interpretability work, and does it depend on the capability being studied? What oversight protocols remain meaningful when the model is more capable than the human evaluator at the task being evaluated?
Long-term GoalsI want to contribute to the mechanistic understanding of frontier AI systems, with a particular focus on what interpretability can tell us about alignment properties. That means publishing reproducible experiments, building toward a research role at an organisation working directly on this problem, and contributing to the tools and evaluation protocols the field will need as models become more capable.
I try to be precise about what I know and honest about what I do not. The most useful thing I can do right now is understand one specific part of this problem well enough to contribute something that holds up.