Ayodele Abraham
Hi, I am Ayodele. I am currently transitioning into AI Safety, with interest in Mechanistic Interpretability and auditing frontier language models. Exploring how AI systems work, fail, and can be made safer.
Latest Writing
AI Safety & Research
My Journey Through Technical AI Safety: A BlueDot Course Retrospective
A complete retrospective on what I learned, what I built, and what changed in how I think about AI safety after completing the BlueDot Technical AI Safety course.
Building an Input/Output Safety Classifier Pipeline (and Trying to Break It)
I built a two-checkpoint I/O safety classifier pipeline using Gemma 3 and Llama Guard 3, then red-teamed it with universal jailbreaks and targeted contextual prompts. 0 of 9 universal attempts achieved a full bypass. Only 1 of 3 targeted contextual attempts did, and the follow-up test revealed why that result is more nuanced than it first appeared.
Continue ReadingWhat's Actually Happening Inside an AI?
Mechanistic interpretability is the attempt to open the black box. Let's try to understand it together.
Publications
Peer-reviewed Research