Projects
Research · Education · Software
Secret Loyalty Auditing
Submitted to the Apart Research AI Safety Hackathon. The project investigates whether hidden objectives in synthetic LLM organisms can be detected through mechanistic interpretability techniques. Specifically, it evaluates how auditor affordance levels influence the linear decodability of hidden objectives from transformer activations, and explores lightweight auditing protocols based on activation probing.
Independent Mechanistic Interpretability Research
Self-directed research into transformer circuits and feature geometry. Currently studying superposition, how models represent more features than they have dimensions, using toy model experiments and sparse autoencoder analysis. Exploring implications for feature identification and circuit-level auditing.
ARENA Curriculum
Working through the Alignment Research Engineer Accelerator (ARENA) curriculum. Covers transformer architecture from scratch, mechanistic interpretability techniques, reinforcement learning from human feedback, and agent foundations. Self-paced, with implementation exercises in PyTorch.
Lens Academy, Course Navigator
Volunteering as a course navigator at Lens Academy, helping participants work through structured AI safety and alignment learning pathways. Working at the intersection of safety research and education, making interpretability and alignment concepts accessible to newcomers.
Lerna
A personal AI curriculum architect. Lerna generates structured, personalised learning paths for technical topics using language models. Built to support self-directed learning in fields like mechanistic interpretability where reading lists and curricula are evolving rapidly.