Amirali Abdullah

“The game of games had developed into a kind of universal language through which the players could express values and set these in relation to one another.”

— Hermann Hesse

Amirali Abdullah

AI Researcher · Mechanistic Interpretability · Representation Geometry

I am an AI researcher at Thoughtworks focused on mechanistic interpretability, activation steering, and representation geometry in foundation models. My work explores how internal model structure can be understood, measured, and leveraged for safer and more controllable AI systems.

A central theme in my research is the geometry of learned representations: how features are organized in activation space, how steering vectors interact with this geometry, and whether the resulting structure is consistent across models. I am particularly interested in non-linear approaches to representation steering and the alignment of geometric structure across model families.

I collaborate across academia and industry on interpretability tooling, psychometric evaluation of LLMs, and multimodal representation analysis. Previously, I worked in theoretical computer science on nearest neighbor search, dimensionality reduction, and information-theoretic geometry — problems whose core ideas around metric structure and spectral methods continue to inform my current work on feature spaces in neural networks.

Outside of research, I make a comic strip with my daughter explaining AI concepts to kids.

Affiliations: Thoughtworks

News

Research

Mechanistic Interpretability

Understanding the internal computations of large language models through sparse autoencoders, feature geometry, and circuit-level analysis.

Activation Steering & Control

Developing methods to steer and control model behavior through representation-space interventions, including multi-attribute and cross-model transfer of steering vectors.

Psychometric Evaluation of LLMs

Applying psychometric methods to rigorously evaluate personality-like traits, sycophancy, and refusal behavior in language models.

Publications

Full list on Google Scholar.

2026
Make Mechanistic Interpretability Auditable
ACL 2026
Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi, Philip Quirke, Austin Meek, Fazl Barez, Amir Abdullah
2026
Refusal-Gated Decoding
AdvML-Frontiers @ COLM 2026
Phillip Howard, Xin Su, Allen Roush, Manikandan Ravikiran, Amir Abdullah
2026
Modes of Sycophancy
Preprint
Shreyans Jain, Alexandra Yost, Amirali Abdullah
2026
The Capability Frontier
Preprint
Bradley Fowler, Ryan Smith, Daniel Thi Graviet, William Myers, Joshua Greaves, Narmeen Fatimah Oozeer, Antía García, Philip Quirke, Amirali Abdullah, Fazl Barez, Shriyash Kaustubh Upadhyay
2026
STRIDE
Preprint
Rishit Dagli, Abir Harrasse, Luke Zhang, Florent Draye, Amirali Abdullah, Bernhard Schölkopf, Zhijing Jin
2026
Riemannian-Manifold Steering
Preprint
Narmeen Oozeer, Shivam Raval, Philip Quirke, Manikandan Ravikiran, Jeff Phillips, Shriyash Upadhyay, Amirali Abdullah
2026
Dynamic Latent Routing
Preprint
Fangyuan Yu, Xin Su, Amir Abdullah
2026
Geometry-Aware CLIP Retrieval
Preprint
Nirmalendu Prakash, Narmeen Fatimah Oozeer, Xin Su, Phillip Howard, Shaan Shah, Zoe Wanying He, Shuang Wu, Shivam Raval, Roy Ka-Wei Lee, Meenakshi Khosla, Amir Abdullah
2026
Spectral Superposition
Under Preparation
Georgi Ivanov, Narmeen Oozeer, Shivam Raval, Tasana Pejovic, Shriyash Upadhyay, Amir Abdullah
2026
LLM Refusal
AAAI 2026
Nirmalendu Prakash, Wei Jie Yeo, Amir Abdullah, Ranjan Satapathy, Erik Cambria, Roy Ka-Wei Lee
2026
DreamReader
Preprint
Nirmalendu Prakash, Narmeen Oozeer, Michael Lan, Luka Samkharadze, Phillip Howard, Roy Ka-Wei Lee, Dhruv Nathawani, Shivam Raval, Amirali Abdullah
2026
Curveball Steering
Preprint
Shivam Raval, Hae Jin Song, Linlin Wu, Abir Harrasse, Jeff M. Phillips, Fazl Barez, Amirali Abdullah
2026
Robust Steering
Preprint
Cullen Anderson, Narmeen Oozeer, Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Jeff M. Phillips
2026
2025
Beyond Linear Steering
EMNLP 2025 (Findings)
Narmeen Oozeer, Luke Marks, Shreyans Jain, Fazl Barez, Amirali Abdullah
2025
TinySQL
EMNLP 2025
Abir Harrasse, Philip Quirke, Clement Neo, Dhruv Nathawani, Luke Marks, Amir Abdullah
2025
Activation Transfer
ICML 2025
Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Michael Lan, Abir Harrasse, Amir Abdullah
2025
Sycophancy
EMNLP 2025 (BlackboxNLP Workshop)
Shreyans Jain, Alexandra Yost, Amirali Abdullah
2025
Distribution-Aware SAEs
Preprint
Narmeen Oozeer, Nirmalendu Prakash, Michael Lan, Alice Rigg, Amirali Abdullah
2025
Beyond Monoliths
Preprint
Philip Quirke, Narmeen Oozeer, Chaithanya Bandi, Amir Abdullah, Jason Hoelscher-Obermaier, Jeff M. Phillips, Joshua Greaves, Clement Neo, Fazl Barez, Shriyash Upadhyay
2025
Psychometric Evaluation
Preprint
Alexandra Yost, Shreyans Jain, Shivam Raval, Grant Corser, Allen Roush, Nina Xu, Jacqueline Hammack, Ravid Shwartz-Ziv, Amirali Abdullah
2025
2024
Feedback Patterns
NeurIPS 2024
Luke Marks, Amir Abdullah, Clement Neo, Rauno Arike, David Krueger, Philip Torr, Fazl Barez
2024
Earlier Work (Theoretical CS)
Isoperimetric Inequality
Spectral NN Search
FOCS 2014
Amirali Abdullah, Alexandr Andoni, Ravindran Kannan, Robert Krauthgamer
2014
+ additional publications at AISTATS 2016, SoCG 2013, ISAAC 2016, and IJCGA 2013. See Google Scholar for full list.

Projects

DreamReader

Interpretability toolkit for text-to-image models analyzing cross-modal concept structure.

Spectral Superposition

Studying eigenspace structure and feature overlap in neural representations.

Comic

A comic series explaining AI research concepts for kids and curious minds.

Mechanistic Interpretability: Explaining AI's Mind for Kids! — Page 1
Page 1 — Mechanistic Interpretability: Explaining AI's Mind for Kids!

Contact

Email: amir.abdullah@thoughtworks.com

Location: Utah, USA