Neel Nanda
AI safety / mechanistic interpretability researcher. Leads the DeepMind mechanistic interpretability team; formerly at Anthropic (under Chris Olah), independent during a 2022-2023 sabbatical, and an AI-safety intern at FHI/Oxford, DeepMind and CHAI. UK IMO medallist (Bronze 2015, Gold 2016-2017) and IOI 2017 competitor; Cambridge mathematics undergraduate who did quantitative-trading internships at Jane Street and Jump Trading before pivoting to AI safety.
- Quant employment verification: CONFIRMED. Identity is supported (LinkedIn 'neel-nanda-993580151', his own site and verified X all place the same Neel Nanda at DeepMind with prior Anthropic and Cambridge roles). Actual quant roles are paid summer internships at proprietary trading / market-making firms: Quantitative Trading Intern at Jane Street (2018), Quantitative Research Intern at Jump Trading LLC (2019), Quantitative Research Intern at Jane Street (2020), all London. Jane Street and Jump Trading are quantitative proprietary-trading and market-making firms, i.e. genuine quantitative finance. The 80,000 Hours story independently confirms the quant internships as real work. The Morgan Stanley Spring Week (2018) is a bank insight programme, not a quant role. The employer-level role records are LinkedIn-index-sourced (index-only); the employer classification rests on those index records, corroborated at the category level by the independent interview.
- Runs the Google DeepMind mechanistic interpretability team, whose work is to reverse-engineer the algorithms and structures learned by trained neural networks; based in London.
- Worked at Anthropic (2021-2022) as a language-model interpretability researcher under Chris Olah, contributing to the transformer-circuits line of work.
- The LinkedIn index (index-only, not independent primary-source confirmation) records three quantitative-trading internships: Quantitative Trading Intern at Jane Street (2018, London), Quantitative Research Intern at Jump Trading LLC (2019, London) and Quantitative Research Intern at Jane Street (2020, London).
- The 80,000 Hours career story independently confirms he 'spent his summers doing internships in quantitative finance', which corroborates the trading internships as real paid work rather than competitions or insight days; it does not itemise the employers.
- Turned down quantitative finance and a Cambridge maths master's to take a year of AI-safety internships (Future of Humanity Institute at Oxford, DeepMind, Center for Human-Compatible AI) - the pivot that started his interpretability career.
- Bio and X profile both describe him as 'Mechanistic Interpretability lead' at DeepMind and a former member of Anthropic, so the research identity is self-declared across two independent surfaces.
- After leaving Anthropic he took a self-described sabbatical (2022-2023) doing independent interpretability research, notably on 'grokking'.
- Authored and maintains TransformerLens, the standard open-source library for mechanistic interpretability of transformer language models, plus research repos Grokking, Neuroscope and neel-plotly on GitHub (neelnanda-io).
- Mentors researchers through his MATS stream, a full-time twice-yearly (summer and winter) interpretability research program.
- International Mathematical Olympiad record for the United Kingdom: Bronze 2015 (17 pts), Gold 2016 (30 pts) and Gold 2017 (25 pts).
- Also a Spring Week Internship in Technology at Morgan Stanley (2018, London) - a bank insight programme rather than a quantitative role - recorded alongside the trading-firm internships.
Researcher on the DeepMind Mechanistic Interpretability team
On Sabbatical after leaving my role at Anthropic. Working on independent interpretability research, currently trying to understand what's going on with grokking. Published Progress
Researching language model interpretability under Chris Olah
Competition record
United Kingdom · IOI