Exploring How Reasoning Models Break Mechanistic Interpretability Techniques

Exploring How Reasoning Models Break Mechanistic Interpretability Techniques reveals several interesting facts.

  • LLMs that can "think" and "reason" have become increasingly popular. But what is a
  • This is a talk I gave to my MATS 9.0 training scholars about the big picture of mech interp - as of Oct 2025, what had changed?
  • Lex Fridman Podcast full episode: https://www.youtube.com/watch?v=ugvHCXCOmm4 Thank you for listening ❤ Check out our ...
  • Large Language
  • Deep neural networks successfully power modern AI, but they operate as black boxes. In traditional software, we can read the ...

In-Depth Information on How Reasoning Models Break Mechanistic Interpretability Techniques

A talk I gave to my MATS 9.0 training program about EuroPython 2025 — South Hall 2B on 2025-07-17] *Hacking LLMs: An Introduction to Take your personal data back with Incogni! Use code WELCHLABS at the link below and get 60% off an annual plan: ... A discussion on the philosophy of deep learning,

Why are some

Stay tuned for more updates related to How Reasoning Models Break Mechanistic Interpretability Techniques.

How Reasoning Models Break Mechanistic Interpretability Techniques.pdf

Size: 14.37 MB · Format: PDF · Secure Download

Download PDF Read Online

Related Documents