I love understanding the inner workings of neural nets to make them safe and useful.
Pinned Loading
-
mechanistic_interpretability
mechanistic_interpretability PublicMechanistic interpretability experiments: raw GPT-2 inference, activation steering, and manual backprop MLP.
Jupyter Notebook
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.
