Goodfire finds sparse autoencoder features capture curved neural geometry in three ways
Original titleCan SAEs Capture Neural Geometry? - Goodfire
AISummary
Goodfire Research examines how sparse autoencoder directions relate to curved manifolds in neural representations, identifying shattering, compact capture, and dilution as three ways lines can represent them.
The team trained an autoencoder on synthetic data containing shapes such as donuts, spheres, and Möbius strips, and reports that real features in Llama 3.1 8B show dilution.
It also describes an unsupervised pipeline that clusters features by firing patterns to surface manifolds in that model.
Source: Goodfire Research · goodfire.comPublished · added here