Dictionary entry Trending
sparse autoencoder
also SAE
Definition of sparse autoencoder
-
: an autoencoder trained so that only a small fraction of its learned features activate for any one input, often used to decompose a neural network's internal activations into more interpretable features
- Researchers trained a sparse autoencoder on the language model's activations to find a feature associated with references to a particular city.
- The SAE produced thousands of candidate features, but each still required testing before it could be given a reliable human-readable label.
Two meanings, one word
In everyday English
Sparse means thinly distributed or containing relatively few active elements; an autoencoder is a neural network trained to compress and reconstruct data.
In AI
An autoencoder trained so that only a small fraction of its learned features activate for any one input, often used to decompose a neural network's internal activations…
Origin & history
Sparse autoencoders predate large language models, but the abbreviation SAE became prominent in mechanistic-interpretability research in 2023 after researchers showed that such models could recover more interpretable features from language-model activations. Anthropic's 2024 Scaling Monosemanticity work helped bring the method to wider attention.
Test yourself
Which of these is the meaning of sparse autoencoder?
Cite this entry
"Sparse autoencoder." AI Dictionary, Dadgogo, https://dadgogo.com/dictionary/sparse-autoencoder/. Accessed 7 Oct. 2026.