DadgogoAI Dictionary

Dictionary entry Trending

sparse autoencoder

noun \ˈspärs ˌȯ-tō-in-ˈkō-dər\

also SAE

Advertisement

Definition of sparse autoencoder

  1. : an autoencoder trained so that only a small fraction of its learned features activate for any one input, often used to decompose a neural network's internal activations into more interpretable features

    • Researchers trained a sparse autoencoder on the language model's activations to find a feature associated with references to a particular city.
    • The SAE produced thousands of candidate features, but each still required testing before it could be given a reliable human-readable label.

Two meanings, one word

In everyday English

Sparse means thinly distributed or containing relatively few active elements; an autoencoder is a neural network trained to compress and reconstruct data.

In AI

An autoencoder trained so that only a small fraction of its learned features activate for any one input, often used to decompose a neural network's internal activations…

Origin & history

first known use2023

Sparse autoencoders predate large language models, but the abbreviation SAE became prominent in mechanistic-interpretability research in 2023 after researchers showed that such models could recover more interpretable features from language-model activations. Anthropic's 2024 Scaling Monosemanticity work helped bring the method to wider attention.

Test yourself

Which of these is the meaning of sparse autoencoder?

Cite this entry

"Sparse autoencoder." AI Dictionary, Dadgogo, https://dadgogo.com/dictionary/sparse-autoencoder/. Accessed 7 Oct. 2026.

Dictionary entries near sparse autoencoder

Advertisement