Source evidence reference

Toy Models of Superposition

Nelson Elhage, Tristan Hume, Catherine Olsson, et al. · Transformer Circuits Thread · 2022-09-14

Identifier https://transformer-circuits.pub/2022/toy_model/index.html · version inspected: version of record as cited

This page is a projection of Maha canonical record releases about one source. It is not a separately certified source assessment, not an independent review, and not a replication. Every finding below is carried from a released record and disappears from this page when that record is withdrawn.

What the source investigates

What Toy Models of Superposition investigates, and what Maha records released against it establish.

Evidence type: section or full-text inspection of the cited source

Sections inspected

  • Definitions, toy models, geometry, sparsity, and feature-interference experiments.

Access and rights basis: citation-with-paraphrase

Findings carried from released records

  • The cited source supports treating neural feature superposition as a distinct concept within the stated mechanistic interpretability scope.

    Released record · Definitions, toy models, geometry, sparsity, and feature-interference experiments.

  • The cited source supports treating polysemantic neurons as a distinct mechanism within the stated mechanistic interpretability scope.

    Released record · Definitions, toy models, geometry, sparsity, and feature-interference experiments.

  • The cited source supports treating toy models of superposition as a distinct method within the stated mechanistic interpretability scope.

    Released record · Definitions, toy models, geometry, sparsity, and feature-interference experiments.

  • The cited source supports treating superposition geometry as a distinct measurement within the stated mechanistic interpretability scope.

    Released record · Definitions, toy models, geometry, sparsity, and feature-interference experiments.

  • The cited source supports treating representation probing boundary as a distinct comparison within the stated mechanistic interpretability scope.

    Released record · Definitions, toy models, geometry, sparsity, and feature-interference experiments.

What this source does not establish

  • Independent replication of any result reported by the source.
  • Any claim from a Maha record that is not currently released.
  • Endorsement, peer review or expert consensus by Maha.

Limitations

  • A projection of released record claims. It adds no finding of its own.
  • Inspection reached the locators listed above and no further.

Typed bridges

  • mechanistic-dependency: undefined
  • mechanistic-dependency: undefined
  • mechanistic-dependency: undefined
  • mechanistic-dependency: undefined

Related released records

Provenance digest sha256:56c8e72f3cd9558fb7d2ab5060c613f7b8a174d693fdb2fdc828673072602a33 · route /knowledge/sources/https-transformer-circuits-pub-2022-toy-model-index-html