published-canonicalmechanismmaha-epistemic/1.0

Polysemantic neurons

Polysemantic neurons is represented as one reviewable unit in the Mechanistic interpretability graph. Its source, locator, scope, uncertainty, and prohibited inference remain attached to the claim rather than being generalized across the domain.

Bounded definition

A source-bounded mechanism record for polysemantic neurons within mechanistic interpretability.

What the cited work establishes

The work develops toy models in which neural networks represent more features than available dimensions under specified sparsity conditions.

Limited to Definitions, toy models, geometry, sparsity, and feature-interference experiments. in “Toy Models of Superposition”; this candidate records the concept boundary and does not pool results from uncited systems or studies.

Claims: urn:maha:claim:mechanistic-interpretability-polysemantic-neurons

What remains a separate question

Polysemantic neurons does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.

A toy-model mechanism does not establish that every feature in a production model has the same geometry or semantics.

Connected domain graph

Typed dependencies preserve publication state.

Only independently canonical records receive public links and relation statements. Draft graph topology remains private.

mechanistic dependencycanonical

Neural feature superposition

outbound connection · concept

Polysemantic neurons is positioned after Neural feature superposition in this bounded dependency sequence; the edge is navigational and does not assert equivalence or causation beyond the cited source scope.

mechanistic dependencycanonical

Toy models of superposition

inbound connection · method

Toy models of superposition is positioned after Polysemantic neurons in this bounded dependency sequence; the edge is navigational and does not assert equivalence or causation beyond the cited source scope.

Claim ledger

Every proposition keeps its own evidence state.

theoretical-modelsingle-study

The cited source supports treating polysemantic neurons as a distinct mechanism within the stated mechanistic interpretability scope.

Scope
Limited to Definitions, toy models, geometry, sparsity, and feature-interference experiments. in “Toy Models of Superposition”; this candidate records the concept boundary and does not pool results from uncited systems or studies.
Boundary
Polysemantic neurons does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.
Uncertainty
No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review.
Replication
Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.

Primary sources

Citation, locator, rights, and boundary travel together.

  1. Source 1 · Transformer Circuits Thread

    Toy Models of Superposition

    Nelson Elhage, Tristan Hume, Catherine Olsson, et al.

    Exact locator
    Definitions, toy models, geometry, sparsity, and feature-interference experiments.
    Establishes
    The work develops toy models in which neural networks represent more features than available dimensions under specified sparsity conditions.
    Boundary
    A toy-model mechanism does not establish that every feature in a production model has the same geometry or semantics.
    Rights basis
    citation with paraphrase · The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.