published-canonicalconceptmaha-epistemic/1.0

Sparse autoencoder dictionaries

The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope. Within this page, that proposition is limited to Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.

Substantial reference · 9 evidence dimensions · maha-substantial-publication/1.1

Bounded definition

The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope. Within this page, that proposition is limited to Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.

Definition and evidence boundary

A source-bounded concept record for sparse autoencoder dictionaries within mechanistic interpretability. The bounded proposition retained by the canonical record is: The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope.

The applicable scope is Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies. This definition must not be generalized beyond the cited source and exact record boundary.

Claims: urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries

Mechanism and technical context

The paper trains sparse autoencoders on language-model activations and evaluates specified reconstruction, sparsity, and interpretability properties. This is the source-bound technical context for the record; no uncited mechanism is added by the compiler.

Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness. The mechanism or method is therefore presented as one component of a larger system, not as evidence for every downstream outcome.

Claims: urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries

How to interpret the evidence

No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review. The evidence maturity recorded here is single study, and the claim kind is theoretical model.

Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract. Sparse features and human labels do not establish completeness, unique decomposition, or causal faithfulness. These qualifications travel with the claim whenever it is reused.

Claims: urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries

What the source supports and what remains unknown

The inspected source supports exactly this: The paper trains sparse autoencoders on language-model activations and evaluates specified reconstruction, sparsity, and interpretability properties. It was read at Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations.

What remains unknown is everything outside that locator. Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness. No quantity, comparison, or downstream outcome is established here unless a separately scoped record measures it.

Claims: urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries

Comparison and calculation boundary

Applicability is decided explicitly, not filled with generic material.

Comparison · not-applicable

This record carries 1 source-bound proposition and therefore has no second supported side. A comparison would have to be manufactured from an adjacent title rather than from a second inspected claim, which the gate forbids.

Calculation · not-applicable

The canonical claim declares no reproducible numerical inputs, equation, units, or uncertainty propagation; recorded uncertainty kind is qualitative. Supplying sample values would invent an unsupported quantitative result.

Limitations and prohibited inference

The claim stops where its evidence stops.

  • record boundary

    Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.

  • record boundary

    A source-bounded mechanism, method, or measurement record does not establish manufacturing yield, economic advantage, safety, clinical benefit, or commercial readiness unless those outcomes are measured in a separately scoped record.

  • prohibited inference

    Do not use this sparse autoencoder dictionaries record to claim that the surrounding technology is proven, safe, scalable, commercially available, or strategically superior.

  • prohibited inference

    Do not transfer a reported result across hardware, organisms, protocols, datasets, operating conditions, or outcome definitions without a declared comparison contract.

  • editorial

    This compilation reorganizes an existing inspected claim and its declared source; it does not add a new experiment, measurement, or independent replication.

  • editorial

    Internal editorial inspection is not external peer review, and no result on this page has been independently reproduced.

Related records and mathematical bridges

application

Neural feature superposition

Declared strategic-dependency edge from this record. The edge is navigational and asserts no equivalence or causation beyond the cited source scope.

Selection: bridge edge

mechanism

Representation probing boundary

Declared mechanistic-dependency edge from this record. The edge is navigational and asserts no equivalence or causation beyond the cited source scope.

Selection: bridge edge

When no declared bridge edge is present, related records are linked by shared evidence or canonical domain adjacency. Those links are navigational and do not claim mathematical or physical equivalence.

Connected domain graph

Typed dependencies preserve publication state.

Only independently canonical records receive public links and relation statements. Draft graph topology remains private.

mechanistic dependencycanonical

Representation probing boundary

outbound connection · comparison

Sparse autoencoder dictionaries is positioned after Representation probing boundary in this bounded dependency sequence; the edge is navigational and does not assert equivalence or causation beyond the cited source scope.

strategic dependencycanonical

Neural feature superposition

outbound connection · concept

Sparse autoencoder dictionaries is connected to the cohort root so source, measurement, and readiness boundaries can be traversed without collapsing them.

Claim ledger

Every proposition keeps its own evidence state.

theoretical-modelsingle-study

The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope.

Scope
Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.
Boundary
Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.
Uncertainty
No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review.
Replication
Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.

Primary sources

Citation, locator, rights, and boundary travel together.

  1. Source 1 · arXiv

    Sparse Autoencoders Find Highly Interpretable Features in Language Models

    Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, Lee Sharkey

    Exact locator
    Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations.
    Establishes
    The paper trains sparse autoencoders on language-model activations and evaluates specified reconstruction, sparsity, and interpretability properties.
    Boundary
    Sparse features and human labels do not establish completeness, unique decomposition, or causal faithfulness.
    Rights basis
    citation with paraphrase · The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.