{"schemaVersion":"maha-epistemic/1.0","evidencePolicyVersion":"mps/0.1","recordId":"urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries","canonicalPath":"/knowledge/mechanistic-interpretability/concepts/mechanistic-interpretability-sparse-autoencoder-dictionaries","contentHash":"sha256:4685df47be76d62b550e5cab12a91f0fbaabc8c6db581c07bc660a59b8c081ca","generatedAt":"2026-08-30T16:03:35.327Z","publicationDecision":{"recordId":"urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries","publicEligible":true,"evaluatedAgainst":"maha-epistemic/1.0","reasons":[]},"claims":[{"id":"urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries","scope":"Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.","boundary":"Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness.","claimKind":"theoretical-model","sourceIds":["source-mechanistic-interpretability-sae"],"statement":"The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope.","replication":{"asOfDate":"2026-08-24","assessment":"Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.","independentReplicationCount":null},"uncertainty":{"kind":"qualitative","statement":"No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review."},"evidenceMaturity":"single-study"}],"sources":[{"id":"source-mechanistic-interpretability-sae","url":"https://arxiv.org/abs/2309.08600","title":"Sparse Autoencoders Find Highly Interpretable Features in Language Models","rights":{"note":"The candidate uses original boundary language and a short paraphrase linked to the cited source. No source passage, figure, or table is reproduced.","basis":"citation-with-paraphrase","quotationUsed":false},"authors":["Hoagy Cunningham","Aidan Ewart","Logan Riggs","Robert Huben","Lee Sharkey"],"boundary":"Sparse features and human labels do not establish completeness, unique decomposition, or causal faithfulness.","publisher":"arXiv","establishes":"The paper trains sparse autoencoders on language-model activations and evaluates specified reconstruction, sparsity, and interpretability properties.","identifiers":[{"value":"https://arxiv.org/abs/2309.08600","scheme":"url"}],"publishedAt":"2023-09-15","exactLocator":"Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations."}],"reviewEvents":[{"scope":"source-fidelity","verdict":"approve","reviewId":"epireview_3ade585ee5244a7f9b901acc82541380","rationale":"Sparse autoencoder dictionaries binds claim urn:maha:claim:mechanistic-interpretability-sparse-autoencoder-dictionaries only to Sparse Autoencoders Find Highly Interpretable Features in Language Models at arXiv:2309.08600 abstract and method summary. The audit inspected version-of-record at abstract-only depth and recorded subject and claim support; the claim remains limited to “Limited to Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations. in “Sparse Autoencoders Find Highly Interpretable Features in Language Models”; this candidate records the concept boundary and does not pool results from uncited systems or studies.”. This decision applies only to record urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries at sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f and does not certify truth, external endorsement, independent reproduction, or fitness for use.","reviewedAt":"2026-08-30T16:02:10.856Z","reviewerId":"expert_maha-internal-editorial-scale-v1","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest.","reviewerKind":"internal-editorial","reviewerRole":"AI-assisted record-specific review of inspected source identity, exact locator, bounded claim, uncertainty, non-claims, rights basis, and exact revision. This is not an external subject-matter credential.","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","supersedesReviewId":null,"reviewerProfileVersion":1},{"scope":"domain-fidelity","verdict":"approve","reviewId":"epireview_5952700f64554eef93855b52f13f589b","rationale":"Sparse autoencoder dictionaries remains within domain mechanistic-interpretability. Its mechanism or method is the bounded proposition “The cited source supports treating sparse autoencoder dictionaries as a distinct concept within the stated mechanistic interpretability scope.”; the record does not transfer that proposition beyond Sparse autoencoder dictionaries does not by itself establish system-level performance, safety, manufacturability, scalability, economic advantage, clinical benefit, or deployment readiness. This decision applies only to record urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries at sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f and does not certify truth, external endorsement, independent reproduction, or fitness for use.","reviewedAt":"2026-08-30T16:02:10.928Z","reviewerId":"expert_maha-internal-editorial-scale-v1","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest.","reviewerKind":"internal-editorial","reviewerRole":"AI-assisted record-specific review of inspected source identity, exact locator, bounded claim, uncertainty, non-claims, rights basis, and exact revision. This is not an external subject-matter credential.","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","supersedesReviewId":null,"reviewerProfileVersion":1},{"scope":"boundary-adequacy","verdict":"approve","reviewId":"epireview_02cc07c1a1d746579b5a2fd9700264b4","rationale":"Sparse autoencoder dictionaries retains uncertainty “No cross-source quantitative interval is asserted. Definitions, operating conditions, samples, instruments, and outcome measures must be checked against the exact cited locator during review.” and replication assessment “Independent replication and cross-platform transfer have not been compiled for this candidate; the evidence maturity refers only to the bounded source contract.”. Its non-claims and 2 prohibited inference(s) remain attached to every reuse. This decision applies only to record urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries at sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f and does not certify truth, external endorsement, independent reproduction, or fitness for use.","reviewedAt":"2026-08-30T16:02:10.985Z","reviewerId":"expert_maha-internal-editorial-scale-v1","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest.","reviewerKind":"internal-editorial","reviewerRole":"AI-assisted record-specific review of inspected source identity, exact locator, bounded claim, uncertainty, non-claims, rights basis, and exact revision. This is not an external subject-matter credential.","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","supersedesReviewId":null,"reviewerProfileVersion":1},{"scope":"rights-and-locator","verdict":"approve","reviewId":"epireview_5556acca7d574fefba3d4cd6807cef16","rationale":"Sparse Autoencoders Find Highly Interpretable Features in Language Models is identified by url:https://arxiv.org/abs/2309.08600, inspected at Method, reconstruction and sparsity objectives, experiments, feature analysis, and limitations., and retained under citation-with-paraphrase. This approval binds only exact revision sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f. This decision applies only to record urn:maha:record:mechanistic-interpretability-sparse-autoencoder-dictionaries at sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f and does not certify truth, external endorsement, independent reproduction, or fitness for use.","reviewedAt":"2026-08-30T16:02:11.067Z","reviewerId":"expert_maha-internal-editorial-scale-v1","reviewMethod":"Each criterion is recomputed from the exact record, its inspected alignment audit, source identity, exact locator, rights basis, claim scope, boundary, uncertainty, replication status, prohibited inferences, and revision digest.","reviewerKind":"internal-editorial","reviewerRole":"AI-assisted record-specific review of inspected source identity, exact locator, bounded claim, uncertainty, non-claims, rights basis, and exact revision. This is not an external subject-matter credential.","targetSha256":"sha256:3c1bf26d79f8b36ac42e1c9e9471cfa0746527abf8683d98f7835b00acaee19f","supersedesReviewId":null,"reviewerProfileVersion":1}]}