Pith. sign in

Paper Citation Record · LEDGER

Monet: Mixture of Monosemantic Experts for Transformers

As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2412.04139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04139 v4

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:49:17.483517Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:55:02.197434Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-09T12:08:23.021613Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48ae1320-e6a8-437e-ac3a-4fed29d17f1b · outbound

This paper cites Nemotron-4 340B Technical Report.

Monet: Mixture of Monosemantic Experts for Transformers Nemotron-4 340B Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.305101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.305101Z digest=sha256:1080aa68cca048b727b6fea0be04eadb513e512c6f788f44802a2e9a3b761e7e

Observation 0b05858b-9258-499a-bb95-2cdf3134d911 · outbound

This paper cites b", nn.initializers.zeros, b_shape) 11 12def __call__(self, x, g1, g2): 13x = nn.relu(self.u(x)) ** 2 14x = jnp.einsum(.

Monet: Mixture of Monosemantic Experts for Transformers b", nn.initializers.zeros, b_shape) 11 12def __call__(self, x, g1, g2): 13x = nn.relu(self.u(x)) ** 2 14x = jnp.einsum(

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.015573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.440051Z digest=sha256:9bee65b9bee2f845e5464858862aa8a7e9d434a2462bec603898fce78a4d3618

Observation f52d3f37-a5a4-4a4a-976d-58e3e238cf44 · outbound

This paper cites Haozhe Chen, Carl V ondrick, and Chengzhi Mao.

Monet: Mixture of Monosemantic Experts for Transformers Haozhe Chen, Carl V ondrick, and Chengzhi Mao

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.319506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.319506Z digest=sha256:ba3e6014cf4d946461ec4395e860c2f46dbed3ef8a8113daf0837a8ea5544833

Observation c3bd332f-a45c-490c-81c8-a25fd29239de · outbound

This paper cites F**kyou!F**k (...)* (16.68%)(...)Snakesonamotherf*ckingplane.

Monet: Mixture of Monosemantic Experts for Transformers F**kyou!F**k (...)* (16.68%)(...)Snakesonamotherf*ckingplane

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:17.904732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.476617Z digest=sha256:d2f827cad8e52b012f49f1a00e65c49bd007e9486044a0d471a01154296515e1

Observation cb4d296f-f8ed-4751-80d6-5f80c23a81c6 · outbound

This paper cites RealToxi- cityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Monet: Mixture of Monosemantic Experts for Transformers RealToxi- cityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.103123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.332827Z digest=sha256:df9c8125ab20b4da488189481f32d7c853fce32a29e1b5aa2557e3583d9e3c85

Observation 50b5e7a9-de52-495d-b6ee-ae6976f0760c · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Monet: Mixture of Monosemantic Experts for Transformers Transformer Feed-Forward Layers Are Key-Value Memories

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.092077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.337104Z digest=sha256:c635047ce38519c25588f64f892cd62aead6d6a7cf0058cada3bf7b654856768

Observation 079ff9d5-2989-4396-b146-5c1b7bf9854e · outbound

This paper cites Mixture of A Million Experts.

Monet: Mixture of Monosemantic Experts for Transformers Mixture of A Million Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.345568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.345568Z digest=sha256:b28ee60301f9e68815a46f25bcdd5180d2409faabd3d25e9ded4c000c7552cad

Observation 7be31826-7cc4-468c-832e-0eaaf53d73fc · outbound

This paper cites An Overview of Catastrophic AI Risks.

Monet: Mixture of Monosemantic Experts for Transformers An Overview of Catastrophic AI Risks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.350721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.350721Z digest=sha256:4d1f39f5ce4fddb66b19a5ffe15ee8253cc4e0cc37df228a839f33d6c6e62863

Observation 282fbb42-89ca-4431-92a4-f850a48bdc8d · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Monet: Mixture of Monosemantic Experts for Transformers AI Alignment: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.355345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.355345Z digest=sha256:1a1d32e5f30bfad8a0e2e1284d10e1bd03d756124e5561b02904c4aca7a75779

Observation a8d396f3-f03a-4a46-8b79-f4fe4f2aae47 · outbound

This paper cites Mixtral of Experts.

Monet: Mixture of Monosemantic Experts for Transformers Mixtral of Experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.359959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.359959Z digest=sha256:6a4203991be96e8572d0d920a2b640f2f7e8829e2cdeb6860d91db269931f751

Observation 5837cd5c-81c5-4297-9768-8331a34d6c41 · outbound

This paper cites The Stack: 3 TB of permissively licensed source code.

Monet: Mixture of Monosemantic Experts for Transformers The Stack: 3 TB of permissively licensed source code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.366448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.366448Z digest=sha256:47936febd76ba348545edec6bb2f637c93473b50c557e3f59159c766e0eae2b5

Observation c909ed0d-cae8-4be6-a2ea-195e35dfac81 · outbound

This paper cites Theory on Mixture-of-Experts in Continual Learning.

Monet: Mixture of Monosemantic Experts for Transformers Theory on Mixture-of-Experts in Continual Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.372648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.372648Z digest=sha256:beec0bf46716bb75a315803c7f0393032b2880ffc89320b4c3d0a8b11948979e

Observation 09b0c3ed-ca01-4bd2-b729-05e8f79176b7 · outbound

This paper cites StarCoder: may the source be with you!.

Monet: Mixture of Monosemantic Experts for Transformers StarCoder: may the source be with you!

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.377588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.377588Z digest=sha256:ad8046ce03c1a8ab7eb5ec3c6b286c67fe2329fc75de62566d6f13e6902cccef

Observation a5f7b35d-b5d2-4d89-b75d-2b9c9eed1692 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Monet: Mixture of Monosemantic Experts for Transformers Scaling Laws for Fine-Grained Mixture of Experts

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.080870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.381840Z digest=sha256:72c10530a8f31883253292316d1eb8e9993796059cc6e0aad65537837c276c97

Observation 27bd9235-fb72-459e-ac69-e41d5c8db21b · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Monet: Mixture of Monosemantic Experts for Transformers Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.386441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.386441Z digest=sha256:6da5864fb5c0417e7f6f3d349691f2698c3ae875598cb542eca6c8f0ed5ad505

Observation ec5e0d1f-8b6d-4135-a950-e7c1c5844321 · outbound

This paper cites Jesse Mu and Jacob Andreas.

Monet: Mixture of Monosemantic Experts for Transformers Jesse Mu and Jacob Andreas

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.066925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.392160Z digest=sha256:9e0b4a5269a934c2f7dcac2b67cfb11b14bf3842ab2bdbe3f516e530929d0c6e

Observation e650a4d8-dfe9-47c6-b120-9261da758d87 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Monet: Mixture of Monosemantic Experts for Transformers OLMoE: Open Mixture-of-Experts Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.396208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.396208Z digest=sha256:f9d4ef71627c4ff22e15415f47fe5e1bfc36b74aa42fb6c9e285db57e6c45ea0

Observation 17cd2fc4-25f7-46dc-b552-0baf645a894d · outbound

This paper cites Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models.

Monet: Mixture of Monosemantic Experts for Transformers Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.400132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.400132Z digest=sha256:ae60d5ecf4f526d823b0146d1da81b17326ea21f1d029eccaecaeee4285d574e

Observation aefa3d9c-84a0-4219-83bf-2b3f5d689843 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Monet: Mixture of Monosemantic Experts for Transformers The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.404980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.404980Z digest=sha256:88191f9d018b836bf5c5b854667e446e16fa3f64f9a4329b2c46ef1ff0387b96

Observation fa7fe330-0eaf-4015-bae2-c6511e5a16a9 · outbound

This paper cites Taking features out of superposition with sparse autoencoders.

Monet: Mixture of Monosemantic Experts for Transformers Taking features out of superposition with sparse autoencoders

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.053942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.410568Z digest=sha256:a55eabd12d8b5b13ffacb1c2e5962f42819b54fd4d5f0ae5682d7359476dd3b9

Observation 8cb3b353-fb27-4997-a56f-98bdb79ee9c0 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

Monet: Mixture of Monosemantic Experts for Transformers JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.414775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.414775Z digest=sha256:b69cf20da0aedceeedb11c4a9c2791728ac9398f3c29e44d545440f42b823471

Observation 4dc2d541-93b0-4641-9cde-bbe132a9a1b0 · outbound

This paper cites Codebook Features: Sparse and Discrete Interpretability for Neural Networks.

Monet: Mixture of Monosemantic Experts for Transformers Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.418586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.418586Z digest=sha256:27fc49c52ac6ae401ddbf79dd97b7d3c3764c7a8622299aab049587b3d992ce7

Observation 2e7dff76-7ce8-4d30-b7f3-bfae6a2551ac · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Monet: Mixture of Monosemantic Experts for Transformers Gemma 2: Improving Open Language Models at a Practical Size

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.422409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.422409Z digest=sha256:6b30e5cf4a24dff5c1d664218af2da430d5fd442ec8f668dbf71d2ff9659ce02

Observation 40c0a478-58f3-4a12-b0dd-135ab96bcabe · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

Monet: Mixture of Monosemantic Experts for Transformers ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.426665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.426665Z digest=sha256:d3e1e0fa2f2f940e5f3e3ce38bb2bb3a8abcb1487125c55e038ffae82e07289d

Observation 88ddfaba-7cce-45f2-ab6d-bcace9afc643 · outbound

This paper cites CONTENTS A Method Descriptions 18 A.1 Expansion of Vertical Decomposition.

Monet: Mixture of Monosemantic Experts for Transformers CONTENTS A Method Descriptions 18 A.1 Expansion of Vertical Decomposition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.042704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.431778Z digest=sha256:15756a1c2e46937e926fda7b5bc28f6bee89cce6a93e983a6ad1ae144f323644

Observation 5787dd54-254f-43b2-bf12-6405cace9c38 · outbound

This paper cites Moreover, the multi-head expert routing probabilities are consoli- dated into single routing coefficients PH h=1 ˆg1 hi and PH h=1 ˆg2 hj, reducing redundant aggregations.

Monet: Mixture of Monosemantic Experts for Transformers Moreover, the multi-head expert routing probabilities are consoli- dated into single routing coefficients PH h=1 ˆg1 hi and PH h=1 ˆg2 hj, reducing redundant aggregations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.030720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.436382Z digest=sha256:30601ccc1d003a0f303bc3180ae666da9442ea557ed484856c1289c086b9562c

Observation b521dc35-7029-4b51-819d-64fea07fada6 · outbound

This paper cites an unresolved cited work.

Monet: Mixture of Monosemantic Experts for Transformers Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:49:18.001376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.444083Z digest=sha256:0277949bf4754c45a1551a0ff751a3853bcb9cf4cb1b46970a7ae3011df0937d

Observation 205d08a4-c02a-4e40-9b3f-dce2b2e89ee2 · outbound

This paper cites b1", nn.initializers.zeros, b_shape) 16self.b2 = self.param(.

Monet: Mixture of Monosemantic Experts for Transformers b1", nn.initializers.zeros, b_shape) 16self.b2 = self.param(

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:17.987633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.447699Z digest=sha256:7edd6bb58d429e8b5f3bd85ae49296c2f2ab70e44780d21f8213b17fcea73e4d

Observation 2155c7b8-ccad-4b50-88cc-fd5497bd2339 · outbound

This paper cites To manage computational resources effectively, we adopt a group routing strategy wherein the routing probabilities are reused every 4 layers.

Monet: Mixture of Monosemantic Experts for Transformers To manage computational resources effectively, we adopt a group routing strategy wherein the routing probabilities are reused every 4 layers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:17.972085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.451221Z digest=sha256:9782bb7d399c55344d9ab112d855e79c5636c146255eea42fcd83a97a25a1e10

Observation 3b6d7aa3-ec52-4887-8481-d7afdbb26e34 · outbound

This paper cites an unresolved cited work.

Monet: Mixture of Monosemantic Experts for Transformers Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:49:17.958821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.454643Z digest=sha256:b0dbf78db1547daac994c407cf14adea55f0907c6226beaa2e57e9975a91fcef

Observation d3aad49d-4168-44ca-8d6c-53e51176b194 · outbound

This paper cites The ta- ble reports the number of experts assigned to each programming language across all routing groups.

Monet: Mixture of Monosemantic Experts for Transformers The ta- ble reports the number of experts assigned to each programming language across all routing groups

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:17.945417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.459871Z digest=sha256:6334adcfca5d415eb9bf4931972609fb6274b0f3dec7ec081c41032cb7742c5f

Observation f0ad5e76-b44c-44e2-acc1-27ddb37e82c7 · outbound

This paper cites "" 12#!/usr/bin/env bash 13 14echo.

Monet: Mixture of Monosemantic Experts for Transformers "" 12#!/usr/bin/env bash 13 14echo

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:17.931920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.466771Z digest=sha256:8d1f19f2d0bd2137b3fa5750be3725519b2c7a7c7cd47a0049dee1df58f02401

Observation 248b5d53-cb14-42fb-a7cf-2921df6104fa · outbound

This paper cites Based on this, we masked experts associated with each language and re-evaluated the code generation benchmark to estimate the model’s capa- bility to unlearn programming languages.

Monet: Mixture of Monosemantic Experts for Transformers Based on this, we masked experts associated with each language and re-evaluated the code generation benchmark to estimate the model’s capa- bility to unlearn programming languages

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:17.918495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.471020Z digest=sha256:398460d3fc03a04fe8a65631a9515654f41f9a5b1b7b888a987c6c00a2d7708b

Observation ff23ca77-efa7-4ac1-896b-d436f02b129c · outbound

This paper cites JULYIV (...)rew (59.50%)(...)TheembroideryreadsinHebrew:.

Monet: Mixture of Monosemantic Experts for Transformers JULYIV (...)rew (59.50%)(...)TheembroideryreadsinHebrew:

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:17.890678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.483517Z digest=sha256:bf2ed5774e864abf9314d54ccd7d03e95b7191d97fc5e2e8924a2bdcdac4719e

Observation 5dfbb1be-ff90-4064-873d-92e3824017e8 · outbound

This paper cites BatchTopK: A Simple Improvement for TopK- SAEs.AI Alignment F orum,.

Monet: Mixture of Monosemantic Experts for Transformers BatchTopK: A Simple Improvement for TopK- SAEs.AI Alignment F orum,

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.126943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.314729Z digest=sha256:7f06751d7bbb9f38c9b8f7dbd92bcb20418207f4900b5aa9fb30713720954a99

Observation 3be18078-dc49-4ee3-8d8c-5a5a606dd494 · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

Monet: Mixture of Monosemantic Experts for Transformers Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.340410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.340410Z digest=sha256:7c44325d1386c44ed2af30ff5acd57821fd64fb59d301046cb0b6871bd5faf7d

Observation 38440118-84ef-488e-9c34-51d3bc740853 · outbound

This paper cites A Review of Sparse Expert Models in Deep Learning.

Monet: Mixture of Monosemantic Experts for Transformers A Review of Sparse Expert Models in Deep Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.328429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.328429Z digest=sha256:81b8c18d6b48542074f2d781a0ec112e8e4723825ff9566b2a05c021c2678522

Observation f391b43e-99f7-4df1-b0be-83ca33bcbf8c · outbound

This paper cites an unresolved cited work.

Monet: Mixture of Monosemantic Experts for Transformers Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:49:18.140201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.310182Z digest=sha256:d4985225d333f5d1a418a64ceeddec4b7bf723f1f7f763846b3ed1ab4b5cb816

Observation a2481438-38c2-41c0-b057-1993764de5aa · outbound

This paper cites Mem- ory Augmented Language Models through Mixture of Word Experts.

Monet: Mixture of Monosemantic Experts for Transformers Mem- ory Augmented Language Models through Mixture of Word Experts

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:49:18.114743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:49:17.323690Z digest=sha256:c7183b91725b28845afd3c613a2d4e8430a7b1254851af9dab8f91ccac6f6de7

Pith citing papers

Observation e3e18910-6f01-4337-ad66-3b0d06010f01 · inbound

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition cites this paper.

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition Monet: Mixture of Monosemantic Experts for Transformers

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T14:55:02.197434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:55:02.197434Z digest=sha256:2f3af98b95e29110d4bb0f3b6da8fa4ec09e012abcafba9a7dbfb956da836b20

Observation a9802ffa-0cfa-4a99-b7e6-efa79ea705d1 · inbound

Studying Cross-cluster Modularity in Neural Networks cites this paper.

Studying Cross-cluster Modularity in Neural Networks Monet: Mixture of Monosemantic Experts for Transformers

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-09T12:08:23.058043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T12:08:22.805965Z digest=sha256:52a1b829288376e4586b4120f0f5153f5cebe71f6f27d4847a8105214babb370