Pith. sign in

Paper Citation Record · LEDGER

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2412.05540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05540 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:21.642571Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:43:25.901211Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T15:51:38.997059Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc710e34-59b4-4380-ab50-e3dab07a61f4 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.490482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.490482Z digest=sha256:895e7e9ec223d8fc5cf9fafd67eac6d7b6503bd333029ceb0c5b9a5592ad415c

Observation e4b5cca8-1401-4f7e-b55f-14dfc5164697 · outbound

This paper cites Zero-shot text-to-image generation,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Zero-shot text-to-image generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.264997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.497115Z digest=sha256:8c35906201974848b3131eae5cb1893d5b0b2a1fedfe5ecca0c980c0cc41bb00

Observation a07b0ed2-e6b8-4132-8281-57948c57aafd · outbound

This paper cites Mistral 7B.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Mistral 7B

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.502618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.502618Z digest=sha256:899e101871492cb8a8f93e52320f613fc511d4e80031f244849277e87f1dd723

Observation 01231065-4c37-4995-a696-df65663d7198 · outbound

This paper cites Unified scaling laws for routed language models,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Unified scaling laws for routed language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.250930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.507634Z digest=sha256:d94915c3d2cfca62d3d284257ed95140270eacd006a3ba9ccd2988156b706d4d

Observation 630c04d3-6f97-4d36-8d79-ef099dfdd46c · outbound

This paper cites Scaling vision with sparse mixture of experts,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Scaling vision with sparse mixture of experts,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.519139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.519139Z digest=sha256:9bde6819aa6b791be3a9497b2a49be31f0db9e47c12afea841dd8c4d443ac7d1

Observation b0469497-c796-44f9-9bbb-8f807b467062 · outbound

This paper cites M3vit: Mixture-of-experts vision transformer for efficient multi- task learning with model-accelerator co-design,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers M3vit: Mixture-of-experts vision transformer for efficient multi- task learning with model-accelerator co-design,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.226290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.526321Z digest=sha256:7ba856579a20883d25b3fea5c5133a5bed8a503b09b0201293a92fdaa08db347

Observation 1dd75d63-800d-4c3b-a0e0-d562a164bcb6 · outbound

This paper cites Networks of spiking neurons: The third generation of neural network models,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Networks of spiking neurons: The third generation of neural network models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.207769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.534114Z digest=sha256:7eda5b76f65dee3a71fa7df961441c429930d9e70127269e18c339cfd0819fbb

Observation d7eabd5c-0990-4a42-af51-2e010b8588c7 · outbound

This paper cites Towards spike-based machine intel- ligence with neuromorphic computing,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Towards spike-based machine intel- ligence with neuromorphic computing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.192790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.540861Z digest=sha256:74557c6009d412d5e500d9b91d6d677602b20403e518c33216b281f3ad1cbc53

Observation abc41cbd-dc55-4c6e-bd83-07b1e49cbee6 · outbound

This paper cites Truenorth: Design and tool flow of a 65 mw 1 million neuron programmable neurosynaptic chip,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Truenorth: Design and tool flow of a 65 mw 1 million neuron programmable neurosynaptic chip,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.179121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.546070Z digest=sha256:e4247ce04e5eb39d7997d72eca71bdd234a41b5821a97e67854a6b9288d97a5c

Observation bcc1ecce-cd5d-4e50-aed4-85d9e761055c · outbound

This paper cites Loihi: A neuromorphic manycore processor with on-chip learning,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Loihi: A neuromorphic manycore processor with on-chip learning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.552958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.552958Z digest=sha256:f97f334b28e00a7b4ed05db9a348ed555c05663528189adf9721588de6f17f76

Observation 4f0bb142-1c41-4a8b-829e-7f5da7b0fb41 · outbound

This paper cites Spikformer: When spiking neural network meets transformer,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spikformer: When spiking neural network meets transformer,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.137091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.562208Z digest=sha256:7306af74830eb8719a74824125eb6c336ae9e7982f5c1d6cd1cf2b19493b1d7a

Observation d3faa5c3-17da-4d16-a6eb-cf0ed2d29e29 · outbound

This paper cites Spiking transformers for event-based single object tracking,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spiking transformers for event-based single object tracking,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.113630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.566987Z digest=sha256:71afce4ecfeca515fe03ec7c73c504275a66d60621348c33e7d66f0c1bf2ad6b

Observation c2a553b6-e50d-42ab-a6b9-3769f2de6c90 · outbound

This paper cites Spikegpt: Generative pre-trained language model with spiking neural networks,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spikegpt: Generative pre-trained language model with spiking neural networks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.093632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.571540Z digest=sha256:9839f1733b9924658561f83cf90cea3de63aaff6e4a7b5ae22079cb3c12e2b82

Observation 77e862d1-3657-45d9-b92b-356ddae6465d · outbound

This paper cites Spike-driven Transformer.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spike-driven Transformer

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:41:21.933526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.576255Z digest=sha256:0a9ea1a905084b34803f232cb0f951afb6075c0c03405552108ff8814a3df369

Observation c7930d23-617e-46df-8b11-b783b97b05eb · outbound

This paper cites Reconfigurable dataflow optimization for spatiotem- poral spiking neural computation on systolic array accelerators,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Reconfigurable dataflow optimization for spatiotem- poral spiking neural computation on systolic array accelerators,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.073022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.582056Z digest=sha256:dfdee5d5be460883c049625c497f08941c59faed32ef85d3bc30d00373f45470

Observation 734d311b-43f5-47c4-97cb-0b1646618ea7 · outbound

This paper cites Parallel time batching: Systolic- array acceleration of sparse spiking neural computation,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Parallel time batching: Systolic- array acceleration of sparse spiking neural computation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.153553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.587299Z digest=sha256:f933e3e9e6ccd064d6059daf9cc2a2ffd598eecd934dfdeb57a31d452b5fe733

Observation f117f8ed-ade5-48d3-86c0-bc289fc712bc · outbound

This paper cites LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:41:21.910319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.592496Z digest=sha256:d2bfa8803e157ce0ea99b2fe9f919bb582f57bf2e7c1f8a33ecfdf6776955f5a

Observation 0496b4a5-050a-4289-a350-2dc3a3a3c674 · outbound

This paper cites Spinalflow: An architecture and dataflow tailored for spiking neural networks,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spinalflow: An architecture and dataflow tailored for spiking neural networks,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.598783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.598783Z digest=sha256:f427d46dd1e6e32c39b5ae1244021347fc5ec70f7f0159135227eab37b9f52c9

Observation 35c969a9-2f04-4a0a-a1b3-a4f7b879adf0 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.604890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.604890Z digest=sha256:a9867f11d5ca2110ae380ae6175faddda556930860d73193c0524bcd40ae5306

Observation df4b7f45-241f-45f0-bb8d-3950f1efd3ad · outbound

This paper cites 3d-carbon: An analytical carbon modeling tool for 3d and 2.5 d integrated circuits,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers 3d-carbon: An analytical carbon modeling tool for 3d and 2.5 d integrated circuits,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.034996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.611982Z digest=sha256:4f10984eded3ee69ce0a6d55d72955e7bcc5d75b6ceca006e700451bd380a1ad

Observation 7031a519-f29e-4358-94ba-36e100f94e80 · outbound

This paper cites Nebula: A neuromorphic spin-based ultra-low power architecture for snns and anns,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Nebula: A neuromorphic spin-based ultra-low power architecture for snns and anns,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.014632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.616423Z digest=sha256:0d1516582f45dccd8cfe8351ad060a7c3caecc533ab4b315fa08dbc9de00e87e

Observation 13cef4d3-103b-4b89-a464-67624381c2e9 · outbound

This paper cites 30.2 a 22nm 0.26 nw/synapse spike-driven spiking neural network processing unit using time-step-first dataflow and sparsity-adaptive in-memory computing,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers 30.2 a 22nm 0.26 nw/synapse spike-driven spiking neural network processing unit using time-step-first dataflow and sparsity-adaptive in-memory computing,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:21.995628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.621426Z digest=sha256:b8688b556bd0da794bfa50979c9c6ae5371391a35cd089e3afeb7db94869b089

Observation 55057ff2-f72a-4ab2-99c3-d6b7bc894bb0 · outbound

This paper cites Design and architectural co-optimization of monolithic 3d liquid state machine-based neuromorphic processor,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Design and architectural co-optimization of monolithic 3d liquid state machine-based neuromorphic processor,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.626932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.626932Z digest=sha256:9abe4dbc7dc2826a425b171f44187b6c7897adc6bb113f9cfded6ff7770326ad

Observation 20363ebb-9580-46ff-af03-082a1fca549a · outbound

This paper cites Area-efficient and low-power face-to-face-bonded 3d liquid state machine design,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Area-efficient and low-power face-to-face-bonded 3d liquid state machine design,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:21.979842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.631366Z digest=sha256:b803b8e4b2b2e57b50b74710a9176d13681b36d5577ce09135c46d55d6ee5fb4

Observation e1c9ebdf-6c5c-46f9-a8b0-d40daf0a8dc5 · outbound

This paper cites The cifar-10 dataset,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers The cifar-10 dataset,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:21.965393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T20:41:21.636773Z digest=sha256:2218fe09d0fea800c08dbaeb3d63f8ddf802e2fda90ba03e57ad0b9c1e7a47cf

Observation c91688db-3921-4659-9f99-8ec3f1d6f87c · outbound

This paper cites Pin-3d: a physical synthesis and post-layout optimization flow for heterogeneous monolithic 3d ics,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Pin-3d: a physical synthesis and post-layout optimization flow for heterogeneous monolithic 3d ics,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.642571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.642571Z digest=sha256:d6031a2988955d80571307f4230bc5d90c34438ca843dee8a5640c6c56ac1180

Pith citing papers

Observation 8aa85490-e6c9-4ede-adb9-6bb8a5c07066 · inbound

SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks cites this paper.

SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:25.901211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:25.901211Z digest=sha256:79122875b3cfcc27927935540164907fb5ad085594785e3deba2039982e17d20

Observation 1a9c8c84-c348-4d5d-8571-33d990d5e2c5 · inbound

Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts cites this paper.

Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:39.113913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T19:10:18.888463Z digest=sha256:377f9c627cc47f2475b88740070a50ecfc699c20ef13c0fda1658260bfca012e