Pith. sign in

Paper Citation Record · LEDGER

Mixture of Attention Heads: Selecting Attention Heads Per Token

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2210.05144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.05144 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:59:01.344329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:11:06.720935Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5365c168-cc5a-44a8-be24-673fe2225cc8 · inbound

BTS: Harmonizing Specialized Experts into a Generalist LLM cites this paper.

BTS: Harmonizing Specialized Experts into a Generalist LLM Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T21:59:01.344329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:59:01.344329Z digest=sha256:e3cdb7a2537a16c71afde38e0f8d9a3949d9abcebc906716154a9a51fc645535

Observation 86c2d868-d342-44ad-938e-c669d2ec080c · inbound

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs cites this paper.

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T17:01:41.240522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:01:41.240522Z digest=sha256:33af8001256e93b4c0f52f1924068a4766c4abcaba9a7fcbd50e9f937be84c0b

Observation 6d3294ea-96b7-4c0a-87d5-33354b077ca3 · inbound

Action is All You Need: Dual-Flow Generative Ranking Network for Recommendation cites this paper.

Action is All You Need: Dual-Flow Generative Ranking Network for Recommendation Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:16.753361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:16.753361Z digest=sha256:04574c47210857cb22a98affb5141b1bdca796af583296086cf81424b2a8fec5

Observation 6067eade-cc87-4b17-9dd7-cdbed2f449e5 · inbound

Breaking Thought Patterns: A Multi-Dimensional Reasoning Framework for LLMs cites this paper.

Breaking Thought Patterns: A Multi-Dimensional Reasoning Framework for LLMs Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:39:54.315679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:39:54.315679Z digest=sha256:8954e9a207a52573095127c30c008724e88b1402f34e914cfc3942b44da75b55

Observation c0620691-263f-49d0-9227-6f2bacb2b558 · inbound

Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning cites this paper.

Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:38.748604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:38.748604Z digest=sha256:44945833f4354099cc5fb0c64d076e58a70e265c1b1ea31929c185f8d30ce2d3

Observation 89e829d2-9a93-4075-bffb-9877a06fc1a7 · inbound

Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time Training cites this paper.

Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time Training Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:53.389688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:38:53.389688Z digest=sha256:c9c0b67ea05f3ffc1a8d1b7c540dc5189ddbd29ee56b5397ca032119ec1017db

Observation cacf03d1-5f62-4c47-b776-6e1ba64a21bc · inbound

SHMoAReg: Spark Deformable Image Registration via Spatial Heterogeneous Mixture of Experts and Attention Heads cites this paper.

SHMoAReg: Spark Deformable Image Registration via Spatial Heterogeneous Mixture of Experts and Attention Heads Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T15:17:05.032901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:17:05.032901Z digest=sha256:94015b65d0102c31cd87f63ef7dd4b25f8351e015e8857416a43d9fe13db3580

Observation a8445be7-52b0-4da4-8f05-1d5fc7ee50e3 · inbound

Multi-LLM Token Filtering and Routing for Sequential Recommendation cites this paper.

Multi-LLM Token Filtering and Routing for Sequential Recommendation Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:11:06.724810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:05:18.323554Z digest=sha256:d642760ad889936d0522f23e77e938f48e402b509a11eda0f7af5379069662a1

Observation 899edbde-80e4-45f3-8cb6-41f16005532d · inbound

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference cites this paper.

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference Mixture of Attention Heads: Selecting Attention Heads Per Token

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:56.432036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:27:55.991919Z digest=sha256:541264781b00832593e0788bc8241b94965f3bd31900c54131d37567f10f5cc3