Pith. sign in

Paper Citation Record · LEDGER

Multi-matrix Factorization Attention

As of 11 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2412.19255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19255 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:55:04.061512Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:52.656467Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:32:36.395278Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcb67969-c01d-48d7-81b1-b21ff41c4f52 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Multi-matrix Factorization Attention GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.742782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.742782Z digest=sha256:ba0186262610f636c39a4684caa517345ac41331d8ae946d772cc27e6d190765

Observation e76f2303-2042-4701-89b5-4dd16375b2e6 · outbound

This paper cites The Falcon Series of Open Language Models.

Multi-matrix Factorization Attention The Falcon Series of Open Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.749966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.749966Z digest=sha256:6e314c7031244d09a43def7e2807e05f4aba8438a14ce440ecf75a00cdfe3805

Observation 3caed963-bb3b-451a-b038-868ea514c13b · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:55:05.201098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:55:03.756583Z digest=sha256:0e70e1697cb205866655503929f2ac4dae5631b21a7ea419b7dfe1e555eb8cd6

Observation 92b67c6b-c2ad-4f6e-80ab-be3fd0fa0894 · outbound

This paper cites Low-Rank Bottleneck in Multi-head Attention Models.

Multi-matrix Factorization Attention Low-Rank Bottleneck in Multi-head Attention Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:55:04.965410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:55:03.763727Z digest=sha256:3a8f6911ff2a514d6c12d6dfab9754ffa2a6c7757cc8a0282259009ceb84d311

Observation ae81c9a7-50c7-4072-a22b-4d00f3febee7 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.772492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.772492Z digest=sha256:f25b2bea5261bd815f5dfcfc4e2e1a717837a5001d061f34b5fe90cb935bc9e4

Observation e6d12d0e-1a79-4b0e-8f87-f82d9fe68267 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

Multi-matrix Factorization Attention Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.777756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.777756Z digest=sha256:4f3a0c9b9c07914263018e686e1287f32d8f163226c1fa1bbc5ba114f7972bf3

Observation 678b5c3c-087f-4b4a-9bef-24fa8d8c3f0b · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Multi-matrix Factorization Attention BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.784450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.784450Z digest=sha256:d64ccc03fa5115f531719657cd7340bae45aac63100dc04ac9623efb38c0269e

Observation c451885a-bfe1-4cfe-b4a3-0e86e51439ff · outbound

This paper cites Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y.

Multi-matrix Factorization Attention Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.789651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.789651Z digest=sha256:affe89d9d74be25b59aecd943367ae07cc040a2f6207c25c860f29cb95706d68

Observation 73001b93-6c2f-4a70-9f99-d319f3a44414 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Multi-matrix Factorization Attention FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.794878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.794878Z digest=sha256:2b4f5787cfb3b2106b089ffdb2cfed8ec50a9b6bc29582d065c81d4d066cbc74

Observation db60e678-aaca-40c2-8c99-d91446af1d55 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Multi-matrix Factorization Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.802359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.802359Z digest=sha256:9a5bc80e76c803a1000481e4824e9aecfed1853beb554fabd419a876547aba80

Observation 7d058330-4dd1-409d-9251-92e2e7dfc3b0 · outbound

This paper cites Fewer Truncations Improve Language Modeling.

Multi-matrix Factorization Attention Fewer Truncations Improve Language Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.808027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.808027Z digest=sha256:3b71d4c8731e4cb646a3b349520b9e288101fffab2970bacab7c75c20cf7e130

Observation 8f0ae81e-355c-4074-a3a8-7d89bebb67d5 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-matrix Factorization Attention The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.812973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.812973Z digest=sha256:b33f2160f87123ec7cbb9551423b30a2bf1037b524a7bb74362dde93ba95ede7

Observation 27d5b3ed-c61a-459a-8222-f31e5773b378 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.819178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.819178Z digest=sha256:7d6c3eb5429b540b40e27fb0c4b25e0a128a49e78572d265a8224fbcb811c8ac

Observation 16a0bccf-9d73-4f53-abfc-ae3a4467ff13 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Multi-matrix Factorization Attention Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.824319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.824319Z digest=sha256:fb3d64159f3132a4381e31ed7bf3d76d2777b04525d95f71a6918cd976556897

Observation 9e1f3c55-eea6-43c0-ab90-debb06b7fe67 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Multi-matrix Factorization Attention Measuring Massive Multitask Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.828877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.828877Z digest=sha256:fe7723e12d4c9730ba4865ae4e7bf44760d1148618f9d44a525c520708827da1

Observation 515c7370-7777-40ac-8362-1158ad466c85 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Multi-matrix Factorization Attention Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.833918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.833918Z digest=sha256:7437981b4df258bd45cd32515034fe681fd3db7d9ec3b6fda212897dcd660566

Observation 1e29b1d1-8265-487e-b149-d23950b0e44b · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Multi-matrix Factorization Attention RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.839350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.839350Z digest=sha256:6086f0856343ac01d768a00e89f183ddbcf91d5fd65b92b9c442751e7c0ecf9f

Observation 8fcc3c4c-8363-4af5-8151-9b398e34449c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multi-matrix Factorization Attention LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.844509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.844509Z digest=sha256:952115bbafe5af9bb6cd483b5f0f8593111f173a523b5eae1be6dc9be02a8d52

Observation a5f8428a-9483-4257-9547-65091a61dd79 · outbound

This paper cites Mixtral of Experts.

Multi-matrix Factorization Attention Mixtral of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.850264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.850264Z digest=sha256:2400b9d224c3b71ed22f959c502283589a8452a99b5e132b3006354cf6f26df3

Observation e29ba3a1-f6a0-474b-a947-cb21c7a285ae · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Multi-matrix Factorization Attention Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.855619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.855619Z digest=sha256:642e18fb05ad655ccfbfda986567f2805194939f94911cb3c0d454c87615be4a

Observation 26de0459-1aed-4a8c-9acb-590f0067ebeb · outbound

This paper cites Weight decay induces low-rank attention layers.

Multi-matrix Factorization Attention Weight decay induces low-rank attention layers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.863080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.863080Z digest=sha256:8d3c5f8176ecd98e07502196884cb5261d466566b192e68abc05ad8dcf086ef7

Observation 0e627618-f2cf-411c-9900-7e6e3932e4b3 · outbound

This paper cites DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation.

Multi-matrix Factorization Attention DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.868518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.868518Z digest=sha256:e7bda5fc53e2824229ad84823c9675f6ef25f7676b32e45acd079cad4f2acd3f

Observation 2d07eded-e829-4e6b-8ea6-878072c60e98 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Multi-matrix Factorization Attention Jamba: A Hybrid Transformer-Mamba Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.873620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.873620Z digest=sha256:3f67c5a709911566e53989dd4a4eb8f52c32376a3f88b2096bf08b37ab8c7153

Observation 7f855f8d-2a7d-47cc-962b-b53268c255a9 · outbound

This paper cites Decoupled Weight Decay Regularization.

Multi-matrix Factorization Attention Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.878675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.878675Z digest=sha256:f962570d56bfea69e6cf1f07c955bb529bce5c82b7129950ee61b56e219684ad

Observation 2a33144a-e033-44d2-8b79-e10b0888ef76 · outbound

This paper cites Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention.

Multi-matrix Factorization Attention Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:55:04.155091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:55:03.884161Z digest=sha256:4f67858195e515318fccdd0b69814e25d74bc11ae7c107552569e515bf64305d

Observation 3e62234d-b87f-4364-82fe-676556fbf698 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.890189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.890189Z digest=sha256:6d63330eceade3dcfa01a0c98769c563cca5fa06bca83607854ad70d96911f8d

Observation bd718645-5319-4ee1-8220-a2a76148f831 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Multi-matrix Factorization Attention OLMoE: Open Mixture-of-Experts Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.895257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.895257Z digest=sha256:1f924ac6cdefb96245608f3e9f1cbf216c04a0448341d6bdd6fbc7777db077cb

Observation 52a26e0c-be81-4f2f-9c4c-ef851eb31be0 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Multi-matrix Factorization Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.901462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.901462Z digest=sha256:83ddc0899b7f4adb922136c7453a82f3b42068c0fa60b60097dfef72d0a53b63

Observation f9372660-e69a-4365-9c66-3cd35823a340 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

Multi-matrix Factorization Attention Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.907073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.907073Z digest=sha256:9cde54d14d387da3ae40670a90005ace4d210f9dafcbafe9bc97f71b104b35d4

Observation eab54852-8589-4d4c-996e-0ade451bf39a · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.912598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.912598Z digest=sha256:3fc6e49afed93a3a40206177ac20b25a57e0e6a4b78e61a6526c2ac5667dedfa

Observation 2439adca-13bd-43e4-bbc1-ec73fe15d66b · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Multi-matrix Factorization Attention Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.918833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.918833Z digest=sha256:0c298153477abb8a29d857f32628b1188ee391099c22aedf6d7a3c217baa1f38

Observation 5a745ef7-1ac7-4647-81af-de117a6f7d03 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.924736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.924736Z digest=sha256:a01a3d85147780a00eb1781feaa5870026ef0dbf4efbaca795f5731c12b901b5

Observation e845b31c-b6c2-492a-acbb-35a364e79dd2 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.930221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.930221Z digest=sha256:8fdbf25011fef333bfc8d1ededc3ec84abaf991f6211be4e7b5c4e4cb35bdf19

Observation c53b8bad-ebfd-4a94-b39d-9071aada4429 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Multi-matrix Factorization Attention SocialIQA: Commonsense Reasoning about Social Interactions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.944483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.944483Z digest=sha256:7bfd1f3c5ebb5bad60c5c91f05f357f062b6aeddbddb6c6e58b812ed8d5ea25e

Observation 044f62ea-03b4-499f-aa1f-934571b1e6e5 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Multi-matrix Factorization Attention Neural Machine Translation of Rare Words with Subword Units

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.953507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.953507Z digest=sha256:aa59403daeec2c2d69864390cb54345dfb8b67524e3954235c532aeb04fdf0b6

Observation 9dc26d3b-df67-433d-a6db-e221fa417004 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Multi-matrix Factorization Attention Fast Transformer Decoding: One Write-Head is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.962019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.962019Z digest=sha256:d9b760a578703c5725da27505f413e4f10aee48a4c1ce761444d683305f108a2

Observation 46481ceb-e592-43b7-953d-e3d8f80a6bb7 · outbound

This paper cites GLU Variants Improve Transformer.

Multi-matrix Factorization Attention GLU Variants Improve Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.969065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.969065Z digest=sha256:1d3953b34794b0754f80466c9b6c0e277099dd62120a5f292cc9a75986a6af0d

Observation f0cebf4b-e2e1-47b6-830e-0458faf94f32 · outbound

This paper cites Talking-Heads Attention.

Multi-matrix Factorization Attention Talking-Heads Attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.976160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.976160Z digest=sha256:d3020035fd29307b3a253ae34a567f4761df5f39df44b9842df3e340c578c8a4

Observation d599db10-43d2-418a-9ce6-335e5f5feeeb · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Multi-matrix Factorization Attention Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.982953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.982953Z digest=sha256:018d413c56e35d8675a1251da4df4b49bafabcdcfdb9bf9c3509ed439032b57a

Observation d1ad1292-b006-46c8-914a-c6e9b46962a9 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.989969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.989969Z digest=sha256:da243a3780c43badfbf09caa466905129f812b45e85350f34e16c464e5fd2a57

Observation 36bc1ab2-a1db-49a9-a90f-bcb34ab95f26 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Multi-matrix Factorization Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.003033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.003033Z digest=sha256:88b99d46c3b7d12a1a27708fcab6176af999aa07594ca39adcd00ae7e56d377c

Observation ecd3b455-2d7e-4233-af32-645607228eeb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Multi-matrix Factorization Attention Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.009622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.009622Z digest=sha256:c1f40af93401e124b489b93ea202e6c2b41adfbbc26313afc98ac728dc9e7b5d

Observation 3ca8e6cc-aa11-482a-a0b4-51a81a288680 · outbound

This paper cites Attention Is All You Need.

Multi-matrix Factorization Attention Attention Is All You Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.016132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.016132Z digest=sha256:c0835fb8f7763053aae27f41443a5f0adb45f8f19acb246d31fff92af2376a55

Observation 3dbcce2b-91ea-470b-b348-b6e9d2e20cf3 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

Multi-matrix Factorization Attention Crowdsourcing Multiple Choice Science Questions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.021750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.021750Z digest=sha256:010b4495423a6b0e645949b0909a0c4570c04031add00b0d229103e80b4510dc

Observation c02ce927-b42e-47cb-8078-db7de2b8cb3d · outbound

This paper cites Improving Transformers with Dynamically Composable Multi-Head Attention.

Multi-matrix Factorization Attention Improving Transformers with Dynamically Composable Multi-Head Attention

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.027465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.027465Z digest=sha256:c91e47e53e113035f8445fff2c034396d1b1f2d25fb8b83248c127ed3c2d6268

Observation aa6f7a4b-58ba-466b-98c5-7c236e74f2fd · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Multi-matrix Factorization Attention LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.032759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.032759Z digest=sha256:9b75ab67ca4b186e0c3e1432583884fdfe487d1de5d455a617d357771ed6f823

Observation 89770ec0-a618-4696-b185-7fcf6e6be6f2 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Multi-matrix Factorization Attention HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.038913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.038913Z digest=sha256:e5615c9eb3f5c035b642622e82ae44fcb8f12f3d3d27df2eebc39d056a11babc

Observation eccf9207-bffa-43db-88ba-eb1282551107 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.044373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.044373Z digest=sha256:ec53cdda0f00420ddb127ab4795a2c080c4d5beed8a35b4cc10dbb050a1b36e5

Observation f6a75ec7-d879-4c79-ab6d-85c446259519 · outbound

This paper cites MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding.

Multi-matrix Factorization Attention MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.049741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.049741Z digest=sha256:f42267f03219feafa25ccf8878eb8836c50dd9036b3ac7e708ffb352b6476d4b

Observation 311f6198-0f32-4797-bba9-abeef5f0edbf · outbound

This paper cites online" 'onlinestring :=.

Multi-matrix Factorization Attention online" 'onlinestring :=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.055304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.055304Z digest=sha256:8ad334b3309bbe3319d31c20277e94caace66b321d0e44f37bd05b9e197aab48

Observation 647f227c-efdd-401b-adef-4d3c79a9398c · outbound

This paper cites write newline.

Multi-matrix Factorization Attention write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.061512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.061512Z digest=sha256:c8a12efa5eb7ee1201ed88da3f07526c830a7ac061a4aba4e87041dbe7207849

Pith citing papers

Observation 8d6772a9-cfea-4d8e-8201-6ebe13e2042e · inbound

UHD Image Dehazing via anDehazeFormer with Atmospheric-aware KV Cache cites this paper.

UHD Image Dehazing via anDehazeFormer with Atmospheric-aware KV Cache Multi-matrix Factorization Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.656467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.656467Z digest=sha256:fcd3838ffc0644859cf9e0ebe82130edbdf123110dad73f675fa0ca35e969363

Observation 52e12996-15d3-485f-bf64-f646d8edfe21 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Multi-matrix Factorization Attention

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:32:36.488542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T13:32:30.968649Z digest=sha256:b22d5cb3a73ee119cc2bbfb9d8513a068fd2926a1e4a4f4b8602f658e8a0e6c3