Pith. sign in

Paper Citation Record · LEDGER

GTA: Grouped-head latenT Attention

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2506.17286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17286 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:53:51.139013Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 00dbff2e-f327-4928-9f0b-8d5c481e0104 · outbound

This paper cites Language models are few-shot learners.

GTA: Grouped-head latenT Attention Language models are few-shot learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.390682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.390682Z digest=sha256:b420d5a59c7627c63a90fc9cba3dbdaed27f2dd98fd9e952b558cd5cdb879e62

Observation 4ebbec90-b89c-41fd-be84-6e272170501a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GTA: Grouped-head latenT Attention LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.501981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.501981Z digest=sha256:382d1052652b918bbd239b7a1aa7d37923acb23970769523fbfa839d61c23007

Observation fc29431e-b635-49aa-bcbb-a996ea5d8ca4 · outbound

This paper cites Attention is all you need.

GTA: Grouped-head latenT Attention Attention is all you need

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.647788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.647788Z digest=sha256:598a236030c25940aef2204405fc25cd970dd5cd10799fbb095310f128231669

Observation 920eb91f-e7c4-4663-aae1-85b1eb9c38cf · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

GTA: Grouped-head latenT Attention Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.764009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.764009Z digest=sha256:399ad0d1078ae89f5cb6323ab3994d454892edf88bb7518d87fb44834ce6cc99

Observation 4ea556eb-1159-4dc4-8ae5-5fb29e61424d · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

GTA: Grouped-head latenT Attention Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.833695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.833695Z digest=sha256:3753d8fc96f3fb1d6b8ef24ab5cf2797d30616016da57efb2ed98c13cf9e0d9a

Observation a8326140-3e94-4c77-9983-a9226041878d · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

GTA: Grouped-head latenT Attention Fast Transformer Decoding: One Write-Head is All You Need

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.967513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.967513Z digest=sha256:d51f5aea52aa7693070accb42b37ad9289e9849c0950d9206ef65dd17e2843e8

Observation b95852db-d43f-4473-af45-d4722c829c85 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

GTA: Grouped-head latenT Attention GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.108508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.108508Z digest=sha256:a2d784cb997bdd13696ffe38364631e0e358d22c7864831d780ec18fede7053b

Observation 307edecf-d678-4982-b0c5-4dbd80e6415e · outbound

This paper cites DeepSeek-V3 Technical Report.

GTA: Grouped-head latenT Attention DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.226361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.226361Z digest=sha256:37959b05f1e196b3728efec682df2060373fb86fde049c06fafec1e304067944

Observation fe511906-48f7-45b3-a422-053b641aef66 · outbound

This paper cites Differential transformer, 2025.

GTA: Grouped-head latenT Attention Differential transformer, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:53.096465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:48.355571Z digest=sha256:39a174e7e552b9f6540adaf1d0accb2a344f8e02afbb36d8402853057ff780a3

Observation b51b9367-075d-4340-8c3f-60a558981f32 · outbound

This paper cites Multi-token attention, 2025.

GTA: Grouped-head latenT Attention Multi-token attention, 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.972928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:48.456181Z digest=sha256:750c6303659e3255ad1c2937fd20480868b5896a5e35d80fb3fc6a76a0a77652

Observation d2968eed-4a2f-4e4d-9508-2ed67aa706ef · outbound

This paper cites Glu variants improve transformer, 2020.

GTA: Grouped-head latenT Attention Glu variants improve transformer, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.518843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.518843Z digest=sha256:185991dfc6a78396ffc938a1f8d9bb99303598e277734ec142f1f9c34e90dff5

Observation c722fbdb-1af5-4fad-bb5c-1a813315cc2e · outbound

This paper cites You only cache once: Decoder-decoder architectures for language models.

GTA: Grouped-head latenT Attention You only cache once: Decoder-decoder architectures for language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.849712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:48.572133Z digest=sha256:34a3d3da5605c448b0d5e3e1c3b50f07a595417cf788bc954c80e3a9c39a2fce

Observation 58ba30cb-17a5-4f4a-9488-88b050c1b950 · outbound

This paper cites Ni, Haifeng Zhang, and Jun Wang.

GTA: Grouped-head latenT Attention Ni, Haifeng Zhang, and Jun Wang

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.684881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:48.674845Z digest=sha256:34e6bc64d1a976a17e3614f9342cc97c2da94c3412aa3623191f242673f109d0

Observation 08ad2c8a-ae4f-4aa7-8d96-49f80ec64a0e · outbound

This paper cites GLU Variants Improve Transformer.

GTA: Grouped-head latenT Attention GLU Variants Improve Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.777635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.777635Z digest=sha256:e6c30be220b1b49810d378cab0b5311030d2e45a543a05ea2fdd48fcdae5faac

Observation 43b62c68-2d45-4000-a3f3-492cd34c5dd0 · outbound

This paper cites Gated linear attention transformers with hardware-efficient training, 2024.

GTA: Grouped-head latenT Attention Gated linear attention transformers with hardware-efficient training, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.535108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:48.880524Z digest=sha256:bae48a1ca5cbf697c93c75652ae4bdcecd3970163cbfd7ba129ae909ca214773

Observation 06bd6350-5d75-4be8-b3ab-1927f5051590 · outbound

This paper cites Hardware-efficient attention for fast decoding, 2025.

GTA: Grouped-head latenT Attention Hardware-efficient attention for fast decoding, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.367762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:48.956459Z digest=sha256:a5cc5a6ce02270450efe6b0c82e7c871fb00220d27c0ee32a949d3d2f0eb8ee8

Observation b0d713df-c335-4064-8255-7beb72fc8881 · outbound

This paper cites an unresolved cited work.

GTA: Grouped-head latenT Attention Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.088623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.088623Z digest=sha256:2f31035e8ad0863d5136bf9ba3828fb6503e81d66786b2fca660bc7e09f5d774

Observation a9eaddfe-0c3a-46f5-8b68-5036f0ad77db · outbound

This paper cites Decoupled Weight Decay Regularization.

GTA: Grouped-head latenT Attention Decoupled Weight Decay Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.214988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.214988Z digest=sha256:3baeeae2102cefc145ecc74792e5c483e919e052a8ca4a80080d14b6ecdfb037

Observation 56d20d05-d48d-4d11-b3e2-25b021c85938 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

GTA: Grouped-head latenT Attention TinyLlama: An Open-Source Small Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.285407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.285407Z digest=sha256:58e0d94e3c3220a6e3ee109273ef2ec609684a874c38e7b187fa7bc896c1548c

Observation dc9c92cb-2732-498e-8ab7-17540c8354b5 · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017.

GTA: Grouped-head latenT Attention Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.369240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.369240Z digest=sha256:6a524b81d0ac95b5983cc0d59b90536a748e9dc63095e2ba0a790a5a609ce61b

Observation 77e9d984-8634-4cb3-8730-055289ba8ece · outbound

This paper cites Relu 2 wins: Discovering efficient activation functions for sparse llms, 2024.

GTA: Grouped-head latenT Attention Relu 2 wins: Discovering efficient activation functions for sparse llms, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.194831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:49.498212Z digest=sha256:1f9c6facf409a006c7193a3f05f955d5a9a5b033de19fcf28e82a667d5c14f35

Observation f4d3644c-b984-4d45-8a23-4bd89a89d1d9 · outbound

This paper cites Smollm-corpus, 2024.

GTA: Grouped-head latenT Attention Smollm-corpus, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.632203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.632203Z digest=sha256:a3db0e0ea3bdd8dd44aba50f2c90ecd6ab5d4bb798b2e47e9e08fef70b3c6fe1

Observation 5c8ba0fb-51db-479c-8c4b-9ebddd7450a5 · outbound

This paper cites The llama 3 herd of models, 2024.

GTA: Grouped-head latenT Attention The llama 3 herd of models, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.063682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:49.697068Z digest=sha256:fadf291a34470f6a93cf4061dbf05e033cc494ae95263f863a7e5f75b7a748b7

Observation c2ef41e4-dfdb-4af5-9f77-e9cbe4ad480e · outbound

This paper cites Mobilellm: Optimizing sub-billion parameter language models for on-device use cases, 2024.

GTA: Grouped-head latenT Attention Mobilellm: Optimizing sub-billion parameter language models for on-device use cases, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.814149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.814149Z digest=sha256:5b2c80acbdf9f346dcdbd169dc3b10b1653fc925a9af8f41936d86012442ee5a

Observation 3cf808b0-f001-4041-9900-5ca672e55c3d · outbound

This paper cites The language model evaluation harness, 07 2024.

GTA: Grouped-head latenT Attention The language model evaluation harness, 07 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.960808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.960808Z digest=sha256:82fb94ce7a1937ea00268bdba0868c30d8cf5ec6e3fb20523e1e864b788b57b4

Observation b0ea5114-81a4-4763-823d-df4477e0d4c0 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

GTA: Grouped-head latenT Attention Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.994896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.994896Z digest=sha256:85b8d0ca448b258305ebf4d1dabf9e4a309691aef4c061f9cebaf2a32f16b64e

Observation 49bdd69e-18ce-4294-a2e0-0dba32490221 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

GTA: Grouped-head latenT Attention Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.081508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.081508Z digest=sha256:3df61e87636805655bf2b8350319cfdc695621230f7237491dbf3efc5982a1bc

Observation 0b0fd502-af34-455e-b466-23472381942f · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

GTA: Grouped-head latenT Attention BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.161101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.161101Z digest=sha256:9e1af48e3abd4d7608afc04de663f0faa552315dab9f54008439ae4408be8dd5

Observation 87aafeb8-9661-4c83-8d04-119cb1a6d60c · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

GTA: Grouped-head latenT Attention Piqa: Reasoning about physical commonsense in natural language

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.214470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.214470Z digest=sha256:a6a23ce344d16d12edd192a8fae650ca6c9d959435cd60f31d23a57b11b07247

Observation 07b4956b-e65e-4d1e-af13-605dc5505eb8 · outbound

This paper cites MathQA: Towards interpretable math word problem solving with operation-based formalisms.

GTA: Grouped-head latenT Attention MathQA: Towards interpretable math word problem solving with operation-based formalisms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.940208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:50.285598Z digest=sha256:f2811cd3570ecd014a275dd259f6a355d58a79f44039830d1fee35ba60529e51

Observation e93a7f2a-d47a-4b22-95f9-99829db64080 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

GTA: Grouped-head latenT Attention TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.363927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.363927Z digest=sha256:d30a477c4f99dd3a0e20e46017352766c23384b981e9573aa789c2681f3325b5

Observation 18f11fec-a649-47e8-ba35-77e51e755d9a · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

GTA: Grouped-head latenT Attention SocialIQA: Commonsense Reasoning about Social Interactions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.475025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.475025Z digest=sha256:dd5140de818f321c90c0215c9bb52235f4d88b0cc0db828657f1746e4512efed

Observation 7ab3499b-47f7-44f7-8964-9f5ad47448e7 · outbound

This paper cites Program Synthesis with Large Language Models.

GTA: Grouped-head latenT Attention Program Synthesis with Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.540047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.540047Z digest=sha256:9f0273af39ee38d28963d4315fe4e633383607306882f38c907f3005c6324cbd

Observation 8740e4f0-43d0-43c6-8de4-6ea2b9f9383d · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

GTA: Grouped-head latenT Attention Instruction-Following Evaluation for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.616735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.616735Z digest=sha256:3f40c5747eeac708a8a5ed951fb13be447346a7ba1430cbd4ea87419493b4400

Observation a82b3e84-a2a8-424b-a46b-a42348b96bd7 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

GTA: Grouped-head latenT Attention LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.701663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.701663Z digest=sha256:570fa3d72f96eda9ce3f67f8a30c0964376b07dc8b3256b5363dd9b3622d63c0

Observation 33c5ad1b-d64a-455b-8b27-754e019856c8 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

GTA: Grouped-head latenT Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.764532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.764532Z digest=sha256:536e9868f81c01ebbb5abb0d975c2ad37f116225aa7a418c3f51a668ab35e3fb

Observation d7d3f5f3-3534-46b5-aca8-647aaa316281 · outbound

This paper cites Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D.

GTA: Grouped-head latenT Attention Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.801241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:50.841526Z digest=sha256:68fee68a37b2223abe597a174f087c765b82562c860552bbd629a0a0091f213e

Observation e837571f-512b-4c22-8492-1a7c3854961f · outbound

This paper cites Llm inference unveiled: Survey and roofline model insights, 2024.

GTA: Grouped-head latenT Attention Llm inference unveiled: Survey and roofline model insights, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.694951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:50.920455Z digest=sha256:48e17990a5c6b69e5654705287fcc527107cad8b5476907e51bbab8d7c00c98b

Observation 9aa579f8-cbfa-4cb2-9c23-1a82b430a474 · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

GTA: Grouped-head latenT Attention Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:51.000364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:51.000364Z digest=sha256:5bf117e338a71a2091095ac473f9a58980b3c5752b4547e4ebe28379b1c2318c

Observation 4b2d1f06-2342-48e4-9fe5-3cc590977f97 · outbound

This paper cites Validation.

GTA: Grouped-head latenT Attention Validation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.571847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:51.079470Z digest=sha256:dc1af940e663e5f0903333051bf01686d62fdebcdf927ca4ab2f63202a7cb30f

Observation 694255ab-4d01-495d-bdc6-466514bcc0c8 · outbound

This paper cites an unresolved cited work.

GTA: Grouped-head latenT Attention Unresolved cited work

Reference 128

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:51.479020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:53:51.139013Z digest=sha256:3215e9e4970a6df5ee7fdae54e35beaedc6d662c8e50353f2882f5bcd07ca37e

Pith citing papers

No inbound Pith citation observations are available.