Pith. sign in

Paper Citation Record · LEDGER

GTA: Grouped-head latenT Attention

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2506.17286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17286 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:53:51.139013Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 00dbff2e-f327-4928-9f0b-8d5c481e0104 · outbound

This paper cites Language models are few-shot learners.

GTA: Grouped-head latenT Attention Language models are few-shot learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.390682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.390682Z digest=sha256:4f8e9ba73ee681c58937df58f843180c3b23c9e7003bf2f2d2ecfc3f3eebbd4f

Observation 4ebbec90-b89c-41fd-be84-6e272170501a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GTA: Grouped-head latenT Attention LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.501981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.501981Z digest=sha256:3810348d1aba85133a49d667309f5d0870569d15af6fc9eb0f00832690845f40

Observation fc29431e-b635-49aa-bcbb-a996ea5d8ca4 · outbound

This paper cites Attention is all you need.

GTA: Grouped-head latenT Attention Attention is all you need

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.647788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.647788Z digest=sha256:e8beff1f687e241f6c208d38da3911a0efb7a6f8470b1f53d9ef23237a416171

Observation 920eb91f-e7c4-4663-aae1-85b1eb9c38cf · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

GTA: Grouped-head latenT Attention Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.764009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.764009Z digest=sha256:47709a8650778c8bc3f31e2d5061b8311f8e57f23691b857678a8a191f70d5f9

Observation 4ea556eb-1159-4dc4-8ae5-5fb29e61424d · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

GTA: Grouped-head latenT Attention Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.833695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.833695Z digest=sha256:28cac18a29f3a2c2ba8032c4822abd934fe7eef20c63cefb85102b01991b9fa4

Observation a8326140-3e94-4c77-9983-a9226041878d · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

GTA: Grouped-head latenT Attention Fast Transformer Decoding: One Write-Head is All You Need

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:47.967513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:47.967513Z digest=sha256:da4aaad4e215ceb7e36e48807d3ab3f8f7bc41308263578b69de0c5e7dcafb57

Observation b95852db-d43f-4473-af45-d4722c829c85 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

GTA: Grouped-head latenT Attention GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.108508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.108508Z digest=sha256:4f1b0bc3248896f9f8ab118b70c3a4418ef5f48eaac09204bcb3ca3c8297d72e

Observation 307edecf-d678-4982-b0c5-4dbd80e6415e · outbound

This paper cites DeepSeek-V3 Technical Report.

GTA: Grouped-head latenT Attention DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.226361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.226361Z digest=sha256:c57a93752377ce5fcdf47817a39999c72e086d1a83390e080126d08cda0fe3aa

Observation fe511906-48f7-45b3-a422-053b641aef66 · outbound

This paper cites Differential transformer, 2025.

GTA: Grouped-head latenT Attention Differential transformer, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:53.096465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:48.355571Z digest=sha256:12f44e9b45c9794f6f3daae2d8f8cfb64a08727b1154cc6cc58f3b1c5f1e9a9a

Observation b51b9367-075d-4340-8c3f-60a558981f32 · outbound

This paper cites Multi-token attention, 2025.

GTA: Grouped-head latenT Attention Multi-token attention, 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.972928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:48.456181Z digest=sha256:b850e0a4c1c39496882fed9730e63c01ff004c467ed0041618b3522141f1954c

Observation d2968eed-4a2f-4e4d-9508-2ed67aa706ef · outbound

This paper cites Glu variants improve transformer, 2020.

GTA: Grouped-head latenT Attention Glu variants improve transformer, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.518843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.518843Z digest=sha256:25dad37ad8dcf82ac8bebdaec7a5140a766b746169c0d25d95d6f094b977030c

Observation c722fbdb-1af5-4fad-bb5c-1a813315cc2e · outbound

This paper cites You only cache once: Decoder-decoder architectures for language models.

GTA: Grouped-head latenT Attention You only cache once: Decoder-decoder architectures for language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.849712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:48.572133Z digest=sha256:132a666a91adb2b1aac2da71a24f1b879adba94a544359538f7ae53ff79120c6

Observation 58ba30cb-17a5-4f4a-9488-88b050c1b950 · outbound

This paper cites Ni, Haifeng Zhang, and Jun Wang.

GTA: Grouped-head latenT Attention Ni, Haifeng Zhang, and Jun Wang

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.684881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:48.674845Z digest=sha256:a956d0c3f54385b38ccf140afe4f4c9b1e03483307ef8ca45dc5e9945890f0c2

Observation 08ad2c8a-ae4f-4aa7-8d96-49f80ec64a0e · outbound

This paper cites GLU Variants Improve Transformer.

GTA: Grouped-head latenT Attention GLU Variants Improve Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:48.777635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:48.777635Z digest=sha256:cee8308b7370d60f35f5467cdf53d1643d99a88eff3ec8fde4881b6db2b7eb6b

Observation 43b62c68-2d45-4000-a3f3-492cd34c5dd0 · outbound

This paper cites Gated linear attention transformers with hardware-efficient training, 2024.

GTA: Grouped-head latenT Attention Gated linear attention transformers with hardware-efficient training, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.535108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:48.880524Z digest=sha256:38368419c4784d7575f73a7fb5800b7c371d3538217bd2473add87fc2a2d4462

Observation 06bd6350-5d75-4be8-b3ab-1927f5051590 · outbound

This paper cites Hardware-efficient attention for fast decoding, 2025.

GTA: Grouped-head latenT Attention Hardware-efficient attention for fast decoding, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.367762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:48.956459Z digest=sha256:91bf8a4890ceb8ab77de533e88c173ba305962ed891ba700ba154fd1b82f93f6

Observation b0d713df-c335-4064-8255-7beb72fc8881 · outbound

This paper cites an unresolved cited work.

GTA: Grouped-head latenT Attention Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.088623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.088623Z digest=sha256:c9d92677f01429700278f8cd379bb69cf9efba653c88c8cf4d64fc6ae3ef4083

Observation a9eaddfe-0c3a-46f5-8b68-5036f0ad77db · outbound

This paper cites Decoupled Weight Decay Regularization.

GTA: Grouped-head latenT Attention Decoupled Weight Decay Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.214988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.214988Z digest=sha256:7b9b692147d49ea4ec0c642ce0ce4efa5aad38b7d40d46b1bf0d5795993daac9

Observation 56d20d05-d48d-4d11-b3e2-25b021c85938 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

GTA: Grouped-head latenT Attention TinyLlama: An Open-Source Small Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.285407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.285407Z digest=sha256:944284fd90caf904bc10077d84b3fa88b71775252b001c7973d04947c06dcbac

Observation dc9c92cb-2732-498e-8ab7-17540c8354b5 · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017.

GTA: Grouped-head latenT Attention Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.369240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.369240Z digest=sha256:94d1e968529927ae5e9c26ed35c903cf2b525c8d9cb9ae12caa7126106b8f683

Observation 77e9d984-8634-4cb3-8730-055289ba8ece · outbound

This paper cites Relu 2 wins: Discovering efficient activation functions for sparse llms, 2024.

GTA: Grouped-head latenT Attention Relu 2 wins: Discovering efficient activation functions for sparse llms, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.194831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:49.498212Z digest=sha256:0551ae26170bfa78f850179c45494f9a9b43716b160b58371cfa08b6a20a233a

Observation f4d3644c-b984-4d45-8a23-4bd89a89d1d9 · outbound

This paper cites Smollm-corpus, 2024.

GTA: Grouped-head latenT Attention Smollm-corpus, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.632203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.632203Z digest=sha256:f1d531d9e9939d6889d5b289470298daf8ca7974aaf0b0d23c5ff7cce59b97e0

Observation 5c8ba0fb-51db-479c-8c4b-9ebddd7450a5 · outbound

This paper cites The llama 3 herd of models, 2024.

GTA: Grouped-head latenT Attention The llama 3 herd of models, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:52.063682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:49.697068Z digest=sha256:410dc02e5ad925fa864a87f83280533db3d8d200f5517b08626858f07466f22c

Observation c2ef41e4-dfdb-4af5-9f77-e9cbe4ad480e · outbound

This paper cites Mobilellm: Optimizing sub-billion parameter language models for on-device use cases, 2024.

GTA: Grouped-head latenT Attention Mobilellm: Optimizing sub-billion parameter language models for on-device use cases, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.814149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.814149Z digest=sha256:9085e726816c6f7f0a6f56bbf35b986eb1b844dfffde471153cd7eb541b7440a

Observation 3cf808b0-f001-4041-9900-5ca672e55c3d · outbound

This paper cites The language model evaluation harness, 07 2024.

GTA: Grouped-head latenT Attention The language model evaluation harness, 07 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.960808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.960808Z digest=sha256:e99207b1587b627f6af1bcd6bfcb2bbfd58cc9368aba6e73993d8a9d51340e99

Observation b0ea5114-81a4-4763-823d-df4477e0d4c0 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

GTA: Grouped-head latenT Attention Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:49.994896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:49.994896Z digest=sha256:437f23b681ee33761b9320e28b675a99bd13184317037bf821f45c3ab9878cfa

Observation 49bdd69e-18ce-4294-a2e0-0dba32490221 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

GTA: Grouped-head latenT Attention Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.081508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.081508Z digest=sha256:bf5b3ad07834476b195ccce695675dddbfbca29f8deb1ab2e7a8240a9c892b66

Observation 0b0fd502-af34-455e-b466-23472381942f · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

GTA: Grouped-head latenT Attention BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.161101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.161101Z digest=sha256:2e39fda7e5b23b79db0e3291c28bb60c9b6ddf3cda1cd0cb83206fc161b7050b

Observation 87aafeb8-9661-4c83-8d04-119cb1a6d60c · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

GTA: Grouped-head latenT Attention Piqa: Reasoning about physical commonsense in natural language

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.214470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.214470Z digest=sha256:feaabae914a9adcef9e5104ef7b319c98bcee1341f7262514c01e1b4e75a7fd4

Observation 07b4956b-e65e-4d1e-af13-605dc5505eb8 · outbound

This paper cites MathQA: Towards interpretable math word problem solving with operation-based formalisms.

GTA: Grouped-head latenT Attention MathQA: Towards interpretable math word problem solving with operation-based formalisms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.940208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:50.285598Z digest=sha256:52b6cf28f586b4f1850b32379876df82375a9b54024d2677d47c0e3bf533d5ba

Observation e93a7f2a-d47a-4b22-95f9-99829db64080 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

GTA: Grouped-head latenT Attention TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.363927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.363927Z digest=sha256:6e00fb219f0c4512af1ed3ae054979e18ac7fd8d8304178dfffa54f5c755b7e9

Observation 18f11fec-a649-47e8-ba35-77e51e755d9a · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

GTA: Grouped-head latenT Attention SocialIQA: Commonsense Reasoning about Social Interactions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.475025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.475025Z digest=sha256:ba9ff6c2203339b39e1727e6885835abf4507ac0bdabc53301af19cb5b19b453

Observation 7ab3499b-47f7-44f7-8964-9f5ad47448e7 · outbound

This paper cites Program Synthesis with Large Language Models.

GTA: Grouped-head latenT Attention Program Synthesis with Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.540047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.540047Z digest=sha256:8a7bfdb8176ce09999de76bf31229fdcbdf9cfcd6d931a03a00b01291396aaba

Observation 8740e4f0-43d0-43c6-8de4-6ea2b9f9383d · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

GTA: Grouped-head latenT Attention Instruction-Following Evaluation for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.616735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.616735Z digest=sha256:6faa617bdb6f63f40b803f0a660fb4f75ea17a021f38df718eddc33dc070658c

Observation a82b3e84-a2a8-424b-a46b-a42348b96bd7 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

GTA: Grouped-head latenT Attention LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.701663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.701663Z digest=sha256:2693e775dac9f40a75b3506e04f5bc20dc0ea3bbb9e6445a4d60a812c97ef4ae

Observation 33c5ad1b-d64a-455b-8b27-754e019856c8 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

GTA: Grouped-head latenT Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:50.764532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:50.764532Z digest=sha256:a7682fac6378d35f047b85e77a5955f76c41d00b7ec69836916fc979e4db9f0e

Observation d7d3f5f3-3534-46b5-aca8-647aaa316281 · outbound

This paper cites Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D.

GTA: Grouped-head latenT Attention Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.801241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:50.841526Z digest=sha256:4846ca92a8fe95810eebd7e76176efdb97ca78d1b7a3315634712f6718f438f7

Observation e837571f-512b-4c22-8492-1a7c3854961f · outbound

This paper cites Llm inference unveiled: Survey and roofline model insights, 2024.

GTA: Grouped-head latenT Attention Llm inference unveiled: Survey and roofline model insights, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.694951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:50.920455Z digest=sha256:fc4899df9a234e76869ebc6ce35602262d6cae6990dfca3af97c24dcf1ae44fa

Observation 9aa579f8-cbfa-4cb2-9c23-1a82b430a474 · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

GTA: Grouped-head latenT Attention Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:51.000364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:51.000364Z digest=sha256:0240e6cdeaa17e75671860d68aefe961639f7910169273567cf0ce05a90afad3

Observation 4b2d1f06-2342-48e4-9fe5-3cc590977f97 · outbound

This paper cites Validation.

GTA: Grouped-head latenT Attention Validation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:51.571847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:51.079470Z digest=sha256:8b9b6478361e647a35d17d98018a0ab150bd09408f9c219c1c43ba2422abde3d

Observation 694255ab-4d01-495d-bdc6-466514bcc0c8 · outbound

This paper cites an unresolved cited work.

GTA: Grouped-head latenT Attention Unresolved cited work

Reference 128

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:51.479020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:53:51.139013Z digest=sha256:eef908e13cf94d0fe1a727307144b36e1478952571748a14f7e39df2c1bd27e1

Pith citing papers

No inbound Pith citation observations are available.