Pith. sign in

Paper Citation Record · LEDGER

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA

As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.20677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20677 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:19:23.112017Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:59:42.128875Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:59:42.230640Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7151ea51-0361-4590-87fa-d8dec4883b12 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.995212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.995212Z digest=sha256:33a1dfe25823e5a5fcf1913d26d0352d2ec8db070bf8e0173759d6a3b25ed062

Observation 7fbde293-76e1-4a51-b55a-3fe3ccb88c84 · outbound

This paper cites an unresolved cited work.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:19:23.525027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T23:19:23.107616Z digest=sha256:d675d42a1762fa0dd2a0879da9740eb753da57f00fc4812c1c5b30375ba6b523

Observation 7fe3f08a-d2e1-4909-a378-f717391ae2a3 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA LoRA: Low-Rank Adaptation of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.014442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.014442Z digest=sha256:07e2955acef3626b1d5861b474df20dcd1296c72da9533b719e56ea5bfe9b39e

Observation 343c113a-3b63-4898-b0d8-2cb8df6d0ffa · outbound

This paper cites Mistral 7B.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.019925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.019925Z digest=sha256:5005f055cbfc02d26019cf313b0a1297d8e662910c4e1912d253ccb34cb52e7e

Observation acebbc93-a900-4ccf-87e8-8f88b18c52f5 · outbound

This paper cites BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.024524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.024524Z digest=sha256:b93fa65a6718c20506819e0d33019a5181556003119ac73f086118daaeb45a84

Observation a489ebb9-8328-4c56-8ea7-5a33c9091de7 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.028493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.028493Z digest=sha256:cab8a397e409167bec1afd0424de8e3e0047678fa6d266cb88e5569a85c21136

Observation 3b049843-b8fb-466d-a225-3f858b14dfa8 · outbound

This paper cites Learning Sparse Neural Networks through $L_0$ Regularization.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Learning Sparse Neural Networks through $L_0$ Regularization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.032946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.032946Z digest=sha256:ca541778846d05120a18e7180e60661ae534254c9566b7e17742ed0ee66ba9eb

Observation b7705666-0266-47cd-b997-559697d8ccb8 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA SocialIQA: Commonsense Reasoning about Social Interactions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.040939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.040939Z digest=sha256:61c07acf8ecb08602223e1d34d87b4a3d4ba1d1291f29a0501f4b884929936ad

Observation 3f5afc6e-3c90-4e81-97b4-c96794f439e1 · outbound

This paper cites Recursive deep models for se- mantic compositionality over a sentiment treebank.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Recursive deep models for se- mantic compositionality over a sentiment treebank

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:23.552556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T23:19:23.049613Z digest=sha256:92091d8f5d6948f3450ab198109290d1eec86f8bc1746c8ac0830478fee989e9

Observation 45404514-2a8d-4b1b-9869-eafcc5300f46 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA A Simple and Effective Pruning Approach for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.053957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.053957Z digest=sha256:e131722339890a9566ec9bb1460741d46a60cc929458726e74d0aa08bd8101a8

Observation d7f1e43b-321b-4635-a518-8a98c59e813f · outbound

This paper cites RazorAttention: Efficient KV Cache Compression Through Retrieval Heads.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.058774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.058774Z digest=sha256:9b4966c9b7a36b44d8e481574a64d496e466192804365f99c9209274907d1eb3

Observation b62e7200-aa4a-4488-a2bc-c7cbe7f77dd2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.063117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.063117Z digest=sha256:9022cc3087d78a2f7ac2459640cb7fc22a12fccb9be696efd01fa3209875602d

Observation a5f7f156-3956-4295-bb7e-95440dce08e5 · outbound

This paper cites Structured Pruning of Large Language Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Structured Pruning of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.067137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.067137Z digest=sha256:113d65baf8a054d8d0f3d40a0326d4f49df2fa49fea04ac904c86d27264bf39b

Observation 62655e4c-e53d-406b-8e09-a916ae36eb06 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.075700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.075700Z digest=sha256:564e9582b1638714f9465e2ae98475df5949208d91f822bd9c0591a32f612594

Observation 948d91bc-580f-4a81-9f94-579e967b72e8 · outbound

This paper cites Qwen2 Technical Report.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Qwen2 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.080312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.080312Z digest=sha256:25112165739910883c99f7b3f5d13d6681a42095aa9d488f584c6ecf8150231f

Observation e49414fa-1196-4d67-903c-894eef64356b · outbound

This paper cites Effectively Compress KV Heads for LLM.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Effectively Compress KV Heads for LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.084595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.084595Z digest=sha256:12478075a58abfd79f9f9354fa37c9b98c8e690280b2eab2a91913c256725f0b

Observation 674a44b1-e114-4054-9095-f3990fc287f5 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.089199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.089199Z digest=sha256:9dc905ba0cd320bf99d5f89594a6f5f7af954360a7464067284746ed2cd29705

Observation 831db064-129f-4930-bd18-0ab67374e58c · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.093364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.093364Z digest=sha256:ecbccdce6960a832e75b635c6adc54bd89043a993301b0c1f8cfe8a585bd8a99

Observation 780554e5-c090-4460-85f1-7df43cb64211 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA TinyLlama: An Open-Source Small Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.097981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.097981Z digest=sha256:d3bfedbc97e14596018120e6a9144ee4f04d85e658e7dd77a569749465c5a746

Observation 61840526-4f67-4146-bf65-83deef55c602 · outbound

This paper cites During the pruning training process, the sparsity warm-up steps account for 30% of the total steps, during which the target size of the L0 masks decreases linearly to zero.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA During the pruning training process, the sparsity warm-up steps account for 30% of the total steps, during which the target size of the L0 masks decreases linearly to zero

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:23.539292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T23:19:23.102598Z digest=sha256:a9008a18d023abd493d831b752833fd4de1638cbb79e18a48aaa00e6a40298d3

Observation eac5dce5-d74b-4b33-85f6-e1199cb4fd14 · outbound

This paper cites an unresolved cited work.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:19:23.511078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T23:19:23.112017Z digest=sha256:d7a825ab2e312c0550951240539229f26f35afbe0a609f229b8c3e0ead59b774

Observation ab5f951b-3fd8-45ef-87d6-b0e6d373263d · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Fast Transformer Decoding: One Write-Head is All You Need

Reference 1966

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.045321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.045321Z digest=sha256:acbd7e3ae3cb8bca102b61553f3109b77aa117bdf19e6d48ad438710e51fb2a8

Observation b3067072-352b-47bb-8ca1-7faa285b5a07 · outbound

This paper cites DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.990805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.990805Z digest=sha256:8e25479de743eb1dc5af671a6dbf137709cbbff02aa775881ca6085640f96998

Observation 9b327448-53b0-48e8-b877-aeb244b96897 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.037009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.037009Z digest=sha256:579bad36a7d86587669590bf0a9073b232312ff805d49e7d42eb50eefa090878

Observation d5bbf37a-8162-4973-9119-fb744a724811 · outbound

This paper cites The Llama 3 Herd of Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA The Llama 3 Herd of Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.004497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.004497Z digest=sha256:4e0f1fb9993a370e396a5db4f70b68499349262aaef1357842ff874106ea008a

Observation 2f2fc995-ba2a-41ca-8213-8f1f88cd251e · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.999949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.999949Z digest=sha256:35e587a8ed1448ba69c59d40ba2df2afdddc4b61d6099054ad29d416f9187016

Observation cf5d0739-4783-4149-ab55-ad05197b6b36 · outbound

This paper cites Lan- guage models are few-shot learners.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Lan- guage models are few-shot learners

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:23.566901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T23:19:22.986705Z digest=sha256:43109342c4fd0269395932fdb5275c1660508899a8f6e0c6ba397c747ff0daaa

Observation ad30bf8b-f329-46cd-9743-077854a87204 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.009235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.009235Z digest=sha256:6263e2bf954e90d416ca2f42bf5c6dedd4f0d2d6fd62617f3e9a0a52c64ec1c1

Observation 6aaea453-9a7a-4c1b-810d-ebbda686b8bd · outbound

This paper cites Structured Pruning Learns Compact and Accurate Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Structured Pruning Learns Compact and Accurate Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.071189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.071189Z digest=sha256:4f8b4457169854a86b22a389cfb0602acf76bbc86d164b7ad0d2d0b4760b35e9

Observation a5f98dbc-0d0f-4192-a531-e065ae2d4a72 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.982078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.982078Z digest=sha256:08f56cc5a479c5cc48a09e7367be91bead457ab6d407ee1ce0216f6b5650fa2f

Observation 2aca0bbe-26d1-4b46-a4b5-fbdb433eec82 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.976195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.976195Z digest=sha256:8c414307824f0967c96ca8efc0c86cf22ed0401ea9771e2c64f2e17e8f1f07e9

Pith citing papers

Observation 6ad80293-41fd-4ec1-b430-229b314136f0 · inbound

Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques cites this paper.

Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:59:42.239517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:59:42.128875Z digest=sha256:4f22284d118fbc04d544318f5ab1bcd997ea11e5dd386b735e4f04e7ff437565