Pith. sign in

Paper Citation Record · LEDGER

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs

As of 9 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2507.11953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11953 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:06:35.243279Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddce7f71-81e7-4bfd-a628-4677a6f16057 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.682069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.682069Z digest=sha256:c71992340a2e2a3141a08b45352a78967181e3807622ae82940baeb0b5f918e3

Observation d27b0413-f088-48dd-80fc-b5682c8635b2 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.753842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.753842Z digest=sha256:ed5c90c938e15abe74b28fd4a2085f6771f5fd0b9c5b6d027caeafd2813ef0b9

Observation f2201d20-badd-46f4-9c6e-bbdf0975a507 · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs bert2BERT: Towards Reusable Pretrained Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:06:35.789463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:06:32.816564Z digest=sha256:02fdf56879c62e5ef7492fc798598a6f91273930821b53da11d40a4c426ba296

Observation 78e2ff80-2383-4ed6-82c2-cf1fc53818fd · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.884394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.884394Z digest=sha256:5dbeb5374d7adbbee1fb2771d7190b33ebedf6669cf1db67a56583fab91a3631

Observation 031f1f96-1911-4e72-ac5b-db09b890c50e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.936687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.936687Z digest=sha256:0b56caa6d50771722c335352b71ff62035ae8f7e4e767cc360cf5e36950b5b05

Observation 4f5aa031-7ccc-4838-b0ae-49e0fbad6755 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Better & Faster Large Language Models via Multi-token Prediction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.019729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.019729Z digest=sha256:38b1f234928cd39d3b4da8d597f70d9f7040c64cbf083fa6204d839cfffc8c7d

Observation 75bace4f-1661-45df-82a0-84828160eb0d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Measuring Massive Multitask Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.090241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.090241Z digest=sha256:1bc3f0f0a9e97e6c4e47e9507c545f0ebbe7340fa9b013ae5b97038cf5fe0e36

Observation 53b5b577-79bb-4a38-8752-585dbd12bc91 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.153280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.153280Z digest=sha256:42bc1684c5cabb47bc13c57ce6516ea99c36bf65f98ad897edc71ec0373248d8

Observation 5f539c6a-4fe4-4b54-b4d8-8b731cb81118 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.257455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.257455Z digest=sha256:b32f6bc0068920409ff69f78e2716953fa4b5ce0aabe7009c3fd2b501a381e21

Observation 8749c43c-2823-490c-ad32-138f9817c6e7 · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Le, Yonghui Wu, and Zhifeng Chen

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.314979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.314979Z digest=sha256:87d4057bd7c474a411b8b4952dcbfec507c35aacd6659cd0d2fb905d8975eded

Observation 7e1a6e80-93c7-4856-b09c-292eedb8e4cd · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.398030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.398030Z digest=sha256:72df509c05aca60bee77892d9c46e6f2e103393a67dedec46403525a1b3df7de

Observation 747a2e47-ebf3-4362-a1bb-c91e9e816657 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.490400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.490400Z digest=sha256:4db27e11788da507a0fa57d358d0f665e7475be41a413c8796dcbbd426d5dbbf

Observation 881c46c3-0aef-4ff4-abc5-f4f2e8855838 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.575699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.575699Z digest=sha256:ed2dd87c561a8304692b4cc054bb387a8ad97c3c1010e2830990d48f8276c630

Observation 3b9396c2-32b7-4d11-8262-2481c8fbf45c · outbound

This paper cites Pointer Sentinel Mixture Models.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Pointer Sentinel Mixture Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.670665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.670665Z digest=sha256:0fb1f64314c0c7e209c08428a6393dc9f5d883f6d7d46f5110956b4744addac4

Observation 5f9fd941-db81-4ada-be82-f42b8d06fc5a · outbound

This paper cites GPT-4 Technical Report.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.745589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.745589Z digest=sha256:39802c9db746a08fbc33a9a25fd291d42f3d69fbd130d5d9e12a1cf27e4ca90a

Observation 122e322e-7d23-48bd-a4de-6696608a97a4 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:06:36.081629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:06:33.821288Z digest=sha256:16513c7f84fb9f085f1d5cec7fef854deed7c5a52143ac4239bfc3db485191c9

Observation 843ab966-ea9d-439f-a421-b05b56f66f4e · outbound

This paper cites Vicky Zhao, Lili Qiu, and Dongmei Zhang.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Vicky Zhao, Lili Qiu, and Dongmei Zhang

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.894797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.894797Z digest=sha256:1ded4bb0ca4be3666d46173f4d36183a89ca3d556b2d4c529865c9cc17f06f45

Observation 42bc596e-6c1b-4d5c-90ba-f02346f3abcd · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.971861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.971861Z digest=sha256:4caab40cdd84f5098bf683de0ce31db89a016e04a480d5f6819d5a79b25e3f6f

Observation 9ff05fd0-0beb-4fe2-b1de-a78e98b0afe0 · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.063713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.063713Z digest=sha256:4a7f9a63b353f588b171af2dd21bdc7445852d44ed2ad936a2057ee97fbe77be

Observation 32e3c357-9e37-4cdf-904e-c874b813e0a0 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.181852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.181852Z digest=sha256:1959a2bc06e63f266ebc14aa6e62527773f70f1f2df2421c554d81b744fe27f9

Observation ef02b930-ad70-4fd8-8ef7-3bc981c325c4 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.234433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.234433Z digest=sha256:55c12546ba53c1d5c004edefe0aada2bf590646a58464d72c83de8ff1e539571

Observation d86a8ddf-d8a4-4daf-ad24-362ef31dba0c · outbound

This paper cites Hashimoto.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Hashimoto

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.307719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.307719Z digest=sha256:ecf6386fb52212d291fb18f25a4c8c21425e248fe42f3c06d46434fd95b09c1a

Observation 0dd07ab1-e78e-4822-afa8-6174f15c2e8e · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.388738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.388738Z digest=sha256:80a6bf239c5d4c3e9a9728d7af80fc8347411ed281889c1c43384595c3cc189a

Observation 98605f00-7b41-4240-9f92-d3f4cf046fa6 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.508393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.508393Z digest=sha256:23ed47f5dfa1e7ae8f8f7a4583b009328a167cd595983518424c9e84e22cdf11

Observation 024ec14d-13af-4286-88b8-edf2bc45f38d · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.603114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.603114Z digest=sha256:4b0fb0de22e13cb93337e938d615a911c2f409159f05ca2c6b7a8deb0c7e0d2f

Observation 2f06354c-eacc-4bd2-817f-72b6ed71ebb3 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Efficient Streaming Language Models with Attention Sinks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.684694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.684694Z digest=sha256:84ce7cc3d9243ac4670b2d795c52e6e6065b3ffb65a9542cca6ec3afacc3ceed

Observation af505138-7ac9-42ba-afbb-1f5133866f8c · outbound

This paper cites Qwen2 Technical Report.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Qwen2 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.765386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.765386Z digest=sha256:4eeb091a5b0e2ed4d71ced84a49063010407f2a94c158c8ac7e29b6a39919579

Observation 264d52de-0845-4177-9c5a-82d06d16a1d6 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.843258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.843258Z digest=sha256:a43a506b8835b253486b2f119173d610365dc9d654e0aba07bd6ffed7f6bce4b

Observation f1b483d5-96a7-42fe-a269-73873527ec8c · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.915721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.915721Z digest=sha256:b72dd584cc58405a5ee415cf3e5d788fc8b525cede611a2c390511c1867b99fc

Observation ce0c4ed0-b0fe-47a0-abd3-c9bf3479c008 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.975662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.975662Z digest=sha256:8ad3fca081287a575fa0407d9599f49414ae25a01324336337fffc832be5322a

Observation 87ce1964-0ede-499d-aab8-8548b5c3b039 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:06:35.943605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:06:35.039483Z digest=sha256:68c235b05735768177f07ffb1096b545e5ad32efd7a5c43aad1e919630b4f6ec

Observation d0917189-3a92-4a68-a37d-d880cde302e5 · outbound

This paper cites online" 'onlinestring :=.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs online" 'onlinestring :=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:35.145905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:35.145905Z digest=sha256:b80b6f54e49a9f4aac26986cdce057b3603d2a50793fe8f9dc9672d8639dda62

Observation 09201d63-7bdc-4812-8495-6600582b6ce4 · outbound

This paper cites write newline.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:35.243279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:35.243279Z digest=sha256:49ca994fcfb87ee4f91b2423e2049f88cec56014c6f38dfb2d63fb5f2c510c04

Pith citing papers

No inbound Pith citation observations are available.