Pith. sign in

Paper Citation Record · LEDGER

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs

As of 14 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2507.11953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11953 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:06:35.243279Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddce7f71-81e7-4bfd-a628-4677a6f16057 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.682069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.682069Z digest=sha256:297ae65e74cc538db126b970ea781ce4b9f7bab9945b7920f7ba58b49d7bacd5

Observation d27b0413-f088-48dd-80fc-b5682c8635b2 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.753842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.753842Z digest=sha256:0324b1b8d4aa0b4c0f2fabaaee1ce04ca850e73d9baa9c5af9e49fff8bc3ad98

Observation f2201d20-badd-46f4-9c6e-bbdf0975a507 · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs bert2BERT: Towards Reusable Pretrained Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:06:35.789463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T17:06:32.816564Z digest=sha256:b4040373b27cfee32374105c4459776964f4db80d56952d5c396e72a058bb553

Observation 78e2ff80-2383-4ed6-82c2-cf1fc53818fd · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.884394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.884394Z digest=sha256:66956ad974c2888c0bca7cce3558ae29452c05c13bb1fb33c0eaedf40ed0877d

Observation 031f1f96-1911-4e72-ac5b-db09b890c50e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:32.936687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:32.936687Z digest=sha256:b4787e7a75605ddf75f830c2525f864e16203be207910e24f7e29d1f9ebe5c32

Observation 4f5aa031-7ccc-4838-b0ae-49e0fbad6755 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Better & Faster Large Language Models via Multi-token Prediction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.019729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.019729Z digest=sha256:192f9d3e1886e63602d790f17cb97f39abccf3ce0cc1aca9cbe6f121ca01fca7

Observation 75bace4f-1661-45df-82a0-84828160eb0d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Measuring Massive Multitask Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.090241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.090241Z digest=sha256:3239c5af1bd6d666708c641aeba68bea9cc5b14dac53d1967c77a32da141729d

Observation 53b5b577-79bb-4a38-8752-585dbd12bc91 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.153280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.153280Z digest=sha256:e7096ebc18c412f2d2e7b1841e1e3208e14bf221c593116263f87c488071c980

Observation 5f539c6a-4fe4-4b54-b4d8-8b731cb81118 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.257455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.257455Z digest=sha256:23156c699e11d5f013c15ec90b1bf4ac5455909751b6b15b92031afb3488266d

Observation 8749c43c-2823-490c-ad32-138f9817c6e7 · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Le, Yonghui Wu, and Zhifeng Chen

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.314979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.314979Z digest=sha256:0db945d6acaf84b40e0982fa18b4ebbe09f3fa31b0f3e6b1e5e4c09ace8aaff3

Observation 7e1a6e80-93c7-4856-b09c-292eedb8e4cd · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.398030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.398030Z digest=sha256:babb6c40565179b83c60c3d4ecae1230ea8e1f1bec273f6c6c69743917b26645

Observation 747a2e47-ebf3-4362-a1bb-c91e9e816657 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.490400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.490400Z digest=sha256:294955f641ce5454278ee0ccf84738b07ec7981b6e5f9ac014e930d1822bb03e

Observation 881c46c3-0aef-4ff4-abc5-f4f2e8855838 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.575699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.575699Z digest=sha256:168a5641d32643355f672ffdc5096d204f86dc9678f2a40fd81d22989929088a

Observation 3b9396c2-32b7-4d11-8262-2481c8fbf45c · outbound

This paper cites Pointer Sentinel Mixture Models.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Pointer Sentinel Mixture Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.670665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.670665Z digest=sha256:da86d8be2bd22f8d057314e5dbf1754c8d878ca7a9a7132ef04e62202427c844

Observation 5f9fd941-db81-4ada-be82-f42b8d06fc5a · outbound

This paper cites GPT-4 Technical Report.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.745589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.745589Z digest=sha256:0ffc5a08e7dd2887b475a96b73ccdcfcaffb593146b74d17540223d0081bad7a

Observation 122e322e-7d23-48bd-a4de-6696608a97a4 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:06:36.081629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T17:06:33.821288Z digest=sha256:d4732d5c985f95c1d9f0548f73df37893cef826dbe54ff6c69a3ea1188af32dc

Observation 843ab966-ea9d-439f-a421-b05b56f66f4e · outbound

This paper cites Vicky Zhao, Lili Qiu, and Dongmei Zhang.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Vicky Zhao, Lili Qiu, and Dongmei Zhang

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.894797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.894797Z digest=sha256:96a53cc813a4ef3d2e65d15b8bec5b6723e975857a3d660da758626992f45413

Observation 42bc596e-6c1b-4d5c-90ba-f02346f3abcd · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.971861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.971861Z digest=sha256:43c5642b49986bfbe735f2543e576718d2c521be6b7f2d6591221bfd16c85c10

Observation 9ff05fd0-0beb-4fe2-b1de-a78e98b0afe0 · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.063713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.063713Z digest=sha256:68d30c6f4b1376d6b71a92a660446e2c8f4961d587806b4a4190d3acdad449ee

Observation 32e3c357-9e37-4cdf-904e-c874b813e0a0 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.181852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.181852Z digest=sha256:3862f79bc83e479a6c05e42753c62994e872dba6a26e3f7c3213b5bfc21d003a

Observation ef02b930-ad70-4fd8-8ef7-3bc981c325c4 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.234433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.234433Z digest=sha256:4837cf4602c9d4c80440e5674cbecbd58e0275fb4f7e23b0a07f6c4f258a934b

Observation d86a8ddf-d8a4-4daf-ad24-362ef31dba0c · outbound

This paper cites Hashimoto.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Hashimoto

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.307719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.307719Z digest=sha256:e8db9dc70b923839e123752ac7b71b67c8671dd28122ce3546f93c357ed7af64

Observation 0dd07ab1-e78e-4822-afa8-6174f15c2e8e · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.388738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.388738Z digest=sha256:dcd589d2b450cc1c4c442d91ae357fd23e67b8127fb9da3b7dcd445772d901fc

Observation 98605f00-7b41-4240-9f92-d3f4cf046fa6 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.508393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.508393Z digest=sha256:c9d9d23a6bb7084f2261a0610b02eb098ea2e38dcb408bdb3b955b0add6576ff

Observation 024ec14d-13af-4286-88b8-edf2bc45f38d · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.603114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.603114Z digest=sha256:6611b0183a4ae908d9af3cf92c4061be11f38747b6909410959656fb2a538461

Observation 2f06354c-eacc-4bd2-817f-72b6ed71ebb3 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Efficient Streaming Language Models with Attention Sinks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.684694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.684694Z digest=sha256:bdad525c040d970ea5aa92185e873e7379b9b92ed25ed1e11f5ce0f79f2dca56

Observation af505138-7ac9-42ba-afbb-1f5133866f8c · outbound

This paper cites Qwen2 Technical Report.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Qwen2 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.765386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.765386Z digest=sha256:67dba513656a164cd4c99bc25270c168cd47b28e4d40ead31de9e0cf7fb1ee47

Observation 264d52de-0845-4177-9c5a-82d06d16a1d6 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.843258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.843258Z digest=sha256:d3322db8c555d1de3af08e548766f55be4b2ba7e62d4dede32bbcbd293503dff

Observation f1b483d5-96a7-42fe-a269-73873527ec8c · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.915721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.915721Z digest=sha256:9a57628bc51c1e270f516dd47cf06394562f4530062215cb4e4ed8911f33b551

Observation ce0c4ed0-b0fe-47a0-abd3-c9bf3479c008 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:34.975662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:34.975662Z digest=sha256:40aec4e6f157958351ec8ebe796f3c82f12ae90a0396627b58569b4fe596aa3c

Observation 87ce1964-0ede-499d-aab8-8548b5c3b039 · outbound

This paper cites an unresolved cited work.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:06:35.943605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T17:06:35.039483Z digest=sha256:145f8e330089181b8103b536c26131a082fcd9b54dcc2c35b84885f789916124

Observation d0917189-3a92-4a68-a37d-d880cde302e5 · outbound

This paper cites online" 'onlinestring :=.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs online" 'onlinestring :=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:35.145905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:35.145905Z digest=sha256:dc483541517007c47e871578f7ced937d956b07f1e0b3e452c53059d571ca94f

Observation 09201d63-7bdc-4812-8495-6600582b6ce4 · outbound

This paper cites write newline.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:35.243279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:35.243279Z digest=sha256:4fb76ae25819c1b59b279699cda0fa609eeb4af9e7b67be1f2a76210ec04947f

Pith citing papers

No inbound Pith citation observations are available.