Pith. sign in

Paper Citation Record · LEDGER

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

As of 22 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 3 inbound Pith citation observations for arXiv:2505.19602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19602 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:20.927090Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:03:03.062322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T07:58:07.594523Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved75
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 293a0609-f139-43f6-8037-2b43a274623d · outbound

This paper cites Longformer: The Long-Document Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Longformer: The Long-Document Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.031316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.031316Z digest=sha256:07968efa6ab4a76109815c8e71e707f7296491a30f260ea6d861348605310811

Observation a6f5049a-cb1c-4012-92fc-54079f5ef1b7 · outbound

This paper cites Improving image generation with better captions.Computer Science.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Improving image generation with better captions.Computer Science

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.089620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.089620Z digest=sha256:5178509bf367ed9c454ab0737e0642859147af97f51a6b18b2a6106a0a9bd308

Observation dffea209-3858-47ca-a9a0-3affff8ba103 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.217814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.217814Z digest=sha256:e01b678ada2607981a60dd0d169698c53a5d8fb903a28f81460980c2a085385a

Observation e9ea3931-fa9a-4d77-8236-6c48319287f0 · outbound

This paper cites Maskgit: Masked generative image transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Maskgit: Masked generative image transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.340496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.340496Z digest=sha256:585cbbfb772a57e7d6ff9a2eed435f3a831392ccc23e6152e31e526b862a2a71

Observation 646555be-1f9c-4df6-9c80-d2e0bfdc8a28 · outbound

This paper cites PixArt-Sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PixArt-Sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:24.117531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:16:12.414680Z digest=sha256:6fb66630ed2222b1e3a8e66380a50d0d46a8c11f3fbc9cae8f70533eb2c6af6e

Observation a14d7039-3002-422e-b24d-f6c3e6c903e8 · outbound

This paper cites Generative pretraining from pixels.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Generative pretraining from pixels

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.543575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.543575Z digest=sha256:dd05114b9092d985b689f4eb5d206d02fdb2ed41af1b4469fd70787d511bb290

Observation aa1f79b0-8e6d-406e-8dfa-1e8a18d509a4 · outbound

This paper cites Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.670870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.670870Z digest=sha256:f26d95a18e4b2fe2a539600ae6bf09e9b9508387a444ec966389882cd21e46c0

Observation 0fad3e3f-9146-4a5f-beae-d73abe0a52dc · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Taming transformers for high-resolution image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.835884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.835884Z digest=sha256:469b7114a6a122d126f2f249afceb6775c959df361b3093e49356dc2f17c28b9

Observation 0a678a28-5f88-44c4-8487-7a8823fc6d0f · outbound

This paper cites TinyFusion: Diffusion Transformers Learned Shallow.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression TinyFusion: Diffusion Transformers Learned Shallow

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.014126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.014126Z digest=sha256:80d065b1aaf3857d8c4fe0f073198657fe17abb0ae96026727dd15483aa5272b

Observation 2e6847e7-cbbb-4746-be02-7c97c4414363 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.138387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.138387Z digest=sha256:48fc463aafbd0e193472701e537d492d80e5d5344cd4f6fc8db7fb0fb14b5428

Observation 8fd9a677-964d-498e-8609-a9f486163db0 · outbound

This paper cites Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.258212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.258212Z digest=sha256:62c328d0f546dc3bd91cd21016c531459f74454f65551605b8750de4b025d31e

Observation 814de7f5-fdf4-4a33-9933-76925e214302 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.413395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.413395Z digest=sha256:472e95b81f6daad8679d24ae8aa8e16e6d5aa9c29b0066e13fdb650fc4fd5a05

Observation dc481280-ccd0-491c-af0e-50b4974a2be8 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.862112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:16:13.530092Z digest=sha256:95c25bd2f2e9e7afd82b5314ccd17f0527cac49966a85c9746637e2c5e9c1454

Observation 22f15021-513d-40bb-a2ad-7b7af0fd1226 · outbound

This paper cites DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.622854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.622854Z digest=sha256:b513616473ca055ee8385c9aec0a5b1e83ccc4a0af6a5e54d3d8d53a345be7ff

Observation 0d51690e-64bf-4e93-9317-19ab2db69ce6 · outbound

This paper cites FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.693484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.693484Z digest=sha256:0dbdffc195f9addb4c8133d36eb0fcb1403ae12120eac7e1a992c15d0f05e97e

Observation c8144f2e-57d3-4053-bc65-a0aff67de974 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.770499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.770499Z digest=sha256:fd8b39def9fa8592550ffc820d5c364f0eae6d4148fae090f29e57bf561ce37c

Observation e9cccbfa-a369-4d8f-bf85-2dddd7291d5f · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.857590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.857590Z digest=sha256:0d3484df08a0206fdfb6081b081db6c8d356e1a86c3c0f0aadadb82c25862194

Observation 90c1a4be-0e2a-4a36-8855-c9899a98340c · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.948063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.948063Z digest=sha256:e7407976551637f34864489cfd03929dab89fd550e410d806c4b775dbd98b280

Observation b128b49e-cda3-4642-9df2-5a338a5297d4 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Distilling the Knowledge in a Neural Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.033748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.033748Z digest=sha256:f6acc6cad8598cc2281a25bbff9d9c3cc2ed40e1832bc958d051656020d86881

Observation a26ea0ad-6364-4f16-92b9-5d354a338c16 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.124213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.124213Z digest=sha256:7c17c7cc17c0f50381efd8e69b3081303a349e3f89cad46bdfe2764087e3f8ed

Observation e4045744-1d75-4e87-b00d-b7d213d10524 · outbound

This paper cites GPT-4o System Card.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.194123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.194123Z digest=sha256:539c7b2ea7ffaf21f681e65d31cbaa30940347c52d7a80b520cd399edcc932b9

Observation 38eff18d-581c-469e-8f75-d9b88ac6bfaf · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.282281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.282281Z digest=sha256:5ef5a81b1fdd3be86e2a09c5ec7f6419123c0a7f785004586e8f047b148ced71

Observation c1402ed7-f490-4546-badd-4fd7f608e0b5 · outbound

This paper cites Scalable autoregressive image generation with mamba.arXiv preprint arXiv:2408.12245, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Scalable autoregressive image generation with mamba.arXiv preprint arXiv:2408.12245, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.444559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.444559Z digest=sha256:57da0886bbf429dbe62a07eb49f66bc891e5006aca9e519f9b35e33fba8c0593

Observation 2b8bf013-ee0a-40cf-bd57-999abc287a7f · outbound

This paper cites Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models.arXiv preprint arXiv:2411.05007, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models.arXiv preprint arXiv:2411.05007, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.611793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.611793Z digest=sha256:636d936e0a4cd60bd08b508d129f03326f8d9dc73cd0c7cf90cd8ad9f0815d01

Observation 06423f51-cfd0-4ae9-88ef-0b4c16487a02 · outbound

This paper cites Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.731278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.731278Z digest=sha256:fb108ace91c92b66ed8126fcb0714a0364ae9bcdc13e70d11747945776c05cdb

Observation cb8defb5-873d-46e9-9eb1-3f9540d81714 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Autoregressive Image Generation without Vector Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.884901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.884901Z digest=sha256:f56b47e7c3c68519f6ae9051d09e2909e951b7583df0a4748c000f2c69b539a1

Observation 62903964-d2cc-40eb-8e34-f2ca28e166b2 · outbound

This paper cites ControlVAR: Exploring Controllable Visual Autoregressive Modeling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ControlVAR: Exploring Controllable Visual Autoregressive Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.022702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.022702Z digest=sha256:6463fff8226035593a3dd5cdc7138cd29f81da26eabb92fe4db99dc30ab61496

Observation 4026be12-2d5f-44b1-a227-95b768acf7cd · outbound

This paper cites Q-diffusion: Quantizing diffusion models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Q-diffusion: Quantizing diffusion models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.125418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.125418Z digest=sha256:089078706d11a3d6a9c0f064ecccfaf70cf26ca421a8d57930b9bbca9e505ce6

Observation 29d6db59-5830-46f1-a198-ef87d091c6e3 · outbound

This paper cites Snapfusion: Text-to-image diffusion model on mobile devices within two seconds.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Snapfusion: Text-to-image diffusion model on mobile devices within two seconds.Advances in Neural Information Processing Systems, 36, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.563170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:16:15.281994Z digest=sha256:dcd0988cef0ffaccc230db3b3186a0a67bf394f2aa59b522722954c9763faa61

Observation 8e0543e4-440f-4b44-b65f-4a42fa18ac23 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression SnapKV: LLM Knows What You are Looking for Before Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.378042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.378042Z digest=sha256:3b31424f816a9fa5168fb5d6243c5afb544d2e46a1f596c0b16344bf035a9ae0

Observation 6fd371e8-a244-47b5-b834-3ef570d38900 · outbound

This paper cites ControlAR: Controllable Image Generation with Autoregressive Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ControlAR: Controllable Image Generation with Autoregressive Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.523883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.523883Z digest=sha256:27d9e53320ebf7242cbfc292e4394a9772dd43b13eb931d10dd7302cca4911fd

Observation a495cb52-b5cc-4ba3-9c05-bbf931b8765b · outbound

This paper cites Microsoft coco: Common objects in context.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Microsoft coco: Common objects in context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.657471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.657471Z digest=sha256:3de7379531bcd0968ef11e3371f847154dde95fd67b15ee74bab7883ecfde32b

Observation 46aec037-8f7c-40db-bf10-3ec3a2194891 · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.805734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.805734Z digest=sha256:699c0427959882b95391f5bee2ea8f24ac40288d9ed946493b8db2ce2920b9e4

Observation 1adb9af2-121c-41eb-81d2-c622f8fed257 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.891755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.891755Z digest=sha256:63980d931227ad05f9aa2d5f25edca7375ceaaeb866342fefddb13a111296b6d

Observation bebf9bab-5a0c-4ab4-b687-370c985bcc0a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.Advances in Neural Information Processing Systems, 36, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.016184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.016184Z digest=sha256:8a00ff6b8004c2b20347a16b4c9873d27d4e319f7cca37314e4ba7ba340cc7d3

Observation 6c237317-f6f2-4a3b-8185-92f5aef59083 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.195264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.195264Z digest=sha256:2870cfa9d233e4cd376206723e6c4cf0b3c367a29bfaba2bf2987cc0474f71e9

Observation f0675899-2855-42d8-8823-4a6794e46234 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.358206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.358206Z digest=sha256:fcbe615dde4b24b93093d94096b3154ab44a23df434609458646aa6834d7d2f7

Observation d5edde25-459b-4fdb-8432-9e94b657d3a9 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.536400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.536400Z digest=sha256:2be0e46653289965efb438e92694829b45c450bbfe1a7236869e24607392ab90

Observation 6133ddb2-fc52-429d-8bbf-4dcf6ebadf01 · outbound

This paper cites Accelerating Diffusion Models via Early Stop of the Diffusion Process.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Accelerating Diffusion Models via Early Stop of the Diffusion Process

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.665647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.665647Z digest=sha256:42ef459cdf52b9e1f4d49b62c6c29b0b89ba66e4f6fc62066dabc3a1e042728a

Observation 4eb1f813-b788-4468-b191-0668b04482d3 · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.780411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.780411Z digest=sha256:581ba602726531fe5f5a60289629849060ef53873534fd4c3359163f2e0f1995

Observation 2eec77e1-4b91-470e-85ac-3b668055466b · outbound

This paper cites DeepCache: Accelerating Diffusion Models for Free.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression DeepCache: Accelerating Diffusion Models for Free

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.890592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.890592Z digest=sha256:e4a3cb56e49f4a036149673228a6b4d510172377f3f6e146374c3986e54c313a

Observation c33b358f-77c5-4223-bb69-6eaff1831127 · outbound

This paper cites Transformers are Multi-State RNNs.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Transformers are Multi-State RNNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.009571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.009571Z digest=sha256:97e033cce7444e2275931c8e19693348fecd780f455301c53b7a3f974aed755e

Observation 3c17afa5-d337-45c2-9e6c-6c4729a30143 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.118171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.118171Z digest=sha256:fa7c6a77eb0f6b6883f8f9dacd5e4a564b544469a995416c4c34a2a0f1512b04

Observation ba685361-ef66-4fa7-89e4-8b276a81e4f1 · outbound

This paper cites Head-aware kv cache compression for efficient visual autoregressive modeling.arXiv preprint arXiv:2504.09261, 2025.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Head-aware kv cache compression for efficient visual autoregressive modeling.arXiv preprint arXiv:2504.09261, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.201417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.201417Z digest=sha256:e0ffa1889903b47c64c56857fe4a67ae0206ad14a03a040f6392bd2101c98580

Observation e9c6dae5-05e4-4040-a6e6-e81f9537114c · outbound

This paper cites Efficient Autoregressive Audio Modeling via Next-Scale Prediction.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Efficient Autoregressive Audio Modeling via Next-Scale Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.266768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.266768Z digest=sha256:70437ccc512f969862058528922ab372ecab7cda1cdcd063863427cb55200bf5

Observation c21181a6-5cea-470f-a97a-05134b04e7ca · outbound

This paper cites On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.384573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.384573Z digest=sha256:350122912d1e1cbfb62cb2505a4bf1b8fa90ffeea8279a92e50c382b31d8db5d

Observation 4a14f778-46c6-4d1f-a4ec-dcbf3ab8e720 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Progressive Distillation for Fast Sampling of Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.490839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.490839Z digest=sha256:6f2de5dbbc4557888188e97f9cfb922ff1f94c34a7294fcf308ef09d12e17a79

Observation aa9df552-5935-49c0-8aa1-7d1032654f6d · outbound

This paper cites Adversarial Diffusion Distillation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Adversarial Diffusion Distillation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.578156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.578156Z digest=sha256:8d89d4fe3dbc1fef474a052a8dab65a28a56c88c992be610cf0162322bd26996

Observation 3d0a08cb-ef61-469a-b213-c19bf6a9ab4c · outbound

This paper cites Post-training quantization on diffusion models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Post-training quantization on diffusion models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.679607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.679607Z digest=sha256:cd47d63de36db76b476740641ac36fa9872a373f503713e09bed056ab08d8d68

Observation a8693224-84aa-45a6-b7c3-dfc107049171 · outbound

This paper cites FRDiff : Feature Reuse for Universal Training-free Acceleration of Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression FRDiff : Feature Reuse for Universal Training-free Acceleration of Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.772212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.772212Z digest=sha256:b407075290d24eb1adf95020c929b377762f7f64ee25164a969984d91cbd5778

Observation 353a2226-eed9-4e36-903e-a49527443e35 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.009369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.009369Z digest=sha256:28acee3d174d06f3e6509830d23ebf01d9c79a7920969231af0db61bd6792ed7

Observation a8d75859-c32a-486a-af2c-e1c44e89e10f · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.127539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.127539Z digest=sha256:a2e7fb53eb8e04dc54fbead6bea0c80d3166c07677a0bd3d2c58c3198fbe032d

Observation 4080b004-8286-4e92-9045-62ccd4f9bc2c · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.201368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.201368Z digest=sha256:988ae7de527dffb7b81f407769776c970fb7b05084de49bdd54df93c77f07cea

Observation 456bd8c8-8727-471a-935e-d761f58b42f2 · outbound

This paper cites Conditional image generation with pixelcnn decoders.Advances in neural information processing systems, 29, 2016.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Conditional image generation with pixelcnn decoders.Advances in neural information processing systems, 29, 2016

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.298675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.298675Z digest=sha256:28068b44574d9023ae6171086cba694b9fdad7cbb4efbb57700ce922c09dbf1a

Observation 1cb61cd6-bea0-49cc-9630-78d1767074d0 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.412301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.412301Z digest=sha256:7f32e6b16cf53825839caa82e21e7c7aa5ec46d334b404982882c34980892f5a

Observation 735fab92-1f44-4700-825d-a90d83fc65a6 · outbound

This paper cites D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.517957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.517957Z digest=sha256:c05b8eac31657dc5bd5118c9801b6181065f0d2e8deef5df0d83b1985d6f83e8

Observation 1cabc8b2-404d-4115-8940-3fe58a98ed21 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.606283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.606283Z digest=sha256:fb6b9245836cdeaf9d936dd17bfaf50eff57699ed5de84d1ead87f1225794e03

Observation aa87dea9-225f-47b8-8879-38d93ae314cb · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Emu3: Next-Token Prediction is All You Need

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.699512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.699512Z digest=sha256:4186b0fab2194e5512586ba740b156c1cb8b7f24a5175606609ee5797098d2d3

Observation 975c5e03-be46-46c4-8191-fea0a7ff031d · outbound

This paper cites Cache Me if You Can: Accelerating Diffusion Models through Block Caching.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Cache Me if You Can: Accelerating Diffusion Models through Block Caching

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.760250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.760250Z digest=sha256:c68008bce02ce79c5fd841c5f4cfef44a6d669569b6a49264d3235e0a5c6ae3c

Observation e9eb3f00-e6d4-4429-849e-e68f1db1311f · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.842698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.842698Z digest=sha256:31fefc1c599d129b7ae0fad8c79e085a9d44e631034963c46fd79cd6bb7ff17c

Observation 545bebb5-615d-4ecf-b2f1-754fcf5878c5 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.961783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.961783Z digest=sha256:b1c266a33b9162613f42a620d636f8f3dc6b7d36ff1a63fe9b209418c03ca465

Observation bacfe7e9-9139-4137-abab-317b6f2257e6 · outbound

This paper cites Efficient streaming language models with attention sinks.arXiv, 2023.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Efficient streaming language models with attention sinks.arXiv, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.047855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.047855Z digest=sha256:7690b258d1778c1dd618e425bbd4530d04aec554a5d6765509c2272236d5b8b8

Observation 8655804e-2349-42e9-9902-0e4d05f68341 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.134844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.134844Z digest=sha256:c4fe9295da10818767d74918a89190599fc3ff8529d3323bdeda78e193b8a16f

Observation f56e95aa-f6be-4e5f-994b-a7f82c45c05c · outbound

This paper cites LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.250138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.250138Z digest=sha256:674ef0c9913e3fa2e79a3f9fe19564ddbc0c7e56c38c242a9e9b5487ffd5bb48

Observation 827a1ae8-2935-4ac9-96d7-895870c2b1be · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.360068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.360068Z digest=sha256:c8922b97fc2cdfc65ba278238791aa370c15b316ac62b0fe1210abab7e48fc5c

Observation 8d6893d3-7a1a-463f-9bef-f80db39d814d · outbound

This paper cites Hash3D: Training-free Acceleration for 3D Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Hash3D: Training-free Acceleration for 3D Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.475008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.475008Z digest=sha256:71bd159a5e4cc5cba1591b762cd79feb6cea8a5c92cf11e008e68fcc451a5c2c

Observation b7a8c97c-e8de-47ce-bb5a-d8e082b59bd0 · outbound

This paper cites Diffusion probabilistic model made slim.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Diffusion probabilistic model made slim

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.209684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:16:19.548827Z digest=sha256:94228b4b742b19ddb042da8501e84a40d0bf37a7965e1bfbb628a5e4c1b35fe1

Observation f523d14b-49a3-43af-9877-e8c9330bc139 · outbound

This paper cites CAR: Controllable Autoregressive Modeling for Visual Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression CAR: Controllable Autoregressive Modeling for Visual Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.608040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.608040Z digest=sha256:6cc20db306cdba035879c28e91fc24b7585764c439cb6c14eff6318117a0bec9

Observation 0f4c3011-d88a-411d-8bb5-7ed536fb36b2 · outbound

This paper cites One-step Diffusion with Distribution Matching Distillation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression One-step Diffusion with Distribution Matching Distillation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.701607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.701607Z digest=sha256:affeeaec62f0dbf205b92f09e274dffe9404080ab28d3abf1f3b1a04c1da5563

Observation 0e775c07-916f-4896-ae0c-6090247f3f33 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.784103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.784103Z digest=sha256:ecc2a0e90a5f9ea69c116ccecb24dbb6105bf165041128424780fa641a17b3de

Observation 3557f87b-5aab-4826-a750-14232c37e679 · outbound

This paper cites Resshift: Efficient diffusion model for image super-resolution by residual shifting.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Resshift: Efficient diffusion model for image super-resolution by residual shifting.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:22.897167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:16:19.898686Z digest=sha256:9e690b7642cb531273b40c6db7e74002a75e1bf0f9a5f0f33beec0ff63959869

Observation affe042c-ddf2-4a81-9f3d-0a82e5985e68 · outbound

This paper cites LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.004397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.004397Z digest=sha256:66a658e90f4041b51eab2d97a0daa12df23216c434734166f8fe598b56e5b982

Observation aaa1b858-7fb0-4d7a-8413-b7a35de4a838 · outbound

This paper cites G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.116154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.116154Z digest=sha256:4af9b4fd07362a3de955bc4f6ba71ee9a5428a014716246dd79f3d04968b3b4e

Observation 91b9fa70-7ee9-4c17-b71a-2ec91cd63095 · outbound

This paper cites VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.208413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.208413Z digest=sha256:c4dc10a513aa148411d9564450d584db9d84da1a82900d25f34e6c760687e237

Observation 9e2adf34-da4a-445f-a409-4db203c50c08 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression The unreasonable effectiveness of deep features as a perceptual metric

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.320593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.320593Z digest=sha256:eea8377c0d8d477b73f31fb8d8706c69b30fcffc5325f630d44e48de7193b1d6

Observation b14eb1c8-94d5-4847-9600-356c9fdc9949 · outbound

This paper cites Faster Diffusion via Temporal Attention Decomposition.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Faster Diffusion via Temporal Attention Decomposition

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.398774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.398774Z digest=sha256:cca7526ba9f4250aa7da0f15dfeaafaabe90bd1e1c17912fc0d7f61434427b25

Observation 55a8508a-c4ac-4be5-804e-d2fb45026751 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Cam: Cache merging for memory-efficient llms inference

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:22.625939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:16:20.509944Z digest=sha256:7e32716f4f3e1eb61953f54b7d70fe128e517b27350f0adeee800897b7455959

Observation c37d9c57-6909-4482-b493-26005cd2d039 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.627739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.627739Z digest=sha256:77ae61e4f68c9ed81661eecf3a9356bbe153f50bb4b2a1d3ee5377879fcd99c9

Observation 30e5a746-e8e1-4266-af0e-8df43ff5936d · outbound

This paper cites Real-Time Video Generation with Pyramid Attention Broadcast.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Real-Time Video Generation with Pyramid Attention Broadcast

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.706471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.706471Z digest=sha256:c146576e926ef051e845200972818a2b6b90a6807709ff51479d0a360fca936e

Observation 76094485-c0db-45ea-b5ce-5a4b36647ba1 · outbound

This paper cites MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.821528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.821528Z digest=sha256:d654db3d51223b7e9ed6b0eb8915a1fb1d8537c3692135f4a049555d05426d03

Observation f972b876-9b38-4a37-98c5-2063619a429b · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.927090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.927090Z digest=sha256:e56ff23dc8de79af3118956489aefb208d1bea7773953f80bf336e8ea6275da5

Pith citing papers

Observation 1a3ca253-3931-42c1-bdfb-79ffa010f511 · inbound

Visual Implicit Autoregressive Modeling cites this paper.

Visual Implicit Autoregressive Modeling Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:18.325747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T15:20:03.981830Z digest=sha256:626a64bb8d886361aafe91748e88a6f351921e7e9c9ef8f8a358a08f57f3986d

Observation 348b6ba1-7a44-4b8c-91ef-076c5be59bb4 · inbound

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond cites this paper.

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:58:07.596260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:57:51.032025Z digest=sha256:d9291dfe2ab96160e951bd30c179195943c1647535b8f2130855df6e554d99b2

Observation 14f8186d-8235-4945-ba0f-c792e5ebda9f · inbound

Token Radius Attention for Efficient Video Generation cites this paper.

Token Radius Attention for Efficient Video Generation Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:03:03.062322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T06:03:03.062322Z digest=sha256:0348e9e2a278343223196413217914b6f3bba784a34fa47937f58a2eb18ffa09