Pith. sign in

Paper Citation Record · LEDGER

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

As of 9 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 3 inbound Pith citation observations for arXiv:2505.19602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19602 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:20.927090Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:03:03.062322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T07:58:07.594523Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved75
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 293a0609-f139-43f6-8037-2b43a274623d · outbound

This paper cites Longformer: The Long-Document Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Longformer: The Long-Document Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.031316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.031316Z digest=sha256:934131c344e779b0f1c0f27df8dffcc03c73312cf30ca4289cd421f8e7ce7b52

Observation a6f5049a-cb1c-4012-92fc-54079f5ef1b7 · outbound

This paper cites Improving image generation with better captions.Computer Science.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Improving image generation with better captions.Computer Science

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.089620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.089620Z digest=sha256:ac3bbd4ca336540a3202fc0f3596d7c460d08322dbbc27a375b102b588cb4dc2

Observation dffea209-3858-47ca-a9a0-3affff8ba103 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.217814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.217814Z digest=sha256:8016d9d965b66b4a8b13d7f720ff42406e35ab877ab0f89d70e26e2c1fee886e

Observation e9ea3931-fa9a-4d77-8236-6c48319287f0 · outbound

This paper cites Maskgit: Masked generative image transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Maskgit: Masked generative image transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.340496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.340496Z digest=sha256:ddf81ebc51ca912ecb9b50af6c842fd8af4899f01484d96a623f7a28cd57a2d4

Observation 646555be-1f9c-4df6-9c80-d2e0bfdc8a28 · outbound

This paper cites PixArt-Sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PixArt-Sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:24.117531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:16:12.414680Z digest=sha256:67326476acadb06fda8cd2d43ab02695a428190f35dd337362185aec41991697

Observation a14d7039-3002-422e-b24d-f6c3e6c903e8 · outbound

This paper cites Generative pretraining from pixels.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Generative pretraining from pixels

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.543575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.543575Z digest=sha256:e9a0c45d5f7e3aa314516f9eab0cb86a3bdc6879a26ccc5429adb98ddb3d6388

Observation aa1f79b0-8e6d-406e-8dfa-1e8a18d509a4 · outbound

This paper cites Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.670870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.670870Z digest=sha256:e1cfe22e4ed6f35e6d5b5a57457405cddd06e3349450f06c3e2e21450ae8d122

Observation 0fad3e3f-9146-4a5f-beae-d73abe0a52dc · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Taming transformers for high-resolution image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.835884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.835884Z digest=sha256:f842960ec485ab68a7d439676f4682d3890815e0325b4fab80f5a5585b0d2f08

Observation 0a678a28-5f88-44c4-8487-7a8823fc6d0f · outbound

This paper cites TinyFusion: Diffusion Transformers Learned Shallow.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression TinyFusion: Diffusion Transformers Learned Shallow

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.014126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.014126Z digest=sha256:4d0aed7c2def9eff6b6727546aa6da8c6691701dd93790ba3e18e6cd2264d385

Observation 2e6847e7-cbbb-4746-be02-7c97c4414363 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.138387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.138387Z digest=sha256:bcfb963d46334290354865f803fd7050859020e73d18bcd6fbfa273fc131b55e

Observation 8fd9a677-964d-498e-8609-a9f486163db0 · outbound

This paper cites Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.258212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.258212Z digest=sha256:b3bf912636f107d0bdb474f6cd6b4dd5764d208ee51ab178c1988977f43b8e13

Observation 814de7f5-fdf4-4a33-9933-76925e214302 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.413395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.413395Z digest=sha256:37835e58e45e678b75ee50e6e9d87676614b2f09aa0aadd89e8b56191105cac6

Observation dc481280-ccd0-491c-af0e-50b4974a2be8 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.862112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:16:13.530092Z digest=sha256:ed554cef471f97de120f9eaf6ae4105f9acfb4d86ea630d7abaa7cfd644be85e

Observation 22f15021-513d-40bb-a2ad-7b7af0fd1226 · outbound

This paper cites DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.622854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.622854Z digest=sha256:9a490bc74a0185d04a7b8640267c2407539485adf5f46b82914305fea106affb

Observation 0d51690e-64bf-4e93-9317-19ab2db69ce6 · outbound

This paper cites FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.693484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.693484Z digest=sha256:8db597cb8f8bf55f541e0b344ef32ad9e47503fff0eba485c9dfa505ac800c9b

Observation c8144f2e-57d3-4053-bc65-a0aff67de974 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.770499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.770499Z digest=sha256:5d386d53751062be8a6e516e0010da47999e456cc038a1226713777fe0505c34

Observation e9cccbfa-a369-4d8f-bf85-2dddd7291d5f · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.857590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.857590Z digest=sha256:e465a07c2243cb5a5b70c61b61b8b6b6a8a5a841aa7b1c9bd4f9f7d6aa6f4af1

Observation 90c1a4be-0e2a-4a36-8855-c9899a98340c · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.948063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.948063Z digest=sha256:52d69d1e4a4ad0111423a725e1c196de5a3f00b55a059ae91ae5b753f924b36d

Observation b128b49e-cda3-4642-9df2-5a338a5297d4 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Distilling the Knowledge in a Neural Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.033748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.033748Z digest=sha256:a197832959eba4ea6a25f5dd33b2cb17fddd11aa4aa2e8587ac6944aa3fd44b8

Observation a26ea0ad-6364-4f16-92b9-5d354a338c16 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.124213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.124213Z digest=sha256:ef3db5962daff06fe8a93ece3ebf26a96bd4824a3a0978033251db8f8097ecaa

Observation e4045744-1d75-4e87-b00d-b7d213d10524 · outbound

This paper cites GPT-4o System Card.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.194123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.194123Z digest=sha256:351ac797332c04a081a0ab0fa06c8d3b73384b63c0908f5c328963e873a7ce39

Observation 38eff18d-581c-469e-8f75-d9b88ac6bfaf · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.282281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.282281Z digest=sha256:4c0c14ab55348d6c44854f1190c4f7a1622ec723f2708bb866eb662bdd4beef2

Observation c1402ed7-f490-4546-badd-4fd7f608e0b5 · outbound

This paper cites Scalable autoregressive image generation with mamba.arXiv preprint arXiv:2408.12245, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Scalable autoregressive image generation with mamba.arXiv preprint arXiv:2408.12245, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.444559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.444559Z digest=sha256:dfe949ba72d10e8e7b7e68af1f37bbd4fd4e97227545cf4ddcb59a03183e50df

Observation 2b8bf013-ee0a-40cf-bd57-999abc287a7f · outbound

This paper cites Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models.arXiv preprint arXiv:2411.05007, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models.arXiv preprint arXiv:2411.05007, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.611793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.611793Z digest=sha256:ce2e4717a957e0482a61dc3aca09f9c525c8ac295b4b81d62beb7e76a97fb39c

Observation 06423f51-cfd0-4ae9-88ef-0b4c16487a02 · outbound

This paper cites Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.731278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.731278Z digest=sha256:079f9ce75d258d8efd9a80bb220dc70e6fcbc293462bc2ddf8b5eacd32adfd43

Observation cb8defb5-873d-46e9-9eb1-3f9540d81714 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Autoregressive Image Generation without Vector Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.884901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.884901Z digest=sha256:1e6ae975b64ee5aa9a986cd0a9fe03bf382a9fb6908dc2bf460c5854e4fa4e1a

Observation 62903964-d2cc-40eb-8e34-f2ca28e166b2 · outbound

This paper cites ControlVAR: Exploring Controllable Visual Autoregressive Modeling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ControlVAR: Exploring Controllable Visual Autoregressive Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.022702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.022702Z digest=sha256:8feefbe19fd0d07f734a8aeb31cdedbe796589c2f1550c22d258b655e3bf80d1

Observation 4026be12-2d5f-44b1-a227-95b768acf7cd · outbound

This paper cites Q-diffusion: Quantizing diffusion models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Q-diffusion: Quantizing diffusion models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.125418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.125418Z digest=sha256:94c8f5ffb3d437e7c7380747d632411d5087d63d0706f3ba012e8f4ffe6004d1

Observation 29d6db59-5830-46f1-a198-ef87d091c6e3 · outbound

This paper cites Snapfusion: Text-to-image diffusion model on mobile devices within two seconds.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Snapfusion: Text-to-image diffusion model on mobile devices within two seconds.Advances in Neural Information Processing Systems, 36, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.563170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:16:15.281994Z digest=sha256:6fec259016934555ece99bfc99934bfc074fdd0d0ee3bff49de370c9f4fcf56b

Observation 8e0543e4-440f-4b44-b65f-4a42fa18ac23 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression SnapKV: LLM Knows What You are Looking for Before Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.378042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.378042Z digest=sha256:bb08eab9fae22027d676513ecb038e017813839de9934057ed3e8d94fbb0a79d

Observation 6fd371e8-a244-47b5-b834-3ef570d38900 · outbound

This paper cites ControlAR: Controllable Image Generation with Autoregressive Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ControlAR: Controllable Image Generation with Autoregressive Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.523883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.523883Z digest=sha256:7205f98bb9080d6e1e8b0b50251aa530734253b93f03fccc7cc2df0170e62554

Observation a495cb52-b5cc-4ba3-9c05-bbf931b8765b · outbound

This paper cites Microsoft coco: Common objects in context.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Microsoft coco: Common objects in context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.657471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.657471Z digest=sha256:82ab2f451a6194f1da40157c1ec8c09fede20fee5408aa7432cd767d42d9030c

Observation 46aec037-8f7c-40db-bf10-3ec3a2194891 · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.805734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.805734Z digest=sha256:fa29b28785befed5da5f59d9d084ceb2af413e07f298710926eff050a4de303e

Observation 1adb9af2-121c-41eb-81d2-c622f8fed257 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.891755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.891755Z digest=sha256:92e4a78ee1ec2d9259dbc106d20e0762798f959e7d3115b1861c58cb990985ec

Observation bebf9bab-5a0c-4ab4-b687-370c985bcc0a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.Advances in Neural Information Processing Systems, 36, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.016184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.016184Z digest=sha256:0fdd4550c82d2911e0ca390e87ebdd1047e9f816c2c6b7027b9fc7cda367adb3

Observation 6c237317-f6f2-4a3b-8185-92f5aef59083 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.195264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.195264Z digest=sha256:537143a7730c059553cf62757e7ade646b414aad363300f123cafb08d3a7c377

Observation f0675899-2855-42d8-8823-4a6794e46234 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.358206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.358206Z digest=sha256:d2131a409ef2373092e77e916ac228a3f9366ed0e185e73e9d653bb4d19f1df8

Observation d5edde25-459b-4fdb-8432-9e94b657d3a9 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.536400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.536400Z digest=sha256:b050861b09ac060956a4c6ccb29b0617c9085980a11fa98831f623a47100b7b3

Observation 6133ddb2-fc52-429d-8bbf-4dcf6ebadf01 · outbound

This paper cites Accelerating Diffusion Models via Early Stop of the Diffusion Process.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Accelerating Diffusion Models via Early Stop of the Diffusion Process

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.665647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.665647Z digest=sha256:7f9a372062f0de517c27ef7f63481af0de3c41ffc5147e4e5ba45a54933710c0

Observation 4eb1f813-b788-4468-b191-0668b04482d3 · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.780411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.780411Z digest=sha256:0d6dda11ad561673ebc7663c643fa1213eac77eaf783ead1efb118cf831988f5

Observation 2eec77e1-4b91-470e-85ac-3b668055466b · outbound

This paper cites DeepCache: Accelerating Diffusion Models for Free.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression DeepCache: Accelerating Diffusion Models for Free

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.890592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.890592Z digest=sha256:203726de68c9500189ecc8449520a73bc06b6f1cc38d95ef6d8aed3729eb19db

Observation c33b358f-77c5-4223-bb69-6eaff1831127 · outbound

This paper cites Transformers are Multi-State RNNs.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Transformers are Multi-State RNNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.009571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.009571Z digest=sha256:d20aa138ea4122128f42d504a98084ff85c09272255104bb4e01619cac47160c

Observation 3c17afa5-d337-45c2-9e6c-6c4729a30143 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.118171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.118171Z digest=sha256:91390845005dde0bced8d2a2a72b90de3603d59a705f4b54e06e4c7d9f2167e9

Observation ba685361-ef66-4fa7-89e4-8b276a81e4f1 · outbound

This paper cites Head-aware kv cache compression for efficient visual autoregressive modeling.arXiv preprint arXiv:2504.09261, 2025.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Head-aware kv cache compression for efficient visual autoregressive modeling.arXiv preprint arXiv:2504.09261, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.201417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.201417Z digest=sha256:781ec443eacd894d16ba9d8e7a7fad3abe18bc869c5d71eedf1d68cbd15dd0cf

Observation e9c6dae5-05e4-4040-a6e6-e81f9537114c · outbound

This paper cites Efficient Autoregressive Audio Modeling via Next-Scale Prediction.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Efficient Autoregressive Audio Modeling via Next-Scale Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.266768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.266768Z digest=sha256:0a4d024e3a121e85895d7d17e06b422bf5c2768bd0e5b4e0bea0194877948165

Observation c21181a6-5cea-470f-a97a-05134b04e7ca · outbound

This paper cites On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.384573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.384573Z digest=sha256:df8c05a2524bfa6cc3094ec355846d91c6385872e8d516b56380b4b16aebaa3d

Observation 4a14f778-46c6-4d1f-a4ec-dcbf3ab8e720 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Progressive Distillation for Fast Sampling of Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.490839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.490839Z digest=sha256:60b6efee12f078b33d40d41f69bf8ec65b257add31cc182b881935943cd5c6c5

Observation aa9df552-5935-49c0-8aa1-7d1032654f6d · outbound

This paper cites Adversarial Diffusion Distillation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Adversarial Diffusion Distillation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.578156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.578156Z digest=sha256:354de92931b301700aa82086aa943402531649fea8365de1046df0aa71743337

Observation 3d0a08cb-ef61-469a-b213-c19bf6a9ab4c · outbound

This paper cites Post-training quantization on diffusion models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Post-training quantization on diffusion models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.679607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.679607Z digest=sha256:1a6138dc1cde3809f6b2e6389d7e792b934725e1ab4bccafb7d767d2ee1fe06a

Observation a8693224-84aa-45a6-b7c3-dfc107049171 · outbound

This paper cites FRDiff : Feature Reuse for Universal Training-free Acceleration of Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression FRDiff : Feature Reuse for Universal Training-free Acceleration of Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.772212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.772212Z digest=sha256:5d7f1855b08008da41186297cb6e80fcf70d72a17b1010d38e81368c68f8333d

Observation 353a2226-eed9-4e36-903e-a49527443e35 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.009369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.009369Z digest=sha256:1580777234761bc3be563bca4f8fd4be0969716ddf0df2fcc4b63b811bf70ff1

Observation a8d75859-c32a-486a-af2c-e1c44e89e10f · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.127539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.127539Z digest=sha256:255a2c57a893001f3c3fd19df568c19a6baf342bd7e8f2a124c62eddbb182891

Observation 4080b004-8286-4e92-9045-62ccd4f9bc2c · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.201368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.201368Z digest=sha256:4b3aef7d4e785708ff2b5048dd231623942483075dc09ebaa24d7cd39aba553e

Observation 456bd8c8-8727-471a-935e-d761f58b42f2 · outbound

This paper cites Conditional image generation with pixelcnn decoders.Advances in neural information processing systems, 29, 2016.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Conditional image generation with pixelcnn decoders.Advances in neural information processing systems, 29, 2016

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.298675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.298675Z digest=sha256:b61c782d82612747c7a186fed4396284e7ec7f54d5ae4face60ab4ce80218f28

Observation 1cb61cd6-bea0-49cc-9630-78d1767074d0 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.412301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.412301Z digest=sha256:d7a9262546e7a2058576aa560a997d4607d1947171e7656964d22d5fac378920

Observation 735fab92-1f44-4700-825d-a90d83fc65a6 · outbound

This paper cites D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.517957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.517957Z digest=sha256:072d065492e794222e4a66191518be53e4caf6e17fc56dcf3caa4cb39d6bebf4

Observation 1cabc8b2-404d-4115-8940-3fe58a98ed21 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.606283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.606283Z digest=sha256:476047c55700e59563e28e790e6b8b1c31a9ecddcef054d017171bd3676a8e8c

Observation aa87dea9-225f-47b8-8879-38d93ae314cb · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Emu3: Next-Token Prediction is All You Need

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.699512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.699512Z digest=sha256:e8237f860d236126ac9b11c16e1782ceba030394d78cb93a928a5125ed23391a

Observation 975c5e03-be46-46c4-8191-fea0a7ff031d · outbound

This paper cites Cache Me if You Can: Accelerating Diffusion Models through Block Caching.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Cache Me if You Can: Accelerating Diffusion Models through Block Caching

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.760250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.760250Z digest=sha256:c9f10c46c8109a90947dc7d6d95f76bb028c921bcf1c1982272f3e4f9fb883bb

Observation e9eb3f00-e6d4-4429-849e-e68f1db1311f · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.842698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.842698Z digest=sha256:aeb27f4d1f3e5438803bc3c4d4bc2485876fd071aec57410dca10eab1b8c9db0

Observation 545bebb5-615d-4ecf-b2f1-754fcf5878c5 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.961783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.961783Z digest=sha256:543694d6007da08360bea0e43c6907891c40f22e66d78719205a836a2d44d1e1

Observation bacfe7e9-9139-4137-abab-317b6f2257e6 · outbound

This paper cites Efficient streaming language models with attention sinks.arXiv, 2023.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Efficient streaming language models with attention sinks.arXiv, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.047855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.047855Z digest=sha256:bd74cc181e704afee6c9dde0c777d18e58cb184054efc9ca00cde7b972e0591a

Observation 8655804e-2349-42e9-9902-0e4d05f68341 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.134844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.134844Z digest=sha256:fb1319c9006d7349ad330161858f7597d1d0e0182f863c9488dfc74229df8ff7

Observation f56e95aa-f6be-4e5f-994b-a7f82c45c05c · outbound

This paper cites LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.250138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.250138Z digest=sha256:1f842091b8c5b177f1ca72702adecb6798d3b27ab803e423658cc88d64475745

Observation 827a1ae8-2935-4ac9-96d7-895870c2b1be · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.360068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.360068Z digest=sha256:2cced933469086763a229cd0d8f08723ef7b9e3df22eb942bc037bab3a43b726

Observation 8d6893d3-7a1a-463f-9bef-f80db39d814d · outbound

This paper cites Hash3D: Training-free Acceleration for 3D Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Hash3D: Training-free Acceleration for 3D Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.475008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.475008Z digest=sha256:6829d3993b4357a23249282d2a1251d8f37312f6ca5be445d668782d11750c97

Observation b7a8c97c-e8de-47ce-bb5a-d8e082b59bd0 · outbound

This paper cites Diffusion probabilistic model made slim.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Diffusion probabilistic model made slim

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.209684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:16:19.548827Z digest=sha256:2acde41b70198f5eebf05ce01bbfab776b039707ccaa9a507e6a1ef6fbc643ef

Observation f523d14b-49a3-43af-9877-e8c9330bc139 · outbound

This paper cites CAR: Controllable Autoregressive Modeling for Visual Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression CAR: Controllable Autoregressive Modeling for Visual Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.608040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.608040Z digest=sha256:654a0bad6133d038b622ac56e9cc820bc4aac6cdc82ca07d50d70503ca7d84c2

Observation 0f4c3011-d88a-411d-8bb5-7ed536fb36b2 · outbound

This paper cites One-step Diffusion with Distribution Matching Distillation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression One-step Diffusion with Distribution Matching Distillation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.701607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.701607Z digest=sha256:422404a2bd7b05425c446150b69b8bc35e7f4ad9659cd63f71fd4d89c965db91

Observation 0e775c07-916f-4896-ae0c-6090247f3f33 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.784103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.784103Z digest=sha256:6ad9babc56aa2260df5e5ec1c417198a4b84b2352036ebcb6d8f7e431eb49ea3

Observation 3557f87b-5aab-4826-a750-14232c37e679 · outbound

This paper cites Resshift: Efficient diffusion model for image super-resolution by residual shifting.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Resshift: Efficient diffusion model for image super-resolution by residual shifting.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:22.897167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:16:19.898686Z digest=sha256:7e5edfd00353d9085dadb422349233328f3e1dd8d4484c1fd8cb3fb17b43c697

Observation affe042c-ddf2-4a81-9f3d-0a82e5985e68 · outbound

This paper cites LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.004397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.004397Z digest=sha256:47c13728f7a6e3cf6dfdbeb99279d94b27a6649ef951adcc69bf3113abdda3a1

Observation aaa1b858-7fb0-4d7a-8413-b7a35de4a838 · outbound

This paper cites G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.116154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.116154Z digest=sha256:885cfed47e06dbfb3f0b1be12722a4c1e14d7a67b498e65d3e78a377311c9362

Observation 91b9fa70-7ee9-4c17-b71a-2ec91cd63095 · outbound

This paper cites VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.208413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.208413Z digest=sha256:c1f64227fc5bbbddc735a5f31e7fd155732c7e13f7e3addf31db576934fee236

Observation 9e2adf34-da4a-445f-a409-4db203c50c08 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression The unreasonable effectiveness of deep features as a perceptual metric

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.320593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.320593Z digest=sha256:abd7a9250ce3d1c444de7cb4194f136b61b14812f7d0b4d3f0c5fb917c009b6c

Observation b14eb1c8-94d5-4847-9600-356c9fdc9949 · outbound

This paper cites Faster Diffusion via Temporal Attention Decomposition.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Faster Diffusion via Temporal Attention Decomposition

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.398774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.398774Z digest=sha256:8dde1cab23ad6e0bae128652d330bc5b5de7dbf05295f6c129481f5f8949ed0f

Observation 55a8508a-c4ac-4be5-804e-d2fb45026751 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Cam: Cache merging for memory-efficient llms inference

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:22.625939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:16:20.509944Z digest=sha256:5c0e06678442ae6aa98138c0d7b20ded4af449f2e5c890e5e6b1b22643deb4c8

Observation c37d9c57-6909-4482-b493-26005cd2d039 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.627739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.627739Z digest=sha256:98bd981a37e5d96877ba29c19ef4334153e9e761599bd3c606518080d9e03cad

Observation 30e5a746-e8e1-4266-af0e-8df43ff5936d · outbound

This paper cites Real-Time Video Generation with Pyramid Attention Broadcast.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Real-Time Video Generation with Pyramid Attention Broadcast

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.706471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.706471Z digest=sha256:116bdbe5137c3a039dc8e2a4c0780c2fe6492afd1b11004da0a98c24792ef020

Observation 76094485-c0db-45ea-b5ce-5a4b36647ba1 · outbound

This paper cites MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.821528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.821528Z digest=sha256:b4c7695d6a70216a403f605dc72dc93468485f703271f5dc628418a50eae914b

Observation f972b876-9b38-4a37-98c5-2063619a429b · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.927090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.927090Z digest=sha256:b023a732a495b51b4d5072aea6172ade3ca1f227bf4c384b618721c73e9c7da6

Pith citing papers

Observation 1a3ca253-3931-42c1-bdfb-79ffa010f511 · inbound

Visual Implicit Autoregressive Modeling cites this paper.

Visual Implicit Autoregressive Modeling Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:18.325747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T15:20:03.981830Z digest=sha256:8fc5c94e4ed2fc5aca22fdd0282112aeb1922ed7c68014dc900abec110323854

Observation 348b6ba1-7a44-4b8c-91ef-076c5be59bb4 · inbound

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond cites this paper.

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:58:07.596260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:57:51.032025Z digest=sha256:70fa591ef743dabddb1023c4849b53c8192a78740e334cf643208baad81a6190

Observation 14f8186d-8235-4945-ba0f-c792e5ebda9f · inbound

Token Radius Attention for Efficient Video Generation cites this paper.

Token Radius Attention for Efficient Video Generation Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:03:03.062322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T06:03:03.062322Z digest=sha256:543072d9abfa9249df899a42f64ea1213b69084620099673e7b370bb22327c69