Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:18:52.564559Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2411.12593.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:18:52.564559Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 95410f75-ccf0-4bc0-a682-563d866b901d · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3d9d42e9-f402-4e97-bc93-0c5f8c977a79 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Visual question an- swering, 2015
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d061cf7e-d4fa-4cf3-b2d0-a8356d2579db · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Flamingo: a visual language model for few-shot learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b99d0214-6897-4263-8098-8ec517658b96 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55fec8ff-e7a7-42be-a98c-2e3885118bba · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5716e1-3e86-4639-b831-6149ab187301 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Chen and William B
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eb0fdabb-c8b5-4f57-abb8-1920dce53d98 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Videollm-online: Online video large language model for streaming video
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6013e716-7fca-49dd-bb7c-1867de592ef4 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Gonzalez, Ion Stoica, and Eric P
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 66ad7158-bae8-4ee9-bb25-68dc384440cf · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Rethinking Attention with Performers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f41f1ff-e22a-43d4-9a83-7a60f7d4d602 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a505b763-7b73-4d92-8951-1ce2e913d1ba · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6689eded-43e4-4b39-9861-ebc55cf1cfcf · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Visual question answering: A survey on techniques and common trends in recent literature, 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09b7155f-029e-4c6e-8626-b2e0e25d88b5 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Im- pact of green human resource management (ghrm) practices on organizational performance
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a153acce-fca9-4ca1-b53d-7701cea2bc61 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Model tells you what to dis- card: Adaptive KV cache compression for LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7611fc79-9f75-432a-a42f-31e24613990b · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Ego4d: Around the world in 3,000 hours of egocentric video
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57164a6-e681-43d6-a38a-5f2995b73726 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6e71ec0a-51d6-48fc-9536-b39f3988c980 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2f81c3c2-6582-4a9f-aa3b-dab495207558 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 94794bb2-b37c-42a2-917c-c7035592f231 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Long Movie Clip Classification with State-Space Video Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 91410906-4c72-4183-9d38-0c265dcd6ed8 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Show, Ask, Attend, and Answer: A Strong Baseline For Visual Question Answering
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f741578-4054-4cd9-8acf-f7168886f12f · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Kuehne, A
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fb7e2a55-122b-424a-afe4-bb892213d4bb · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b12c74e-bce0-4c6e-9be8-3a0251bff48f · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 66babe02-ab24-4c5d-b0fb-fce9938a2204 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Swinbert: End-to-end transformers with sparse attention for video cap- tioning, 2022
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d31e14a-4413-4cfe-8992-22ece5de9a18 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Learning to recognize procedural activities with distant supervision
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aea4ce7f-9553-4caa-a556-7dc4e9212327 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Minicache: Kv cache com- pression in depth dimension for large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 65910ce0-ec8e-4968-afa0-be4568b96b69 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Vil- bert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ecee83cb-8a8d-4eeb-81ca-8fa4ceac7922 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Univl: A unified video and language pre-training model for multimodal understanding and generation, 2020
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 075d7e8d-422a-45b4-a6c6-f24acd4974da · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7492252c-79da-4045-be69-edd961913d41 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction GPT-4 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8217b22-e0ba-4cb3-ae3d-0647861bbe59 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Training language models to follow instructions with human feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72fea359-d661-4c8d-ab36-520cfa98c52d · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Keeping your eye on the ball: Tra- jectory attention in video transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56c991cf-ec6c-4fcb-8895-cd15d0bd106f · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4133aced-a1b2-492c-b9e7-5279d56d9b51 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Language models are unsuper- vised multitask learners
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7b98fb64-484c-49ba-9865-0d3af1ea7b2b · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Moviechat: From dense token to sparse memory for long video understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b4c4f3-f47a-4e99-aecd-871d1b2de118 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e5e8fc-7a81-4004-ba67-7218589bec66 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Plummer, Bryan Russell, and Kate Saenko
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0b872129-f3a2-498d-8e96-d1a7efca6d50 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Coin: A large-scale dataset for comprehensive instructional video analysis
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2a97bc42-4581-462b-bfed-7a8a149d692e · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Llama: Open and efficient foundation lan- guage models, 2023
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 894dcfd0-3ec2-4ef6-9207-e8d736315a54 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction CIDEr: Consensus-based Image Description Evaluation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de244629-3e21-438e-baa8-6cba81bd515a · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Git: A generative image-to-text transformer for vision and language, 2022
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7714a6ad-3b9d-4ccc-8727-cf6881322b8d · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Selective structured state-spaces for long-form video understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7f0e9205-dc34-4621-ab15-816f2b720674 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Selective structured state-spaces for long-form video understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d4e1cc1-1f14-423f-b893-7fca84a06a2a · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Temporal segment networks for action recognition in videos, 2017
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7ac3aaf5-c131-4121-81f9-b70d974197cd · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Towards Long- Form Video Understanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 03a8e9b9-470b-4925-95b9-42734523eb9f · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Efficient streaming language models with attention sinks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5db191bf-582b-42a4-aae7-78b4d46c826c · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction mplug-2: A modularized multi-modal foundation model across text, image and video, 2023
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c07de36d-fa67-4162-94ed-11b3435b0afe · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Msr-vtt: A large video description dataset for bridging video and language
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7e9436db-1f67-4dac-b7ed-f23d50b24f5d · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Videococa: Video- text modeling with zero-shot transfer from contrastive cap- tioners, 2023
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 407d2704-b126-4b9a-b957-b6880cc39b26 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Stacked attention networks for image question answering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2857f9ec-ec6f-45bd-953d-08344b617ec9 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Scaling vision transformers, 2022
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0333f01d-dfe9-4fb5-89eb-902d1b851174 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a3e1a175-0111-4a5d-83e1-c1902a5a5706 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73a199a-10d1-476b-ac92-6badf7163f57 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fd4519f6-10fb-4168-bcd2-244b465266f9 · outbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Towards automatic learning of procedures from web instructional videos
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
No inbound Pith citation observations are available.