Pith. sign in

Paper Citation Record · LEDGER

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction

As of 15 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2411.12593.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12593 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:18:52.564559Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95410f75-ccf0-4bc0-a682-563d866b901d · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Keyformer: Kv cache reduction through key tokens selection for efficient generative inference

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.279414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.328144Z digest=sha256:3d24513b4a297c2976cbf8b8a3ddb3d364a379729611f2cc664f20a4e67ca036

Observation 3d9d42e9-f402-4e97-bc93-0c5f8c977a79 · outbound

This paper cites Visual question an- swering, 2015.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Visual question an- swering, 2015

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.267667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.333521Z digest=sha256:3be1c1008b34e7634a0e083dd05ee3fc1abbea01380287fd33fb71ad4cbe2662

Observation d061cf7e-d4fa-4cf3-b2d0-a8356d2579db · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.337833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.337833Z digest=sha256:7c1e760eb61374f1072e3ee555508daae6cc4d2095d748f774b86686eb25b810

Observation b99d0214-6897-4263-8098-8ec517658b96 · outbound

This paper cites METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.343027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.343027Z digest=sha256:1773cb9fbb54d45c21ec2635c2efd1b6edda6a6af18cdad173ad514baa378b1e

Observation 55fec8ff-e7a7-42be-a98c-2e3885118bba · outbound

This paper cites Language Models are Few-Shot Learners.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.347320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.347320Z digest=sha256:d03d9d1bee25a68e70e798465c8daf73ab3e85762e0dd57d1eda5bfa91732f02

Observation dd5716e1-3e86-4639-b831-6149ab187301 · outbound

This paper cites Chen and William B.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Chen and William B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.241155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.353574Z digest=sha256:cde2744b194e441c07b4a80da0c6c382abdec2c91860291485f53c6a17033d21

Observation eb0fdabb-c8b5-4f57-abb8-1920dce53d98 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Videollm-online: Online video large language model for streaming video

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.226887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.357940Z digest=sha256:6bb9f164646b7bebe8d76fe511b1dfed95d4cb5a8f31128825379dae5118854a

Observation 6013e716-7fca-49dd-bb7c-1867de592ef4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.213152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.362071Z digest=sha256:8df3aaa32c06540c691f7df2203e29f7991de81e0c2ffc10014530b4227eedb5

Observation 66ad7158-bae8-4ee9-bb25-68dc384440cf · outbound

This paper cites Rethinking Attention with Performers.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Rethinking Attention with Performers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.366010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.366010Z digest=sha256:22e50b43b57bc3e643971de9e1d6dc3ab98dd114aef3bad18d0fd65b178fc87e

Observation 1f41f1ff-e22a-43d4-9a83-7a60f7d4d602 · outbound

This paper cites Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.198912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.371189Z digest=sha256:faeccc3fbdc9379fbf5e260ba923e1c752a1a69a20b4db388bfcb5a7552509c6

Observation a505b763-7b73-4d92-8951-1ce2e913d1ba · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.182886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.374097Z digest=sha256:165e44a048934a59932e9508b4462b28756c1e6a15b5f4583d5bc9a3732b979e

Observation 6689eded-43e4-4b39-9861-ebc55cf1cfcf · outbound

This paper cites Visual question answering: A survey on techniques and common trends in recent literature, 2023.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Visual question answering: A survey on techniques and common trends in recent literature, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.170278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.377614Z digest=sha256:6069a1224d9bf7a21d3f29aba341fe4feed20c70dfdd1fea9aee34dc0c424343

Observation 09b7155f-029e-4c6e-8626-b2e0e25d88b5 · outbound

This paper cites Im- pact of green human resource management (ghrm) practices on organizational performance.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Im- pact of green human resource management (ghrm) practices on organizational performance

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.157152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.381038Z digest=sha256:5f23808cf2fec60f275ceddee41ef4ea40e6dfe8e09097d0e4d4e13b773982e4

Observation a153acce-fca9-4ca1-b53d-7701cea2bc61 · outbound

This paper cites Model tells you what to dis- card: Adaptive KV cache compression for LLMs.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Model tells you what to dis- card: Adaptive KV cache compression for LLMs

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.145790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.384972Z digest=sha256:159c57050c152a3dd553554c5a1b7f9e5efc7c79a1714213100179828946f3de

Observation 7611fc79-9f75-432a-a42f-31e24613990b · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Ego4d: Around the world in 3,000 hours of egocentric video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.389678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.389678Z digest=sha256:abc5f4d2bbcc7dbe0236c125c36f2cff201aa4c0db5f87ffa2a4ad87e8f53c04

Observation b57164a6-e681-43d6-a38a-5f2995b73726 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.127218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.393366Z digest=sha256:c563cb3fa412828eef36ab89c3cb8e1c2d099f35118e61af63bd33190ab2d2c3

Observation 6e71ec0a-51d6-48fc-9536-b39f3988c980 · outbound

This paper cites an unresolved cited work.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:18:53.114321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.397069Z digest=sha256:13c18742c62bc79e53eac4a6edf36268d7982e03edc66213bf2e7a67d4c45396

Observation 2f81c3c2-6582-4a9f-aa3b-dab495207558 · outbound

This paper cites an unresolved cited work.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:18:53.100621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.401642Z digest=sha256:1c4eaf81a5cb8da06e45b5988a622b78a8c99597bd9f45742872f5fae13f6acc

Observation 94794bb2-b37c-42a2-917c-c7035592f231 · outbound

This paper cites Long Movie Clip Classification with State-Space Video Models.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Long Movie Clip Classification with State-Space Video Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:18:52.696425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.406219Z digest=sha256:e6dd4ff758365e8407e822842221ec501cf4b65050c5d9162b2ca4f4206a03a9

Observation 91410906-4c72-4183-9d38-0c265dcd6ed8 · outbound

This paper cites Show, Ask, Attend, and Answer: A Strong Baseline For Visual Question Answering.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Show, Ask, Attend, and Answer: A Strong Baseline For Visual Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.412135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.412135Z digest=sha256:524a5685ec5a7dca1c493e72ace431db33a490a6db8979de09903a041d313a97

Observation 3f741578-4054-4cd9-8acf-f7168886f12f · outbound

This paper cites Kuehne, A.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Kuehne, A

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.085920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.417858Z digest=sha256:6cdb624c9b60a6fc4e75081e6ac395601b1cdebe267344f71c41b539bde679b7

Observation fb7e2a55-122b-424a-afe4-bb892213d4bb · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.422734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.422734Z digest=sha256:1225e69ac9149069b81a33620bd0292ab318eba503aab389a35feda0c7488d78

Observation 2b12c74e-bce0-4c6e-9be8-3a0251bff48f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.068418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.427704Z digest=sha256:511e8d389ff3ec48066f7aa9a80479c7a76b405b77baae54f326a136fa1a47b4

Observation 66babe02-ab24-4c5d-b0fb-fce9938a2204 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video cap- tioning, 2022.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Swinbert: End-to-end transformers with sparse attention for video cap- tioning, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.055301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.432267Z digest=sha256:76baa15e90d5521d3af2260ae157dddf8ca2bfcfe2bcff3d909b948dd98510d1

Observation 6d31e14a-4413-4cfe-8992-22ece5de9a18 · outbound

This paper cites Learning to recognize procedural activities with distant supervision.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Learning to recognize procedural activities with distant supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.040153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.437823Z digest=sha256:ae93cad14900927f9034e353cfab59109d6e3e05a75b6d3c159b7f4f745d31d3

Observation aea4ce7f-9553-4caa-a556-7dc4e9212327 · outbound

This paper cites Minicache: Kv cache com- pression in depth dimension for large language models.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Minicache: Kv cache com- pression in depth dimension for large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.027625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.442787Z digest=sha256:6185e91009d3ca5ebca39e595df49ad6a2a4490415d3238c1f061a2190e75056

Observation 65910ce0-ec8e-4968-afa0-be4568b96b69 · outbound

This paper cites Vil- bert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Vil- bert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:53.013773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.446307Z digest=sha256:ffd58e8e1c2f3067782712cd8e734e4ce44893c75d865a971fdbc0a450c4080f

Observation ecee83cb-8a8d-4eeb-81ca-8fa4ceac7922 · outbound

This paper cites Univl: A unified video and language pre-training model for multimodal understanding and generation, 2020.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Univl: A unified video and language pre-training model for multimodal understanding and generation, 2020

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.997496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.450111Z digest=sha256:71e90a2c73d549c3d80dc422575fe08db8501c2b9345536f8c153bc76fe9481a

Observation 075d7e8d-422a-45b4-a6c6-f24acd4974da · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.981415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.453664Z digest=sha256:29c7036f2a8ce65d4547dc3b557866e7702599fe6a2b90aef48c2f7fae11f34e

Observation 7492252c-79da-4045-be69-edd961913d41 · outbound

This paper cites GPT-4 Technical Report.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.460526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.460526Z digest=sha256:690917644801277e7b17f97bbd58b491a3d12943700be2f2f2ea03e23e1761d4

Observation a8217b22-e0ba-4cb3-ae3d-0647861bbe59 · outbound

This paper cites Training language models to follow instructions with human feedback.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Training language models to follow instructions with human feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.464083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.464083Z digest=sha256:4f86b3fe24f5d2c077c0a4f10ea6da57b6c36c83ecf6b751ab0778701b6b1b26

Observation 72fea359-d661-4c8d-ab36-520cfa98c52d · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Keeping your eye on the ball: Tra- jectory attention in video transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.469522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.469522Z digest=sha256:16944e5dcdb498b1fcd6a227c46d214673c6404dc5aaeb71470f15bed1729fd8

Observation 56c991cf-ec6c-4fcb-8895-cd15d0bd106f · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.473471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.473471Z digest=sha256:0ef161d06be665a2ac8882fb095a4ce8a7813b2eb27928052be6cd9ffcd55f1a

Observation 4133aced-a1b2-492c-b9e7-5279d56d9b51 · outbound

This paper cites Language models are unsuper- vised multitask learners.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Language models are unsuper- vised multitask learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.963845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.478290Z digest=sha256:1d0f4036ce444103f9cc038fb40e8510077c1e5302fd46e7f7f75b6bf5adc4a0

Observation 7b98fb64-484c-49ba-9865-0d3af1ea7b2b · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Moviechat: From dense token to sparse memory for long video understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.483532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.483532Z digest=sha256:b487286e914d32f87ec4517661d80d339d6c8f1c51eb44443791df18fe19a53c

Observation a3b4c4f3-f47a-4e99-aecd-871d1b2de118 · outbound

This paper cites VideoBERT: A Joint Model for Video and Language Representation Learning.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction VideoBERT: A Joint Model for Video and Language Representation Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.489724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.489724Z digest=sha256:b9656e2104bb384f54d9274d25090a24667b9764fe328be7ef0c9661e9bac0c4

Observation c3e5e8fc-7a81-4004-ba67-7218589bec66 · outbound

This paper cites Plummer, Bryan Russell, and Kate Saenko.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Plummer, Bryan Russell, and Kate Saenko

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.946562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.494322Z digest=sha256:709fc727da0a868d9dbba698fbee03f3dc81b94411786c85107f079bc9eec0c5

Observation 0b872129-f3a2-498d-8e96-d1a7efca6d50 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.932846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.498051Z digest=sha256:aa92138758e8f185457efd42bbaeb5168d3f2fa8a37010185a871104d3acca5d

Observation 2a97bc42-4581-462b-bfed-7a8a149d692e · outbound

This paper cites Llama: Open and efficient foundation lan- guage models, 2023.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Llama: Open and efficient foundation lan- guage models, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.919833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.502098Z digest=sha256:b42c512db5af346b057fc32c7588f29e099dc5125384f46da695dfe8aa654f91

Observation 894dcfd0-3ec2-4ef6-9207-e8d736315a54 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction CIDEr: Consensus-based Image Description Evaluation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.506611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.506611Z digest=sha256:1901460cb568ca575bc14ef45cc6d2bb1b4f8ce09096cfc6343edf61cbe0508f

Observation de244629-3e21-438e-baa8-6cba81bd515a · outbound

This paper cites Git: A generative image-to-text transformer for vision and language, 2022.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Git: A generative image-to-text transformer for vision and language, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.909115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.512572Z digest=sha256:8f952ebf0abd13bda1d5d3dbce16219a030631285091df3f86cdf0912a4ebcfa

Observation 7714a6ad-3b9d-4ccc-8727-cf6881322b8d · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Selective structured state-spaces for long-form video understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.897742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.516806Z digest=sha256:72e9e2de9ce9b0909e7a647ea373040d1681a7288f3c26a851ff68faace1eb1b

Observation 7f0e9205-dc34-4621-ab15-816f2b720674 · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Selective structured state-spaces for long-form video understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.520356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.520356Z digest=sha256:35238db595b7a279cc1e43a6c52defe9fd6a83757a141da8df8321033d7f5088

Observation 1d4e1cc1-1f14-423f-b893-7fca84a06a2a · outbound

This paper cites Temporal segment networks for action recognition in videos, 2017.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Temporal segment networks for action recognition in videos, 2017

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.876048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.525464Z digest=sha256:5827155a951008f57327a3330fae78e49e8762e342406fc7658bdc9dfa4d813e

Observation 7ac3aaf5-c131-4121-81f9-b70d974197cd · outbound

This paper cites Towards Long- Form Video Understanding.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Towards Long- Form Video Understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.863050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.529834Z digest=sha256:c80a78e58b1f684e3cbc72e2b7f354c8670ef1e06be4a47dada2bd9e83b34ff1

Observation 03a8e9b9-470b-4925-95b9-42734523eb9f · outbound

This paper cites Efficient streaming language models with attention sinks.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Efficient streaming language models with attention sinks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.852767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.533263Z digest=sha256:b857d7394d485d7165a74b757acf3fff52a9e7c97d269e8e5756e7b38a9e43ff

Observation 5db191bf-582b-42a4-aae7-78b4d46c826c · outbound

This paper cites mplug-2: A modularized multi-modal foundation model across text, image and video, 2023.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction mplug-2: A modularized multi-modal foundation model across text, image and video, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.841104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.537052Z digest=sha256:74ee359756fc648a60319149911e03c9b202cb7b6e06e0153cbe6478056c87fb

Observation c07de36d-fa67-4162-94ed-11b3435b0afe · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Msr-vtt: A large video description dataset for bridging video and language

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.829041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.540635Z digest=sha256:920fb66e17f104e8b872686a46fb9eb1fec45833b02fcc054e1bfe3652ca393e

Observation 7e9436db-1f67-4dac-b7ed-f23d50b24f5d · outbound

This paper cites Videococa: Video- text modeling with zero-shot transfer from contrastive cap- tioners, 2023.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Videococa: Video- text modeling with zero-shot transfer from contrastive cap- tioners, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.816592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.544018Z digest=sha256:baf5a1fb5a9b1f5759721b97525c9978702f90029ae2c97cb7b33f27e7511767

Observation 407d2704-b126-4b9a-b957-b6880cc39b26 · outbound

This paper cites Stacked attention networks for image question answering.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Stacked attention networks for image question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.804610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.547004Z digest=sha256:a47dbda2790525adfba02a10a9d4e9c50ae4303c3e352d914c67a968903bc828

Observation 2857f9ec-ec6f-45bd-953d-08344b617ec9 · outbound

This paper cites Scaling vision transformers, 2022.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Scaling vision transformers, 2022

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.793904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.549987Z digest=sha256:023df6dbded204e68c8eee73aae6f5da4e0f27994d472d7c4e0c45be9105c3e8

Observation 0333f01d-dfe9-4fb5-89eb-902d1b851174 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.780995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.553106Z digest=sha256:87ac05f4cd0b774ba32fcd811c8dc2b78015a657b162d02a2c32a6d1d7ad4735

Observation a3e1a175-0111-4a5d-83e1-c1902a5a5706 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:52.556066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:52.556066Z digest=sha256:0b22715e513d1dae457839ae5810950774933fb523e1cb78674f34f532b1ba0c

Observation a73a199a-10d1-476b-ac92-6badf7163f57 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.767157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.559752Z digest=sha256:25d35d42a30561e41d8906cb60a7d4e0ea103590e829506bf60c659a9816fbee

Observation fd4519f6-10fb-4168-bcd2-244b465266f9 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction Towards automatic learning of procedures from web instructional videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:18:52.752733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:18:52.564559Z digest=sha256:5f4a9c4168972b9d042837d78acac124ea89dbbdf0abc0d23fde6f0a3f1571fd

Pith citing papers

No inbound Pith citation observations are available.