Pith. sign in

Paper Citation Record · LEDGER

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

As of 5 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2605.07897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07897 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T02:12:20.651111Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:21:47.718375Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact25
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ef130f4-3c4d-4368-a1d6-c2729db0baaf · outbound

This paper cites Qwen2.5-VL Technical Report.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.728343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:d73a98735ff79940ec9398e1b6dd0353947ae2b2a6dbddb726a99afd492a35ee

Observation aca26757-b3c8-4ee5-85ca-89377c2e702f · outbound

This paper cites Token merging: Your ViT but faster.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Token merging: Your ViT but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.296353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:16333cef2748f07df4bf00a3ff8e7565e361202970c29853e0ee0cafa7946d0f

Observation 11bb9116-a31f-4db5-9986-2dea5053d1e1 · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Videollm-online: Online video large language model for streaming video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.282447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:09affbba23f6e2553e9020d43dada84c7920aa9a49839ea4fa0c3a48952ab5bf

Observation e39ad353-654d-49ad-8495-0f4fb42c4eea · outbound

This paper cites End-to-end autonomous driving: Challenges and frontiers.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding End-to-end autonomous driving: Challenges and frontiers.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.275673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:f1c1afa1ecbd8c9c7668c646d7662185570318e14d9525aed54abecada06f50f

Observation 67294575-ce6b-4214-b2d7-2e3d96f52576 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.335451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:7fe302e16455bcc0b051d9b0a4457398096dfd64f9d9483d39fde6c019e10a97

Observation 73cbad4a-5406-4008-a837-973b8350428e · outbound

This paper cites Streaming video question-answering with in-context video kv-cache retrieval.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streaming video question-answering with in-context video kv-cache retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.279141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:da57061e81fe37a7b3d7db7f72b40fe1beb8006c9c405c0097bbda26e795646e

Observation 5a0e7e1d-e4e0-4693-bc82-fd3ad1ff0c3b · outbound

This paper cites Contextnav: Towards agentic multimodal in-context learning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Contextnav: Towards agentic multimodal in-context learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.682389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:5f572829d309ecc613ba3973b5aabde9ec8e17d9f56acccf534790b0fca4ea17

Observation c3c5ac3f-d520-4fac-bc78-2f09cf970c3a · outbound

This paper cites Vispeak: Visual instruction feedback in streaming videos.ICCV.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Vispeak: Visual instruction feedback in streaming videos.ICCV

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.329770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:30be69b50529fad4602ec6731f289db733929f3a39fd663d193bd7c4ee6b7adb

Observation 7c5c557c-67ee-4cfb-b736-c68cf4ec02c9 · outbound

This paper cites Framemind: Frame-interleaved chain-of-thought for video reasoning via reinforcement learning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Framemind: Frame-interleaved chain-of-thought for video reasoning via reinforcement learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.358339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:b9704fe74d3d77c9074ccfe9f9c86a5df935e3551dbe8b4aa82d8ba66aeeead0

Observation 08fafee2-43ab-4884-8e57-c6f90e5993d9 · outbound

This paper cites Online video understanding: Ovbench and videochat-online.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Online video understanding: Ovbench and videochat-online

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.321534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:bbe37bc0acd775d505c310a744c79a3f4f369e66e819dfb6761608086b5cc9ee

Observation 164185d6-5052-4626-a537-f2f29a40f0bd · outbound

This paper cites GPT-4o System Card.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding GPT-4o System Card

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.600709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:0905611293256e1de41c8b18dc2d823aca0d97a38deecd87a785ee65d9d680c7

Observation 856863e4-259a-41bf-a02b-b435c1916619 · outbound

This paper cites Colbert: Efficient and effective passage search via contextualized late interaction over bert.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Colbert: Efficient and effective passage search via contextualized late interaction over bert

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.293819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:42b41e4bcb06572ce12241947691c264551982121b47e63961396b6f449c9a20

Observation f7674042-f15d-49a9-988d-8babf5d5f3af · outbound

This paper cites Interaction methods for smart glasses: A survey.IEEE access, 6:28712–28732.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Interaction methods for smart glasses: A survey.IEEE access, 6:28712–28732

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.310617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:a3725654f50893f262dfa11ed01d7e28255026e97f08f8f1f8f879ec973a12a2

Observation 12cfcf3d-8cf4-4513-bd55-77f4657c2a65 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.660336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:c35dff693936a69cfe369a0d2963d8d26876ec08cef027a3ff19fd64088bef80

Observation 66471926-6853-4272-a1f4-7383aab5a4e5 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.707856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:d0ba65614bfbfed64adf86eb55b7a3cb4c9e8c6d8019945a3248bbf6852737e8

Observation 930384d9-d9c0-4594-9ade-d05441dec74b · outbound

This paper cites Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.304421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:23f0d603f883e179038d55c231bbcfdd7ea907ad5ae7065aed099bb8d9306746

Observation 2da1e5ae-2c09-4a3b-bac5-f7af9f1b4ce8 · outbound

This paper cites Thinking in streaming video.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Thinking in streaming video

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.723910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:d44358d36407ef8298c2429900fcb98f1ea42e20461aa5540e4dd1624b53b6f4

Observation 33c38f6c-0c47-437a-a022-5ef812e39d04 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.306839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:45f291f56c6b40782b8711751620a94c86610be7d83d5e7a3925fa60499b868c

Observation 834c531f-c251-499a-86cd-301799397b71 · outbound

This paper cites A Survey of Context Engineering for Large Language Models.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding A Survey of Context Engineering for Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.795434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:ad96f283c687703ea2556b0c51b4643efbfd9742bc181e09eb2e7747b93c21fe

Observation 4077ec18-0598-48a5-97b2-22f5156a1595 · outbound

This paper cites Gated differentiable working memory for long-context language modeling.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Gated differentiable working memory for long-context language modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.510940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:5dd22ec2fed82b822bd469c6d97e7d6bb283b868f26f18678ee2f4946aaf3eb2

Observation 4721a065-4802-484a-8aa8-555228acd816 · outbound

This paper cites LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.573338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:7aa174e39c9fc61f6e952d9a644af28e5d1e3808a071d722f24a737c1007388b

Observation 8a3a5fa0-23f5-438e-8bd8-5f15150b23e4 · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? InCVPR, pages 18902–18913.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Ovo-bench: How far is your video-llms from real-world online video understanding? InCVPR, pages 18902–18913

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.337915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:6c49164776d8d030b2872189eb61180e1ca9e94c94c3aaf2a5e929e1f7f9e5aa

Observation 394ee2b2-0a37-46b4-9afe-d094c66f3237 · outbound

This paper cites Athresholdselectionmethodfromgray-levelhistograms.IEEETransactionsonSystems,Man,andCybernetics, 9(1):62–66.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Athresholdselectionmethodfromgray-levelhistograms.IEEETransactionsonSystems,Man,andCybernetics, 9(1):62–66

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.288427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:b481710acaab2ae7a603e15aec00ea6142db7f276ba3148d7821126eb66c25df

Observation cf0aef41-8c40-469c-8371-14da1d47333b · outbound

This paper cites Streaming long video understanding with large language models.NeurIPS, 37:119336–119360.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streaming long video understanding with large language models.NeurIPS, 37:119336–119360

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.285159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:90a77c8612f4807698b6758f3e42124063cb0f90814a195d9eeff0f585744eef

Observation 05bfd455-20aa-43be-ac4e-e18d0848ae8e · outbound

This paper cites Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.319311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:2bbbc34a49f09bfbe5f36f067b1c5b7b23d13b4ee5b379f483c992603b850401

Observation a3b26720-147f-4ae9-a996-392e139c8168 · outbound

This paper cites Longvu: Spatiotemporal adaptive compression for long video-language understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Longvu: Spatiotemporal adaptive compression for long video-language understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.317309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:ce0b3096463756decafbaebf601d44bc110782f3e085986a40e9b8da06ccb48f

Observation fdfbf315-9d46-48a6-b71f-c2678416f48d · outbound

This paper cites A Simple Baseline for Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding A Simple Baseline for Streaming Video Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.672127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:c6f008f281dd6399f58488f28335aa36cc59d06a85b31c8acc6b714616e3c23f

Observation e23b966d-1cb3-40e9-8228-87b3ede4758b · outbound

This paper cites Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.327400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:bfb74e8e9559972452890174660370aca81f38cb1c4b53f2d53c1ee7ef8ba5c4

Observation f2dbf2c8-d2ff-4443-b3ef-e4526cd558b5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.640450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:542ec63e9f74ca48dc1b5e4726dde004f8daeeb291830bbac8c3f5238272012d

Observation 214a7323-1d43-4b65-a7c7-11d81a1ec7cc · outbound

This paper cites Streambridge: Turning your offline video large language model into a proactive streaming assistant.NeurIPS.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streambridge: Turning your offline video large language model into a proactive streaming assistant.NeurIPS

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.301870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:e5aaa7042aa9f99a5c2b0f272189ca9ec317f24f34ba125f4320f10eb89e1f7a

Observation 46f6a05a-b353-4bc2-8862-e30fd90d321a · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.473686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:afabcc01b2a3734de5b4c6fd668540fca285cbf3f6a9b06794f6a74623772f1d

Observation 5ce4d7eb-816a-44b9-acb1-7698013bf443 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.484333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:db0bdc97d5b194ba602037aa6f7099f661c51427464180ec2f4a456de21ed741

Observation c2a8e64d-bb3c-445d-94b8-46bc820064af · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.415498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:55922001f3fcedae29577a800e2b86f3750936e302900a6a6d0103dd638a413b

Observation f258556c-d103-4168-afd5-85ef3680fb2f · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:cd4d07a4bdf6087e17cd2899dc97f5a9055ad7f448993c3584c879650e102309

Observation 049a96ef-ffc7-4c3a-82ae-51cae100a626 · outbound

This paper cites CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.493339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:9c5918a554b3eba2e0140c5fd98c2b685b2f1d17aa11b551666862a97728e9e9

Observation a1a9b116-c700-434b-bd40-d91c97c3ea38 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.NeurIPS.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding.NeurIPS

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.313099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:6990a1d2af4609bbe0502c79ad55fa34c526bbd5a30af1d1a045a79ef226a2f1

Observation ca0e2e08-160f-4b92-ae6b-17df6761d35d · outbound

This paper cites Fluxmem: Adaptive hierarchical memory for streaming video understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Fluxmem: Adaptive hierarchical memory for streaming video understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.434561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:76749566d3214dde40411efbcd9ac514b132ddb10569ad1ecb5c47b700678bb8

Observation ca912523-693d-4cc0-bbc8-1a4248e47317 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:51:33.451476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:b733ae925625e01aaeb47ac0e6fde3b9cfd40a7ac9845b9e28b7839929219adb

Observation 4ccd50be-7e92-49ec-9989-e43b071e86f1 · outbound

This paper cites StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.370849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:59d71961b0408f93021c06c06befb392707c8f191bbc823160a1bb99ac7f1758

Observation 71c3f810-d55d-475c-8b6d-829ad8fbdb2f · outbound

This paper cites Timechat-online: 80% visual tokens are naturally redundant in streaming videos.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Timechat-online: 80% visual tokens are naturally redundant in streaming videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.324135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:99beed62df244d4c4f47868bbbbc6def3c6ba3ff13dd45366e115d50953d1ef5

Observation 7972e66d-952c-44ae-879d-643026da7ba4 · outbound

This paper cites Streamforest: Efficient online video understanding with persistent event memory.NeurIPS.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Streamforest: Efficient online video understanding with persistent event memory.NeurIPS

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.315244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:897beca759c174c9ea2dfa08b5ef9037a25b3b954889de8dbea8b3f77a98f695

Observation 75411568-414b-4639-92f5-0aef2823a9bd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:02:01.218694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:0bd8f100bbdc088358b8e053017b711f8ede4e73ac3e2dc9bfcae2ee716a18b1

Observation ed007fe1-729a-4b8d-83c9-ede42513274c · outbound

This paper cites Flash-vstream: Memory-based real-time understanding for long video streams.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Flash-vstream: Memory-based real-time understanding for long video streams

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.332328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:e396b07b92c90581b49eb934e2ad66191093665295b07af7e8458e8f76e1eb6a

Observation 9888cebf-3702-4571-b03c-8ec4b3f3ee2c · outbound

This paper cites HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.380773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:9ea13ba94090950b55e1beba2e836d4193c8eebfd473227d52ab8a30fe1cca46

Observation b1dfc6ce-be28-46e3-aaa0-58a0929a12f8 · outbound

This paper cites Long Context Transfer from Language to Vision.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Long Context Transfer from Language to Vision

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:64f09c53e8b07cfdcf30caeee14c0a38aee58d66b2fc92a2397e4bb7f3103436

Observation 02b79e42-e07e-4354-8171-c48e8b5cede1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:50:56.454106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:5f6ffd99d64025c2883aadb3f483bb9b0c7982cdd58581d659ef38bf3a2b8d03

Observation 8c780156-d464-416c-b990-d3138c9315e9 · outbound

This paper cites Weavetime: Stream from earlier frames into emergent memory in videollms.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Weavetime: Stream from earlier frames into emergent memory in videollms

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:56.530833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:826a99c23b1b9695b05e21041a914a384d6e67f6f3b7af6e0a9eee397bf28169

Observation 4d6215a8-9efa-4e93-b82e-05d7f8b4d6ec · outbound

This paper cites Whatobjectsarevisible?.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding Whatobjectsarevisible?

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.298959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:b5e29131e299a27455567eb85624b195c49429d14d621774fe68ba119384cc79

Observation 22a250f7-bf45-45ab-bbd4-efe61ed1d5f8 · outbound

This paper cites What objects are visible in the scene ?.

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding What objects are visible in the scene ?

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T13:50:58.291504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:12:20.651111Z digest=sha256:410eb930b28cd01e8814d9c074bfb81573e3046a0eec19a86837509a7f934d62

Pith citing papers

Observation 113d77bf-03e9-452a-8d1b-4a14e641c4ac · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:47.718375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:47.718375Z digest=sha256:a3a83689b730a5953af8b2d664fd737de360d792d4c0ad86b0ec20991577517b