Pith. sign in

Paper Citation Record · LEDGER

LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2311.17043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.17043 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:41:44.374259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:35:01.039481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 973d328e-1357-49cc-b90e-1af090db8bb9 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:55:20.473273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:acde40268e8800f4e4ac523a49c9b27141e77a62aca87630a62aa10067879694

Observation 73a3776b-4a93-41a3-a408-726ba5d7b8c5 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.731922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:ddcedbd8c9065558ee476792328abbead1a3e7370d758ede3725cb8fcc2c0862

Observation 081ced19-1148-469e-8efd-826183496418 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.497607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:1097336baef7875062cd2adbbd887bb4936c92aa2dfdedd87b7691fa8d325a1f

Observation 61fb005d-633c-49e3-a414-192694e8cc4c · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.974849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:fab5810bd804ffaecf74cd0b18068bd1912578dacb1fa227f4d796e358075221

Observation d26af7f3-d15e-46d6-8cfc-cf9220c4f338 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.389857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:21e85375f862899359e1715dbcf485e82baae258081dc9b51d2c89e55325b95b

Observation 31eeb163-b9d7-451e-ad84-70b55c64fd6f · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.195081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:bdbe7f6bf781d0c471ef3b590a5e177b4fd3738c52755c47a7b82c8c3be16d2e

Observation 7b089da6-9c35-461c-bc2c-368e96b13c4d · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.654562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:b2315a7f9988eb979d05dcfe642cae001a138e88842d332ad497256a5243e57c

Observation 075a7d8a-4684-4e04-b3e0-ef9b26d2067e · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.037741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:18fef078eec48f35ce211bb9a06be08a829fa6a2d312ae5fd0c3adf7e7f84e6e

Observation fc8b33c8-ddc1-4220-9ece-ac46ead134ee · inbound

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction cites this paper.

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:53:33.157114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T22:51:08.753650Z digest=sha256:29718723827aeac89924568fda3749e81c12cd98998219b5579fd767630ca278

Observation 170e0518-604d-4cd6-a103-28c2ad1cb0fb · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:51:25.511626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:512576bea90e785aaa3a99cf23d4f59948eb9285dc59a860e527427c17922712

Observation 81baf91d-9190-46e9-a25f-7aca2b3d8d6b · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 193

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.978337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:77f09b62af598d349a33b62de2bfb1870ff1ba25a59a3b3d05bc78e668038df1

Observation e6fb5e39-367d-4e45-a0c8-deebae6711f9 · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:51:36.333783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:70b72e4a9000e6f328100127192a1c4d50d6c3276e1531d9ee9e5d4c71b5b152

Observation a3e050ee-0325-4f06-816d-1974903bb38e · inbound

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation cites this paper.

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:41:44.374259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:41:44.374259Z digest=sha256:67e1daea7a7dd5a16e5fbdecd5eaa2b35ea474bcd2d1076b890d694f49db382a

Observation c4fe22d9-0890-46e7-b489-e5439882b89b · inbound

TrackVLA: Embodied Visual Tracking in the Wild cites this paper.

TrackVLA: Embodied Visual Tracking in the Wild LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:40.004022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:40.004022Z digest=sha256:b5746590012266ce56186635552f57142c53881c68624a242c383232102689a5

Observation dda92986-97d9-43fb-b6ad-ba90a10bd94c · inbound

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model cites this paper.

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:37.635595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:37.635595Z digest=sha256:a390508d4e4bed4968d7f3a2963c34fcdd1c0bc0aab1ac084d77e0f95962f7d0

Observation 02ef2d55-2f23-47e2-b7f1-85fa1b1795f0 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.537007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.537007Z digest=sha256:beb8914c17c0bd767fc9d238f9f560fca25c7eb37faa70987f07a0b94beb240a

Observation fe0280de-c759-4176-aace-871854251637 · inbound

STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset cites this paper.

STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:41:45.444716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:41:45.444716Z digest=sha256:97679773bd8786389da5bbe790c13f8eac7cf161a0bbad8fffc2a9096193c573

Observation dc8fecce-635a-4350-823e-a604ad8268a5 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.416913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.416913Z digest=sha256:4e707e70bffe60cfcaf7d009f8788e41dbd73307125af2764916c94433a9d9c8

Observation fa43ef72-8084-4d98-b812-6fccee3b5343 · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.326679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.326679Z digest=sha256:336cff7f0731a879d8456043d31d8067a33ef4bf593198cb9f02a708c1a989e8

Observation f9bd112b-a33b-408d-b03c-4849703a8fb1 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.193817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.193817Z digest=sha256:78e3810cf477d91a2056b25ac903a872fdb0576dd2f136069e9cdd31ec15a7ca

Observation 3de1bfdc-6e02-41d3-98f1-9a17c5964c27 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.514961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.514961Z digest=sha256:d5caaa1bc26512b062655700e044e8403084351f41f78458199e464a976b7720

Observation 2efa0987-9240-45d2-aa06-4a09a803cc06 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.144931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.144931Z digest=sha256:8ad811dfeb6ed8c1a4009bf37a9b17d784cebb520becc6230063c96624bb12d7

Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · inbound

SV3.3B: A Sports Video Understanding Model for Action Recognition cites this paper.

SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.740370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.740370Z digest=sha256:0e2217f834fa633ed986e6121a8a4f21e40e62ec8013d1f63db8f5891da99519

Observation be12beba-f4bf-4e8c-96dd-2118725ef2ad · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.230172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:d0f14eb058509e95afdce881681aa2e8e0a52b13144c801792c039468c43eb31

Observation f98a0e98-a01c-4b5a-a1a9-de828b50673f · inbound

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding cites this paper.

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:01.099735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:01.099735Z digest=sha256:03d52021dae090271edf980e8358585f025d3b8c4337cbad71cdb86c423f14d8

Observation 68cc46ca-92bd-4dd7-9f09-73c833ef94d6 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:08.173803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:08.173803Z digest=sha256:26cb839410039cc4b216aa7b1f56aa18dadb6d372c764f7f9ffbe646b2f30791

Observation be4b702d-4fc2-4de9-bbd3-c0e28db84eaa · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.347146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:ca1584211dcae5957c6d835145e3937ad3304a56b0012453ba08c612517e0ddc

Observation dd36a994-d8b2-4431-b287-a07b92f577df · inbound

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation cites this paper.

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T21:37:47.904757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T21:36:43.971666Z digest=sha256:fef9ed1dd7543b258ea3776e1bacfe2ca093ee69bb0ac905be1e43ec0f157df8

Observation bf2a9c55-f062-401e-9692-ed22c84b8528 · inbound

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation cites this paper.

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:35:01.040957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T19:33:24.558488Z digest=sha256:ec6a66950707df912d13794b05279866d9982d6de385b70a9fc7fb265f26c520

Observation bc6cffe9-4ff9-4f63-b216-8efe10b6a225 · inbound

AffectVerse: Emotional World Models for Multimodal Affective Computing cites this paper.

AffectVerse: Emotional World Models for Multimodal Affective Computing LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:48:05.768202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T06:46:33.612905Z digest=sha256:428a16eaf7164ab630690bda846a8bf3bb38fed15a1111b60833380818aed1b8

Observation 35d8a661-3f92-4a48-94f2-fc65384c5157 · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:6153d9b57b7bfb78b2b798f468c9fdfb0ad5d83215d1d3230fb9a216977cead3

Observation fa248f94-9bc2-4a6a-8b45-563b728ff2ec · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 188

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:7eaabe1f6bb1ebcf89dd30f53ef13c2996d6f573d89d4262fedce06a0fa11cd6

Observation 6da9c10c-6546-4c8f-b1d6-032321ab34ad · inbound

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding cites this paper.

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T17:18:41.284513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:18:41.284513Z digest=sha256:b59608bec18bc178679fa0b578288d691016ab87f21f3ea7a253450032d6b7ef

Observation 4228e8d1-8124-4ab5-b3d3-f7c6d59913a9 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.729612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.729612Z digest=sha256:ce5fc55082acac4f097c961ae18b6e181c39f36f4359ba7cf6cb53a9dda0f893