Pith. sign in

Paper Citation Record · LEDGER

Apollo: An Exploration of Video Understanding in Large Multimodal Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2412.10360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10360 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:08.568366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:02.918791Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 51649a7d-f4f2-4ec0-8e87-21456cbedbd9 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.833571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:a4581d34fcb4a63bc92117aef824d4dc04f505f175648efe1a1eb7f17a2d33ec

Observation 3ca642b1-742d-46e3-9920-106d019fb4c1 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.181502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:5e903e8e0991eb50a75ec95e8722f0c4cf08dfc2c3bfc5996a90835bf930da2b

Observation 73db5c87-2b89-4b6e-abd9-9da8866b837f · inbound

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? cites this paper.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.568366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.568366Z digest=sha256:0468d37f774724681a0bb2505acfa4db918dced44b707d513047a1a2c2fd61d5

Observation 208b340b-771f-4eda-aa3a-b4fde362b2f0 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.930004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.930004Z digest=sha256:d8867a32786b8801a65f8c7f04bf74d847c5293fe3399e337e36d3212c0c4aac

Observation 3e01e98a-c288-4951-bc8a-ad6f1c0fdcf2 · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 120

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:14:47.811975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:47.811975Z digest=sha256:2bc1be6817137605f56fd87c3a2b81f91a7e8f166218978d144c9c8868379850

Observation 447bbede-b0be-4fbc-9feb-22827ba5a7bd · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 84

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:37:24.588694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.588694Z digest=sha256:9226a2aec6f3cf3ef4afa9581954b4ec43d654505ce5df7f04abfc5f44467eda

Observation 0ff19277-4d81-42ea-a6ca-e0b1fdb91644 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:59:11.354129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.354129Z digest=sha256:ffaf46d96c552297efd5ba478ca4fb56f95424a24c600ec937af4a2532454f32

Observation 1ab9b6da-ae28-476b-9d98-ef9bd5fe884c · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:16.401147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:16.401147Z digest=sha256:8080d59ed9801e2ae90cdd39b18dc4343f36bd2df84d488f6b4334e4d61a40bf

Observation 1645717c-29a1-4eb3-8e9f-e7e2978c2325 · inbound

ARGUS: Hallucination and Omission Evaluation in Video-LLMs cites this paper.

ARGUS: Hallucination and Omission Evaluation in Video-LLMs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.827734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.827734Z digest=sha256:dbc1c56aedacafffe88bebccc31881c6df9f51319494ca455f6c55a6a8cd5d9f

Observation aff8f26f-bed7-41c0-a3c2-a8a286242dcd · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.976093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.976093Z digest=sha256:1d66db5b1dcfe653bd55f6843e3d964da20993ce529cbdcfdd9a1a940d3318be

Observation 0a20b61a-a85a-4611-93dc-11fbe6e68ff3 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:57.031289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:57.031289Z digest=sha256:b08d991b40ea28414f9ff947ce666cab6b96076e9143313f766c0d7f45891d8a

Observation 5f9be675-8665-4051-ba81-0ff362f1473c · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.498173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.498173Z digest=sha256:aaad387ac8831e61e446d3e1d248f3f7392396226391c5e41abfd00ca781a509

Observation 44079027-0ad8-4f25-8acd-0306ad0d4fb3 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.575145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.575145Z digest=sha256:abeba33f4bebdbe831dc548b2dbc3df6ef613afc20a2014c1d58d102cafde3a8

Observation 6cf8de07-4242-40f0-a445-23d5848dd7d5 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.512195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.512195Z digest=sha256:36a03a1f1406e3dca5a1deda3f0c7880a03026426b8c195462ace8a28e8eda26

Observation 4ebc1d82-7cb1-49f5-9264-9aaf78b8426d · inbound

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? cites this paper.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.145067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.145067Z digest=sha256:f2e2d7776fdf0d713c8dbcb9dc945cd7532d460cc4ecd4ea5fef615871b48665

Observation bf372c0d-9bca-492b-a2ab-78d4e97b3cb7 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.860880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:7d3cbbf6df40f125475ffdf0541dd4f944279baf6b86e5411a84f0659dbd4788

Observation 318777ff-820e-4174-8411-3e46e8365194 · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 79

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T06:46:37.520929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:519da4f9d214b60cbfa5887040fab74e4464bcd27400e1387de3f318f6a7b493

Observation e639179c-8052-4189-8f40-7fd587b4e7c3 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.639077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:02760299940ba4bebb4f9ff907134f06cbcf9e566bcd75857d7e76c1f5d8f476

Observation 0a37124e-a73b-4485-b966-0c8d1371122d · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:91477aa8c21dcbbf9c40639ad21016ba25589fd7145adddfda404866b6597a11

Observation 3113e5cd-019c-4b5e-9caf-e93928e905c7 · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T15:34:37.002001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:34:37.002001Z digest=sha256:e6a68d838bcebebc688d26051e504c1245b517fc9cda666bc12bf7242e33a350

Observation 842d6468-15ae-4852-8695-f90881a5b040 · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T18:39:24.915547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:39:24.915547Z digest=sha256:06b47a8441cb98b255cc153644b2fdcdc86210d4253301c3c8b96b2569168b28

Observation e62a9e4e-ff01-4793-b538-68d25de14127 · inbound

PEEK: Picking Essential frames via Efficient Knowledge distillation cites this paper.

PEEK: Picking Essential frames via Efficient Knowledge distillation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.314765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:46:53.704892Z digest=sha256:60903ca64e4f66d13ec31bc877c8a6b5781d1b66f5650012673b4497b7c63759

Observation 1b33d8c6-4bb1-49b1-bbfa-ae9e82327cbc · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 263

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.920378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:6282dcc14fd753f8736d66cf47328daebaf061d9327dff47779d4d2e61ba822b

Observation c9dbe76f-b26c-4668-9b61-2ec3dea2696c · inbound

Agent-Computer Observation Interfaces Enable Dynamic Computer Use cites this paper.

Agent-Computer Observation Interfaces Enable Dynamic Computer Use Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:21.211164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:59:20.295818Z digest=sha256:68fed0be5a25151152eba2b66aefa99ac5274781ec579f52ac9487f3db20ea0d