Pith. sign in

Paper Citation Record · LEDGER

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2403.15377.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.15377 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:29.351108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.518508Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ea28090-2a9b-430f-a743-afc02aa2e0ea · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.244814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:31ddb747669ba9a86f27db1f404f705e07db4fa061bee9a1945ba3169234fae7

Observation 6482212f-2fea-4ee8-af46-15f76d04a960 · inbound

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation cites this paper.

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:40:00.050951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T14:39:59.870039Z digest=sha256:260065831e7c9d185f7498061c8185cc59d026823f5a389056d865a4cb3ebfea

Observation 67648cd9-4a98-42e9-9df2-3427496945b5 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.707536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:41c5bb873e2a94381acc0f4d34b6ba81124dd8d3d66f0ed8058777943565e427

Observation be600ca6-544c-4333-806e-450eb4add321 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.143903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:588859b34347e9eaa288eb86d17c1577e1eda41b974a7cb984d4a5ddff1878ff

Observation 9b3a2c62-41ad-4638-849f-bff192cadec6 · inbound

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization cites this paper.

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:02:41.738013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T06:57:50.897865Z digest=sha256:d9edb89ff8a9bf6ef0d627ea9fcaa6b61eb96522e5dbba65f93a49825251c193

Observation d4ddbc8d-88bb-4485-b83d-33783b4f2547 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.358964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:a838932059433cf8927663131cd64dedee18acb9cda9e3e9c99849d02064c8ee

Observation e76c4e79-5a2b-4e07-83b2-a35350833ea4 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.554003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:af02bbfb7d699c3aea9f697482594009a885a75fe7e1fcf06555db0a05b6820d

Observation fe182bdf-a1eb-452c-ab36-07d2a8315a47 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.742495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:6294c758df7541cf10054e6988c84ec793ec202200256b21f077c110bb7440c0

Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · inbound

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks cites this paper.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.351108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.351108Z digest=sha256:9936e9fbfadb48ca57b8ddd3ad68417f7c451ee368140b8d1f017f834534f066

Observation fbc88a1a-335d-41cc-9183-f0aab1822dcd · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:10.374002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:10.374002Z digest=sha256:01c24f170abb0b5b7d973651b0891b2b40b3396a49228e554af9948c6df88a1e

Observation 4f82f8a4-75f2-4b0b-bd9f-12cdf4e124f5 · inbound

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 cites this paper.

HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:47.260706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:47.260706Z digest=sha256:20f353c87dfdf4c915850d4f598bf8bc114a624688607646f54df7de6ee4bde7

Observation c51851bd-3b04-4f8f-9f48-14366c20f7cd · inbound

HuMoCon: Concept Discovery for Human Motion Understanding cites this paper.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.401308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.401308Z digest=sha256:cd96fad32343cae9a6c8b5c680403364f9bad25c2e86f40322ba44c2842f96e1

Observation 5c096c6d-ac30-465c-a225-b79619168202 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.499105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.499105Z digest=sha256:c844b1c6e4a6ea4ab35109b0a9f7a3ba911802a93550e1dff147ee5db25fc589

Observation 34c1518b-6a96-4e58-ad3a-81998148f13c · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.926631Z digest=sha256:941f4838cc3ec58572f276300f016758e7e02d16b76fddb99614113813b703dc

Observation a206936a-a89b-4446-b0ca-5da22eb7bbec · inbound

An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management cites this paper.

An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:59.061572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:59.061572Z digest=sha256:bd3b01165e2190e866ea27991977027a3ade9a5adb13052676369e99f9df3a34

Observation 30471e61-9c6b-4718-a4bb-197fbf6e0fd0 · inbound

DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification cites this paper.

DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:56.851307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:56.851307Z digest=sha256:312e1d7bb7962778b503b74cdb33d9deada8c98929a0b3d99bbf0c72d0dd9b0b

Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · inbound

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization cites this paper.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.161920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.161920Z digest=sha256:1a0700ba85a80c810ff9b64fdc26541947744c4fece99017be79ac59b2aef0e0

Observation 26244bf5-0094-4a2f-a175-3efb65d83ac2 · inbound

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering? cites this paper.

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering? InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:05.964749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:05.964749Z digest=sha256:81ebd4f622a8cf96e5493a2a306b8e6e67db1ed609370f5b37db8bd47bd08093

Observation f2293dc2-fba5-4555-8a7b-352c33d71e45 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:33.036457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:33.036457Z digest=sha256:465a0584620fac5ea1955fd4ab33e21f774e81626f9971d93ed9a82fcc5d2c7f

Observation e683b406-530d-4883-8e6f-29049fa30f24 · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:10:15.130757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:b6853d97d9bb8478ddd1dcba2e515a7a648749b5ef983de09c7d7ddda7cfa91c

Observation 3f95cab1-e444-4ba3-b563-c49f4618a8e4 · inbound

Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction cites this paper.

Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:24:58.979468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:24:58.979468Z digest=sha256:f6effbb705362f7bf8a1f3ac655f690dcebc9a1a373b43f7f737ea74141762c8

Observation 49664a06-07a5-49c3-a551-aeaea6f87f3f · inbound

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly cites this paper.

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:04.588143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:04.588143Z digest=sha256:3180021f8a388a52213e7526ca973e0af9d6c5e1dfd83afaafe1f095806cb1e9

Observation 63b3f5bf-aa89-4242-8139-ecca80d15c9e · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.886005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.886005Z digest=sha256:2417784339ee5eb5565493a79d6c43f45cca42de308de9445cb677c662084381

Observation e4d6e6dc-1985-459c-a7f7-9e3438c90669 · inbound

StreamingVLM: Real-Time Understanding for Infinite Video Streams cites this paper.

StreamingVLM: Real-Time Understanding for Infinite Video Streams InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:51:33.436812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T11:51:33.345812Z digest=sha256:29846da86441bf7303c73fb7c4df98144269050d79522c287d282502ed771a1c

Observation b4991dee-1341-458f-860d-3b3bdb084ac4 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.848489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:8ceecc1b33c2006f7c7461fb57bf60e681ee3176c4a2feac9a2d65511a885561

Observation 9d8ecc9e-0f2c-4a83-b438-04eb4b7f16cc · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.219911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:c0a330ec3ae8bfae17bc6a101e0c4096dcbc01468eee61e13424756d314d54cc

Observation c27af26f-fadf-4167-a8b4-78c00b2f34a4 · inbound

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding cites this paper.

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:47:53.818862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:43:29.123615Z digest=sha256:a0f3f33868c30b5160303b8236c95d70529fc9de6a727d567e6022695a73ee3d

Observation 3d550601-08e8-4fef-bd0a-d450e29c01e8 · inbound

When Vision Speaks for Sound cites this paper.

When Vision Speaks for Sound InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.852107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:12:52.160596Z digest=sha256:63284ce215ba30127ce784c519668f82d956ae3a7c2eb9d4fe3e6a01a77aaebe

Observation 825cb563-f5c1-4de2-8225-143668a3cd13 · inbound

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval cites this paper.

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.375723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T19:12:19.056273Z digest=sha256:b9ca6e07fc26c9fb8cf5ca894a73cfb364952978fade2957538c617351e56376

Observation d949fff9-4f47-4bd9-9828-14cf0d8d1e52 · inbound

VidMsg: A Benchmark for Implicit Message Inference in Short Videos cites this paper.

VidMsg: A Benchmark for Implicit Message Inference in Short Videos InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.063976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:25:06.594946Z digest=sha256:6e4d2f05803bccdf89c4f8bf6385ff7078b9a00dc9231f8e295780e44bbdf693

Observation d76c4a66-6383-49dc-b854-8df5bd1deb6e · inbound

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention cites this paper.

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.379175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T02:22:03.908592Z digest=sha256:b85eb62f02c1eb8d4334d2984e69ba2ad230444b005f57c678c159d8d6a6d88a

Observation f705bd2b-15de-4e9e-a1c3-f31dd6f8147c · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.519902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:c88dbef293e759680dd1d28399c9f2ca7e0455015d38ad2b5b6858a47019cd27

Observation 32b0cf74-2c11-4ecd-993b-6de5dc48818f · inbound

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration cites this paper.

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:50.568899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T04:18:02.341742Z digest=sha256:439cc15f7c0e133151420430213cd1010113f72a9e16301c5690bec41c51a1d7

Observation f73a5272-7603-4603-9b33-02aa4aa36cda · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 141

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:8cf9daf1f4f716931ab4e5ff4d3100ae498e99c52d864f0848e13a96c1bffd57

Observation a5be6167-3f15-4bf3-8fd8-5b7bbee536d4 · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:04.287534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:04.287534Z digest=sha256:217daf000b11d6185645850c3d133b4be567fc93764bb30df6cff3388cb6cceb