Pith. sign in

Paper Citation Record · LEDGER

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

As of 20 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2504.17821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17821 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:00:03.666933Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.121274Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a610f35c-47f6-427a-bd63-3ff7d1c30a28 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.440183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.440183Z digest=sha256:e834a7cb7951f77ec07fe87f1e445bf2bfe96ae461c0a33e27c45f213b1b603c

Observation 9482bb42-850e-4de0-9b9e-9238f6446b3b · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.445687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.445687Z digest=sha256:be2c9fcc4bf3980dc4b23c2687781e7db52fa2710ca120cbef5b9d59db2cbf48

Observation 1c8ee396-9cb5-4dc6-abfa-f7f77019b6c6 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.449989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.449989Z digest=sha256:af033b3630083af94464f8122fe82cc3dacd579d02bb2ab69ab60a154a74f142

Observation 127a5aa1-6d62-4f48-9d1c-ceaf86b6781e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.454622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.454622Z digest=sha256:ac12c6c3e9cf8713e951ab53ce37ae7d37b74adfb87aaf42405bafe0697971c2

Observation a35ac029-0e78-4a41-a0c9-74f845b6895e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.459045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.459045Z digest=sha256:921cd7cc9cbcb5c79cbb4b2143ede4a8ea0e30e1e5172df66fd93befaaf0a025

Observation 70026d46-fc90-435f-ae7e-87553b969b7e · outbound

This paper cites DeepSeek-V3 Technical Report.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.464005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.464005Z digest=sha256:acbdbcf346479879b4c2e30f85aa2f7fb5860bedd09df93fb782215d7dc72162

Observation 9051d4e2-6476-4cf0-ae96-b12a1a39242e · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.469725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.469725Z digest=sha256:6e5a1122655c7e3f0e6ccc1431284afba9701ffabdfe3a5101f0cfafba4fe9a3

Observation a386e1c8-8df6-49f4-a812-084e235ae6a5 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.473696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.473696Z digest=sha256:dc3753028d7bc422d217f990c8692783899eb696b6f5428ee5371dd88a964cce

Observation e44d968d-0695-420b-a545-115b3aa54e79 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.477776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.477776Z digest=sha256:7704b85574b798738caca4249f9931cc975230d52ec0904440db6e55164f4f1c

Observation d3f1e497-125c-4a58-8be6-1cd561de75fc · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:00:04.453645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T11:00:03.481652Z digest=sha256:b659d5262041a00f22c648128942af893bfa0db63b4b39d83fa5d121a893a204

Observation f5acbb57-9f48-4202-a2ac-48a5373ea403 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension VideoChat: Chat-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.485560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.485560Z digest=sha256:12659e2b5cbe5605ff1652f969ef24eb80d4741a5a2c39394a8720831097819f

Observation a021d248-de30-4269-afba-17d8a32ce5e0 · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:00:04.439864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T11:00:03.489725Z digest=sha256:9db0f76be46428a38acd3c33688724d703726c2f3cf35fae3375a61daa19e72a

Observation e93ce6d5-4f91-4416-ab47-4a9b399245ab · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.493355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.493355Z digest=sha256:e9ce42429bfe390e02de8162b6f97c459d6a80ee732191391659c9918654c6e5

Observation c9a7165a-78f0-4522-9a70-64d9b771a3a3 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.498647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.498647Z digest=sha256:e02d229c3c96f1489072d7fb7aca5cfdbc243b5a85460a94800669496dd4e938

Observation 8e0c5624-eb6d-406b-a04a-1e8fd04b1e0c · outbound

This paper cites Temporal Preference Optimization for Long-Form Video Understanding.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Temporal Preference Optimization for Long-Form Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.502576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.502576Z digest=sha256:5dedd9ecf4546f787628c858add320d645e28080071c240644c7b015595ba0df

Observation 89e06abe-9ca8-44b5-ab3a-7049a627d3e0 · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.506606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.506606Z digest=sha256:bdb09861183dba871d5d43a6d7d0c879074463a053f0a9ab35e09954f43ef093

Observation b5505f57-9cca-42c6-b025-27e7cb73779f · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.510166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.510166Z digest=sha256:5bf046053a05b7cd2cdd826b811f915e3274ddbb2e3dc1bebd260e6b448f6840

Observation 10500438-ab93-4423-97cc-d255bad68e1b · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.514208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.514208Z digest=sha256:732053cd698cbce1a54ae1790f2ab9a3cc6f5f9ef6f431b4397969c5216adbdf

Observation eb9c21e1-c7a6-4149-a535-4833a2a5dd9d · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.518752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.518752Z digest=sha256:91e68ab673be494360def8d370e1a9290ea76f1c820bb22c0493fa4828a0ee67

Observation ea5d9e1c-ebf7-4f3a-b88b-fe80723433de · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.523123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.523123Z digest=sha256:b5f3ccf3396db6759d0b309c0e144d13cc790ce8930555ea7ccf48b9c8e8a67c

Observation 12d8bbdc-4979-4dcc-ac72-5ff9ec7a77f8 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension VILA: On Pre-training for Visual Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.530163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.530163Z digest=sha256:15c32930e0f094128648ed9b54f231c85184736e246bca45379b94c96d70f30a

Observation 46c87361-23b2-4fd7-ad94-a7b55e24fd8f · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.535087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.535087Z digest=sha256:65492366277caa04d12ecd2dd34c6d4993d1d8c366d4233253ef99d1bade8061

Observation 63487eb0-6cb8-4223-97b3-7b856b3017a8 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MMBench: Is Your Multi-modal Model an All-around Player?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.539457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.539457Z digest=sha256:996cbbba2c6579d5cb84678bf0b2e30dfd0014d5496a1a0972bf90cae5e27e05

Observation 702622c8-0b25-4d84-b597-9cfd11b1ab63 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension TempCompass: Do Video LLMs Really Understand Videos?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.544988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.544988Z digest=sha256:c1986b0d6b50b35cfff1dacbd17b53d208fd3c98b5ba3e320098b98428d5576c

Observation a73585c4-cc32-4249-9408-473bea2aa5c0 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.549842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.549842Z digest=sha256:30c5b04debec3ea5d7e3d6e700c6e1fde41d9617d688ad2c64e7cf7b471f1f95

Observation 50579325-f441-4c1d-974c-e54258189fb8 · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.554564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.554564Z digest=sha256:07108cf6e96982ddf96b704cb60e2e4dfc559ab6eb0c472dbe969a09d5af62d3

Observation 25b8f88d-331a-4099-b8d2-76dc7acbdbcc · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:00:04.409641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T11:00:03.559148Z digest=sha256:9816d821b2f030fd259d48d791b5f097eb56b9cd5d09324d2237591e643f4ada

Observation f417843c-2461-4162-bd5e-d0bd44bbbe88 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension SAM 2: Segment Anything in Images and Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.563688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.563688Z digest=sha256:c0a0ed162c7355ff6c97564ea7ea9d05ec991c74c1e935e2035f88f457aff584

Observation 19103175-033e-4212-adf8-646344080015 · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.568528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.568528Z digest=sha256:83d281df15f0ea45003066a6df0d80dcabd2d4f01c2cd84143ae0a723ec04151

Observation 3eff4938-6bef-476e-852b-4d37ada7ae00 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.575717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.575717Z digest=sha256:e39918ca3df6d916426d543f189ecec6d3b46dc83c89e159b1923c8d2a7f64fb

Observation 182fda96-18ca-4532-8673-092e54e48300 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension LVBench: An Extreme Long Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.583634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.583634Z digest=sha256:03f1d0cecf0d6ca792c35feba2e13418ecb406cf4577a3fb4028e94ed81e49a9

Observation ab42477c-3080-486d-b3a8-83686e267024 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.588993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.588993Z digest=sha256:33602e4f332e1caced2258f6927c4d0681c9f0e9ce0bdd56985b1a7f7de01374

Observation b91cced1-9938-40a1-a83b-b77f4acdead4 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.593515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.593515Z digest=sha256:283a8db14a8dfb487bd6145a654827b00260dbcba5ece71bcf3689d58b9ec140

Observation 0f9d9766-4e47-4ead-b88f-c70b87755de8 · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.600002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.600002Z digest=sha256:ba72d07e10cdfd7a14b9555432ee04abb1fa5ea7eeae914f0577f186369c411f

Observation e0106d1f-4328-4876-97be-e29f4f2c3df0 · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.609657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.609657Z digest=sha256:0811f7a2b4eb1b77dc660b8542a3aa03c433b3528a4557688bf723cffbfde476

Observation f85781bd-525f-4f60-b09c-7855a3af871b · outbound

This paper cites Qwen2.5 Technical Report.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Qwen2.5 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.614556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.614556Z digest=sha256:fce23a693699b2632cb9ea4397357b4ce6b066417f095bbd86c8804a14f3dc7e

Observation 98e7b97f-5559-4f2d-be8d-6e98cb97d891 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.621553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.621553Z digest=sha256:f8499e79da05fbfc0b261cb8ceb68dcfb1d92e8279abc2764362fc26bd5f68ee

Observation 9bc97747-c0e7-447a-abec-d1b61d007bac · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.626587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.626587Z digest=sha256:0f852328b7ea2c42249e018fcb321ef62de294107d426f8cc4e738b2c0affbc8

Observation 6433f426-bd9d-49c7-a665-7ed466125357 · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:00:04.370958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T11:00:03.635224Z digest=sha256:fd4cf3622d9f440060322b6a1c802dc12c7fdfb9929622b007472ef0f0188724

Observation 57b827dd-babf-4a19-a9f1-3071a60e5b2b · outbound

This paper cites an unresolved cited work.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.639558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.639558Z digest=sha256:258938bae6b81a70407ad359cc123ac8d53c751fb3242d68a52e4b4f48561481

Observation 5e98e299-6735-4e30-bb5a-371982ee797d · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.644348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.644348Z digest=sha256:63d3b56adce4c3397bf1801b47afff49427b76ba11c24c0f0e0c4a3015e9cd3f

Observation b4b56268-dd2f-4396-9f16-2f5dbf29d58d · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.648627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.648627Z digest=sha256:e4d5018129e5658161e775c7198dc03ffe55a4fc0496a1d122305dbfbc378bb0

Observation 52da37b1-e70b-414b-94f3-fe1c7633f1bb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.653483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.653483Z digest=sha256:8936209143f42997ad11dd51821b78bc7207f1ac0892f3f828f77ec6b6ec0d74

Observation 212dc451-1990-41ac-bf03-02aa60471bf2 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MLVU: Benchmarking Multi-task Long Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.657796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.657796Z digest=sha256:d5c6337f65bf33a6829060fe3662cb380cf663a35253ec16367094341862800e

Observation a3d67e98-a8b6-4673-b287-a5421ce19db0 · outbound

This paper cites online" 'onlinestring :=.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension online" 'onlinestring :=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.662387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.662387Z digest=sha256:f4bb07f0b825abd43f44a9ad18956696eb97fa1c09c9c25c04cb897f8676ff6b

Observation 8e31cca6-be2a-488e-877d-d0c6123a751e · outbound

This paper cites write newline.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.666933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.666933Z digest=sha256:0cce4bec6911a6746d93d56bb1cc38632d0974737a27967117ff7ad11cbbbd3d

Pith citing papers

Observation 23230567-8d6d-4e3e-bfb5-008657d004eb · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.121274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.121274Z digest=sha256:f5637b5593bd43109793340042111e68b6c4db7eef44a542c5ed29d978575e56