Pith. sign in

Paper Citation Record · LEDGER

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2505.23922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23922 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:40:47.324806Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-14T22:05:07.326202Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T22:08:04.286737Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved33
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b71eb6e3-1d48-453b-80f9-2d8b3b01e1c0 · outbound

This paper cites Qwen2.5-VL Technical Report.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:43.662910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:43.662910Z digest=sha256:0fc5ab05f71e731133e55769b6607fd8b94fe6a608d62b42db6feb34f2418f04

Observation 682d56f4-7066-4455-b7e2-e3b4f148272c · outbound

This paper cites Chandrasegaran, A.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Chandrasegaran, A

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:50.470397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:43.782302Z digest=sha256:cf2472eee27f241076648a41924e3fcc6c5a97032ff2c5803cad19c99e76fbf5

Observation 25d24c42-a878-4a4f-a6eb-056edf6a179e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:43.917741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:43.917741Z digest=sha256:22be8625975f5d97c68fa939f7072f645e0bd0e179b2d66f2f0a32c15efae9c3

Observation 28fc4a54-d057-4084-b47e-3b63eb7dbdff · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:50.323261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:44.012787Z digest=sha256:b4e358120b78bac5f36ce2688f9ae29ddfbcd498eaa06adcde287b482ed66dac

Observation 480c7562-1d43-4583-a763-24ae623d6714 · outbound

This paper cites DeepMind.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding DeepMind

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:50.166115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:44.131188Z digest=sha256:32bb2a09d1628e8bc0c4d5937c41c7180c9d7d2dac656ac57c64eb215ee78678

Observation 8ad5a41a-2573-4dd1-af4a-aafeba535142 · outbound

This paper cites Doubao 1.5 pro, 2025.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Doubao 1.5 pro, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:40:49.972577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:44.232566Z digest=sha256:7920a1f059917f0d31c47511dc562c8b1f00bf76c1089f8a09fecf875596de3e

Observation 88063761-34f5-4ab5-99e3-3dde934e0a84 · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:49.762946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:44.308846Z digest=sha256:60bed025ad5b5049613468d4edf90e4b47a4d0e531dc03f4a1905adfc6dd52d4

Observation 5392d8dd-7c20-4b75-b402-f2267f201a07 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.409794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.409794Z digest=sha256:b8dce64bb4130d9b9ccdf74379c0d3a8fcbf49b3dd997532adc2904ea2f77aed

Observation 4cc616f9-4037-4362-bdce-c9897d81045c · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.530844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.530844Z digest=sha256:69215bc77927f5de1fafd2d187495bd3e65285d43091b8e0013e60ca7e3e33e9

Observation e6ef9a37-2ca8-45b5-9f62-31b1ad98f46e · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.609839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.609839Z digest=sha256:43863c489ebeafea5b8c5165a9d5533e4bc944682d23cea7665766bb21e4eb30

Observation 3f2aa3f8-f99c-4ae7-b191-a46e1b3bc57b · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:49.608571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:44.732999Z digest=sha256:a661bee87db39260cf213c564b1da4775ff92fbbc68fa06c8a38cba0ddcb1a75

Observation b38eac44-3252-46f8-9521-79bcaa3502cc · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:49.453066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:44.844576Z digest=sha256:70ae4925a4410336852aecdc4203df464c0590fcbaa19694af49f8bd11aec2e0

Observation 6f5a593c-f7dc-4764-85f5-da4346e18514 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.943565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.943565Z digest=sha256:9eff59753e9c296ca3eadd9447daeda574f3a54fa86dfaf2b44eb86f1eafb9b7

Observation 8200610f-ffde-49ed-9473-22cc28cb3372 · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:49.244299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:45.040839Z digest=sha256:4e3461858eb3e12a60c1addb2ad329d27d473b0bd3f57a72f950e495d01d09c6

Observation 39f4d80e-f723-41c7-b215-a26c00515798 · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:49.089735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:45.142616Z digest=sha256:04296b87a9c8c73b9a205012021fc61f94bafc5d45a1a87abdc8e6a64354ee72

Observation a3eb682c-f034-47ce-92d5-f53219be48f5 · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:48.872389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:45.223692Z digest=sha256:43b42b91da70fece79e200946fef9d2f6be7a214fa0bfc8012faf7c0d25eff18

Observation 6091c6ff-b78b-479d-ad4b-24aee1a96c6d · outbound

This paper cites Visual Instruction Tuning.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Visual Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:45.339709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:45.339709Z digest=sha256:f607c9a49230f19613333921c7ed6b28d2728313b0b022a76a583623630d4a1b

Observation c26723eb-e2b1-4f10-8436-71652711cd66 · outbound

This paper cites IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:45.454890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:45.454890Z digest=sha256:dd55a5d3877c6c6fe7522103e08f3a7b7f6d7836fa87903c9eb94a67403fa06a

Observation c586b8cd-c0c5-4e73-8af8-033fda347a11 · outbound

This paper cites Mangalam, R.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Mangalam, R

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:45.527933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:45.527933Z digest=sha256:5fc4719fc94be458cdd2d63798dd98e94e74b35015470bb9afbc96509995a1d0

Observation a1ae3d67-2822-4574-8b7f-1ab190ac2046 · outbound

This paper cites Gpt-4o, 2024.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Gpt-4o, 2024

Reference 20

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:40:48.601869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:45.624965Z digest=sha256:1e11709cf7f27fc5688b81e08f5725275c9dd5cb25b74d066602474192e826cf

Observation b2beb961-e52e-4d7a-b0c9-53ffdea1b977 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:45.721408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:45.721408Z digest=sha256:5c00f30be14c666d0275503ed1d211091424d5d6514af072b54c3360baeac2f8

Observation 5e64fc14-c5b2-4a83-87f8-e1a9ac9ca6d3 · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:40:48.321293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:45.806683Z digest=sha256:5b96a232a4b404b4ef1bee1ea8ad8fc7a59f94f7bd51f5f6b79c9360f051edfe

Observation 701044dd-99e8-4a15-9a2a-0502f07eb7c3 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:45.892788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:45.892788Z digest=sha256:1724d59550930ddb2a5190ee4beb0ebfc1121a20d06cf74cd9980cb5ac6016ae

Observation edbc9e1d-c28f-4f74-b94e-bb6de5dd1ec2 · outbound

This paper cites Kimi-VL Technical Report.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Kimi-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:45.986582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:45.986582Z digest=sha256:850e1928762d5de3dae4f2e2f80d2eaeb4e3026fde89518ef5310954100d284c

Observation f14f1403-e947-4140-9265-c2c4a249b393 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.071081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.071081Z digest=sha256:83cb29631f45308c1817b2f2029e9a99b1c476092b845802d2ff19a3bf21da16

Observation 6cd89460-8510-40d8-80c7-13b5adff7d22 · outbound

This paper cites Paxion: Patching Action Knowledge in Video-Language Foundation Models.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Paxion: Patching Action Knowledge in Video-Language Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.163790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.163790Z digest=sha256:d1156977f585b72ef8f5a9e128b4169457f484c170f93f31c8699c70c43d5927

Observation 7ee1a5e5-7b2c-498f-90ae-a196f9785ab4 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.242481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.242481Z digest=sha256:94cd35624f3484d041dc80ae79e4d27c5c6790484307e09667e63b13424561e0

Observation a2bdbbba-6147-4a80-b3e4-0fcdc8e64677 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.344645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.344645Z digest=sha256:75ecc192bf1d3bb99a69f090ffaeb637bea04b147cc813dfe9d46b7a23b39a2e

Observation bb6b502e-3114-4cbc-a0f6-61b4979945a2 · outbound

This paper cites an unresolved cited work.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.432941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.432941Z digest=sha256:ac786b325dc47da588c2f68135e894a45d13aa29227880a45ce5bfe18cbf0d54

Observation 9e49b836-8fa0-4952-b710-181d1be72000 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.514841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.514841Z digest=sha256:b7628887b60b8088db63b6724916dca33a33381a618b09aabbde6b53f8662592

Observation b536e117-562f-402c-b04e-e5d6419eef9a · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.590863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.590863Z digest=sha256:0be48f9aa216384ddacefb3f6dc8908ad8c25ba09c956351999e97bcbd61c332

Observation 2136ca96-d3e7-47b8-a3bb-a1d841930026 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.656155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.656155Z digest=sha256:00f9babda1bc00904975b39ec56420ddba9b3b19c238cc8186d4ce4e95dd710b

Observation ec5f4923-2264-490d-828d-3ffec31585c1 · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.777607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.777607Z digest=sha256:c70841e72219d51af4acee9a4c891ecec11b3fbcde5c635ba867d7b51cf2c83d

Observation c6df2d43-09d9-4618-97e4-0c342718411e · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.869024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.869024Z digest=sha256:43e1dfd46f3456375508717986bc9c6e52505db108f135f0786fbd2468388e3b

Observation 540e02ff-9660-4c5e-84af-462ad7b9eb2c · outbound

This paper cites LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:46.965574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:46.965574Z digest=sha256:907d00be303ddf768e6703f50ccbbd406f95947ff0fb64a94295b264a1185f6f

Observation 0d4370c3-ae89-4f96-8f7a-f774b3ba986c · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:47.090823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:47.090823Z digest=sha256:4a5f56887a7ab3dab26c725a8274cc7ca9348a31fffc283f42bbeb2f921d2b55

Observation 7c83732d-0799-4372-a423-e1939376ee40 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:47.180213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:47.180213Z digest=sha256:1114c29a8111c873c6b1ea6307b6a4e30a21b4519c2a271c60e5ea129658f3b6

Observation ceccd74c-0d78-4a0a-b4b0-e2360e424580 · outbound

This paper cites HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:40:47.648879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:40:47.324806Z digest=sha256:6e5289a58acd745e3f807e0413876226b9862f77327cc440c834bef16f682335

Pith citing papers

Observation 9337d054-7d63-4276-b9e7-359e2604d5e6 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.290536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:f34371025c99d80b91e12b07d960a9dfc67261dbd32d3bc4714f4b58da0e9f24

Observation 2c925b8d-354c-42b9-9c6c-f07824e1caf9 · inbound

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos cites this paper.

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:30:58.542836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:37:40.373211Z digest=sha256:5dd1287e22bd235a3764a6f0e7c994cbd5bafdcf9e2e6043498d69bbcf267725