Pith. sign in

Paper Citation Record · LEDGER

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 19 inbound Pith citation observations for arXiv:2507.20939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20939 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:12:41.078055Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:28.047281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.275005Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09faad0a-b366-49d8-9cb6-2caf2d9a1b54 · outbound

This paper cites Qwen2.5-VL Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.360764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.360764Z digest=sha256:40d4c1a8a61d40e456c325a675165ec4d568c9b10a4dfc6302f1240b204dc4ba

Observation 7548790d-90d4-4c51-8750-4ca490328737 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.953417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.953417Z digest=sha256:74729329540b48f5fdd879d785ed0c945f89a307c9dba8159a1ea3591d0569ee

Observation 9b6c2843-4cbc-4135-a943-43880b395e63 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.150961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.150961Z digest=sha256:2f3241790e1e0e43bb1fd140bde45ff891ff098938ea9bb9800c3b49336ff14d

Observation b4a0a645-4785-4ba4-a8bd-573947db35c1 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Audio-Visual LLM for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.249511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.249511Z digest=sha256:fbe87d2768673784d975e63aa8db7605fa3f306bd84386acc60e1e7f0335e2d2

Observation 121a0ef3-6a52-41a4-83b9-bcc9bd4d0e63 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.387502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.387502Z digest=sha256:6918d71298e40765f1071bdf61b0ac20daaad2b3c111730b6667174ab0b6232b

Observation f9929eab-f34d-4ddf-99e4-742369172e99 · outbound

This paper cites video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.512824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.512824Z digest=sha256:c46093c4579ad4e868a78c5fa34557ee6d554605ff8134126f90db755ce85e82

Observation cd8c78b8-5b4e-411d-a750-cbaf2e8ec76f · outbound

This paper cites video- salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video- salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.594915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.594915Z digest=sha256:9fd778580a3d483114596135e0ffa2bfecb5d3e7a72cdd76e359c785db1ef66d

Observation 36e2cfcd-788a-4d8c-85b5-7d2114fb4bea · outbound

This paper cites Kwai Keye-VL Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Kwai Keye-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.675805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.675805Z digest=sha256:01e137083f3c522a5997d29a14df05220309399eabd5f80e25cb1fc0a4889486

Observation d90faea1-3e4d-4f84-9b8d-1851eb24d599 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.870798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.870798Z digest=sha256:19ac5f709bee218784d3d11b64903057f5d53c5e3e814f72b02f8f662ec256ba

Observation 7e81c92a-5039-4fec-b214-1c1dc739921b · outbound

This paper cites Self-supervised product title rewrite for product listing ads.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Self-supervised product title rewrite for product listing ads

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:12:41.674677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:12:40.971685Z digest=sha256:a8c92daba2c92b471c842f43076013494c0d55e6428f68ab4fc6b0b0ffd58a38

Observation 0330b3b7-885f-4f0f-b3ed-c34643bedbe1 · outbound

This paper cites Pre-trained language model based ranking in baidu search.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Pre-trained language model based ranking in baidu search

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:12:41.484033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T13:12:41.078055Z digest=sha256:d605cbecd6c5f7028fbe96354b48b88260e7b0e15a16cdeb659868c60deda3a9

Observation 19e9010d-dacf-4543-9715-976def34db69 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.428823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.428823Z digest=sha256:26151d58e0216a547356a2d6e38af67a18f351af3fee88fcfb2bd3a803bdbc81

Observation 83b99e30-bf69-4500-bbe8-019717c45630 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.878853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.878853Z digest=sha256:f9dd140fe0b8d373e8fe8ba1018d7f899043a7f18bc4625e4b00ad68c00cd3cf

Observation 8adc1884-5841-4d92-9623-872b8dd0785b · outbound

This paper cites CLIP2Video: Mastering Video-Text Retrieval via Image CLIP.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.766732Z digest=sha256:7089788c041b77579a9fc23e3d33f1fa79cc2924ea6a7f4709fea10d5a7ce85a

Observation 4bb1c417-c44c-49eb-b3bd-7193a9df94e0 · outbound

This paper cites VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.063831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.063831Z digest=sha256:4bf30b68242c7ebcc26619d8b10bfe6d9d9d90b05f54a1366f8c56d643f4ebd0

Observation 929118e3-4139-417a-8134-569457a97108 · outbound

This paper cites Qwen2.5-Omni Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Qwen2.5-Omni Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.757999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.757999Z digest=sha256:2a7816878f02714725f870c6bb25c5b0e156da544690ef01f90cb5e84f103836

Observation ef3698a5-928b-43b8-b734-8e05bca71011 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.499439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.499439Z digest=sha256:f160f001d8adcb0be458337831975ca157591bcf840fa5786cbfd7c9c290cb26

Observation 17226eaf-2214-427b-b571-910ed26a187a · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.601562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.601562Z digest=sha256:c437444122abbd6b5522349d288d1d639905286c0ca09cddc365bfc9ad57852d

Observation 7da6f7e7-2ed8-4976-b4ca-a5e4d7ee1be7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.679158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.679158Z digest=sha256:fada45bdb8658c121e6c567ca934df83702f592c408c3f57cc7693c4e6d29111

Pith citing papers

Observation d82c0732-c3fe-44a9-bd50-e5bb579ddea6 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:28.047281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:28.047281Z digest=sha256:a25eaf9a6a2909b9c404e2801403685b7b742b9f73102f6d05dae7c976ade40b

Observation 016ec883-792f-4898-8ee7-42fd8c75a35d · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.022646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:27152f47c7b4b2838eff9726c22c25118667549123031f68e03efa001497c7ef

Observation 431cd667-1fe8-4634-a1e5-be3605422d64 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:34.444944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:e959052234f914fb6e99e330bb522ad0902c1bd0e6f459b02a288ef3f4cc7b59

Observation 895f2136-3282-4deb-91f7-27e2bff74873 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.809969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:4329258720d4dec09fee42587117141c16b866301294701e6b2d3e9ae8fecd95

Observation 30240866-491b-453d-8a29-00fb300b939d · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.361256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:8eab023c0c1578a3d8d760c8fca79b1f89f8dc9758991f48bd81d6f7279b7f57

Observation 0d5a81a4-f4a7-4ce7-9d4f-f670f02cb94f · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:23.065987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:23.065987Z digest=sha256:c35221907bc1a5f2aa4e77d2c86d1b339ebf5b0a1b8a12cef5323101ae7777d6

Observation 2f3689b5-1689-4621-8177-8e9a4589739a · inbound

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video cites this paper.

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:13.095482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:35:23.842691Z digest=sha256:8d9401cf4bec36f16f604b47a4e1041f05d102c8a6cb07daa23da54f824e1ca3

Observation 110d5eda-42cd-4a0c-a706-ced4bcebc4c6 · inbound

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning cites this paper.

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.636648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:04:15.238840Z digest=sha256:6c1fde7cb3d5582db25057723a433076105712ec85be94857671dd2c2608304e

Observation 87b80d37-3a54-4c85-9892-6f7cf8472eb7 · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.273746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:fef0faaf3b1d1c3714bd678adf12a93d6752d9d22d59b0903238e14825f66098

Observation a522d834-08cb-4dae-914a-d4ee04358f2d · inbound

Stage-adaptive Token Selection for Efficient Omni-modal LLMs cites this paper.

Stage-adaptive Token Selection for Efficient Omni-modal LLMs ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:08:04.999290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T06:05:44.739490Z digest=sha256:323f80dcebdfc32f2f630fd0ffcf7a7173ace08a12caf4c2bddb8bbb1ff7b273

Observation 0f3610da-8dd6-4be3-aadc-32053dd315be · inbound

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding cites this paper.

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.865844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T18:36:29.376048Z digest=sha256:d68d25bd9a4db86a88943e9a8fba18ec4901a46ebc709401a2af8be15e561b3e

Observation f4ef4a4c-eac4-4efb-a123-54a804582313 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.276794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:2a65c9d29801c37f99956f5ac85fdef501915eef6d505a0abd301dcc00497e32

Observation f5d9aaa1-58f1-450d-aa96-2d51e362a60a · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.578572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:2bd919b749426ed86f82aa1abdaf4fe03d1783cbbfee79ee08f4f63ac5d116a0

Observation b1162d31-7bdc-41ae-b96f-3fdc4ef55fc5 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.611522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:d72d01ba659fa1f11ad3792e4f0c37543e3c4cd7eab6f4dab0bd0ce6f0f80bc8

Observation 182b0ee0-427f-43da-aee5-3e1c50ac9f1e · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:8ccb8ec724cfe3f18ed1f17a51439a7693e73c35df0e5f569abd74c712c18f4a

Observation e773c642-522a-4375-8026-1b9b52077a05 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.165633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.165633Z digest=sha256:0b0929279a945bc2b718f549b3ab31ec755c1482c6000ed939d4a18e2d430602

Observation 6d25759c-5f23-49ae-9276-95bc2c367031 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.828792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.828792Z digest=sha256:5580703095b29c1ff9c3f0f21c73b213b225beada95b757459aa71222eba75e6

Observation 70c59f4e-51a2-44fa-bca1-f081081e2100 · inbound

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs cites this paper.

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:37.839045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:37.839045Z digest=sha256:630de57d412ae9cb97fada8de3c90de8fa4fc433e094b9d86531a5c3b6e1aae2

Observation df338285-baa2-473b-a284-b677b5a98a8e · inbound

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models cites this paper.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.785575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.785575Z digest=sha256:0d7a74b6591d66a929fc35645f3e778c4c819373277f2073e3412234ce90fdc5