Pith. sign in

Paper Citation Record · LEDGER

Valley: Video Assistant with Large Language model Enhanced abilitY

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2306.07207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.07207 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:22:51.754421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.409899Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 897e4d1c-875a-441b-b2f8-26c72af5d579 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.650439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:c0d3a3748ce352d3a13021c41ecdaa40f9b01e96aafd277a4755cac4daa16266

Observation c342be09-1bd7-449d-b81d-5e098411cd7c · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:08:01.242531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:9d263aa07c6a09867d2d5b4024d00089e0e97362dcca66641a64781ba617ce9c

Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.014045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:8f8397f00b9944b90bc5c09b47cd260de64184d8ffb75d8d013dbcccd11fc0ef

Observation da2205a2-3c68-494f-bd58-63204b3bf502 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.769669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:25e2c5f499f000237ef2f29ec5c0875b10ceb33a1b940599ae5c05ce8684c228

Observation d8272357-eb90-40ef-b937-06bdfc05800d · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.884272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:aa52079b8e32db3c6bcf4eb476daad4fc294b6f38cefced642087c0d43b52571

Observation 634a0eee-6bd6-4860-99f0-19f0893c45f8 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.675561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:7c796af3afe40d484dd320e915169a05d6ebc3b784e466a674b7f9e51903b79c

Observation dc8993eb-df65-4539-8b6b-4bd24028e0e2 · inbound

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance cites this paper.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.706077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:ff1e918c7275afa629e183c5765fb523c8f9231877f8fbec7ac276435eef0140

Observation 8f5e0e6f-f569-4ead-809a-2a3b4e6d6d3e · inbound

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos cites this paper.

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:12:43.870984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T08:08:01.889675Z digest=sha256:7914dfd52f42cdc47fe9f35490333e709dcc761b1434a9754f4806e99cd7d3c0

Observation 08271f24-be62-4741-8711-197e91098711 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:22:18.746828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:ad6f3eadbb2634cbe9d0a9daec3e29182471ccae9e0c40bed1dc4159db85a748

Observation b589f826-7aa3-4d32-9b16-895eee05959f · inbound

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling cites this paper.

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:22:51.754421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:22:51.754421Z digest=sha256:f563220052b2272133a357311647a83d9db8eb93112b5677c11a2417552f21be

Observation 13fc8891-7b8a-4efc-b6b5-fc1689b8a69c · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:13.167183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:13.167183Z digest=sha256:d171c0a746b2e2ff11d11d5e8761b813cc1d5d0fca8816db75e5827cc6b5265d

Observation c1b43079-4c42-428b-a47a-aff1f18e0f89 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.929275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.929275Z digest=sha256:41af8307a1623ed30639b6edf72ffd861701390df082428fa5262bb0000f06f1

Observation 1646c9f3-a123-4005-9125-201509576e94 · inbound

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model cites this paper.

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:38.671225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:38.671225Z digest=sha256:60c164cc657647cf342b93616f50a77de772fa7fd1d00e3fc90bdb8db1d2067c

Observation cf23a219-89bd-4daf-9ae1-0fcb6e0aa809 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:29.648522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:29.648522Z digest=sha256:d1fde52daf44bc750d4a3ddcce0952a08cb665eb21b574ffbf9a4430c0ded827

Observation c4229669-f0ae-4e90-a50f-ac08f48c9807 · inbound

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding cites this paper.

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:47:10.369499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T07:43:26.736409Z digest=sha256:d0136a8df13a1ca9719d3be003f2bc2cdef0a6e14bfd1a5816b62dff21f2ffaf

Observation d25c8878-8950-4f13-a1ec-3ad356218ae2 · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:37.503819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:37.503819Z digest=sha256:e9f66dfcb53c53114bfb7ef02622f383e6f79fb3cd0cccbf1ae42d4331725ac8

Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.176899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.176899Z digest=sha256:7e1b13f8e904731260459afbf62609fba6bf63543f8418af7f21637cb7408d6e

Observation f21d3b11-8dd8-4f2a-8183-1df5bf4165df · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.652829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.652829Z digest=sha256:0e820ce463001f52285f447177a04e0e9315e62bc779318ad9dfb78de0a78bab

Observation a6630497-dc85-42bc-83c5-efaa1c28942c · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.875262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.875262Z digest=sha256:c903749ee5d7ab68f144d77707c0da4b3d6febdcfe2d9769c4dd36cad4ccf0a2

Observation 34b2886f-ccd6-447f-b1c7-55ef0e380930 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.248728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.248728Z digest=sha256:615a556584efdda0e46fc4b73e965b5b81a2617c8f4f669805e37317137581b2

Observation d5431206-c977-4ddb-9a17-869ded7abf2b · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:47.855446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:47.855446Z digest=sha256:a67e2893b46ae19a76d244bcd859087e0788c5209b00922df4fc662a5f8e5693

Observation 0536ee5a-5b72-4bb6-b5f3-d8f5ecdfcc87 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.906964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.906964Z digest=sha256:d9efd6e02a5858d2ee9d9939487ca8c091c9d01832642593009178e17e3210e3

Observation 3780dad6-ab2a-414b-8f9d-9e81fe75ccdc · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.556812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:4b0a91a3c36cc6c79aa860d30e5b8f8a0ea15052f3abdd9862ad90a46d973e1b

Observation 89f4997e-db72-495e-a46d-391c9c729afc · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.044127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:567676a68e162e433c33840e3c1afe2e0308a1b2f98327a89e9632e3c90aa4ac

Observation 053cc833-2807-415b-b9d3-2620cdc741e8 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.924293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:ac5d474921df664b4174f859d1dffa8c2bb6941d76e9208ca73c4cacefc4656c

Observation 6d2e3842-bbb3-4e92-84ad-800e932d1c89 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:9ee70a74efe68024e083c1c7adf99948053bb1e644db76488342f1ff29a88569

Observation 6ef3b4ba-347d-4592-8f41-d38510568d4b · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.566195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:881d22ef36166eddbdf9e8aebeaed8e015180f2dc6dfd015351e6e0f03d0cbf3

Observation 9a223249-f3da-42a3-af03-839e912aeace · inbound

ClimateVID -- Social Media Videos Analysis and Challenges Involved cites this paper.

ClimateVID -- Social Media Videos Analysis and Challenges Involved Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.448614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T05:25:05.760276Z digest=sha256:73c9f145b7a662d8a36b03019387592b690652a82f233146f8044072530fdfff

Observation 696d60da-b03f-4d11-a459-80b296cf68c3 · inbound

Dynamic Model Merging Made Slim cites this paper.

Dynamic Model Merging Made Slim Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:18:25.408077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T15:16:24.651868Z digest=sha256:c27b936e0816dcf867402a25537a64e6a49ee18d47b4d8f71551a62971280a8d

Observation 54e5d202-bf23-46d1-98fb-d3bd45514b1f · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.379995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:376eabaea604c52073103e59e2824880755330cfe9db525f5021b049ed223329

Observation 3b7dfae4-d4b9-41b7-8a07-ff1bc762f762 · inbound

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning cites this paper.

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.879950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:53:53.520545Z digest=sha256:815b6f5a647c27ea5b2a9d2c05a84f263a44d5987bd75400a33b4ff8794d9ffe

Observation c2a7f144-34cd-4dad-9773-a3100e589982 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.979640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:996905bdb971a6b4996fba955529e3d42b7cf5bfd4217b587c519f104f49f5ec

Observation 0b1ceda6-c948-4963-b7a4-e89a65659738 · inbound

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning cites this paper.

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.125759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T11:13:49.266859Z digest=sha256:82c71377ab62a4222aaf45aa9d3c4081a6dfa16679fea424bc9b112b125b8682

Observation 2bf24e0f-0851-454d-97e7-f1d6c87a3583 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:06.411854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:52176e9267fcb8f7b08d68bf8189c7f43c2b543d36b16e6d492bdda2d95806c8

Observation c1801018-9997-4953-8a27-fcc10b024224 · inbound

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models cites this paper.

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.685914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T13:44:48.250838Z digest=sha256:1f3e5d5f1cc8783c4b65c387f6209faae0dc2e5905eaecc0abac6ab66000f20a