Pith. sign in

Paper Citation Record · LEDGER

Valley: Video Assistant with Large Language model Enhanced abilitY

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2306.07207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.07207 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.158259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.409899Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 897e4d1c-875a-441b-b2f8-26c72af5d579 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.650439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:a6676b9d934d11caeac43a69baac5d4367411047109a6cd3c1d60c1515b292fb

Observation c342be09-1bd7-449d-b81d-5e098411cd7c · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:08:01.242531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:276a9a9a5cec38292c998faf7583c8958a991a764f8451b8ccdd8d647d88cd45

Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.014045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:8bc547326586b5694e5491b15146eaef258ebbb69f7835cbfb5ad9a40801376e

Observation da2205a2-3c68-494f-bd58-63204b3bf502 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.769669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:817bce307774bd4b8bc81510036c3a73fc349577c7047f2c1bc3157e696a51fc

Observation d8272357-eb90-40ef-b937-06bdfc05800d · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.884272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:958c742bcd1cc2f05102cefb6e3547383ca4fbfe373656fb1f05650545f47012

Observation 634a0eee-6bd6-4860-99f0-19f0893c45f8 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.675561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:4d37eadbc6ea731930e5f5327b83a325b3e37c954e99067e303b973119845500

Observation dc8993eb-df65-4539-8b6b-4bd24028e0e2 · inbound

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance cites this paper.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.706077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:4ed374ac5851429f2e8fb9091683d3cd1bc139f81b1539737fa1bc4ad6a06fc3

Observation 8f5e0e6f-f569-4ead-809a-2a3b4e6d6d3e · inbound

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos cites this paper.

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:12:43.870984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T08:08:01.889675Z digest=sha256:2d2a1859ffe5025a5bab926b21fdce277e9a63c91b0cd7b74a12eeea45ef8bdb

Observation 0e1d8574-032b-4117-ab75-eb7050bd2c38 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.158259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.158259Z digest=sha256:55898a1d42a3e9374399be64bfe9c57aef152613bb386c0e3404a272d2bdbe97

Observation 08271f24-be62-4741-8711-197e91098711 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:22:18.746828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:25f9a09e67593e0753470a058e0eae02c6ab8c6a0167dad27d26a14cb903182b

Observation b589f826-7aa3-4d32-9b16-895eee05959f · inbound

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling cites this paper.

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:22:51.754421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:22:51.754421Z digest=sha256:51e81142aa9cc3af8f5260091b2a3f108757239c849a83c48da283ea4be9af35

Observation 13fc8891-7b8a-4efc-b6b5-fc1689b8a69c · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:13.167183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:13.167183Z digest=sha256:29f4af09c85fba7ca07380440ad064a306f1db09b38e26048638c09bdce796b2

Observation c1b43079-4c42-428b-a47a-aff1f18e0f89 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.929275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.929275Z digest=sha256:b76e2bce2c51f5db7fa843d5c5ed8425e4349e924c70ef1edcfeeeff892656f2

Observation 1646c9f3-a123-4005-9125-201509576e94 · inbound

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model cites this paper.

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:38.671225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:38.671225Z digest=sha256:d804b94b5fa2fe5ab0cea4e1e12491c0abd523ec61494defd2708d1a6c07ee97

Observation cf23a219-89bd-4daf-9ae1-0fcb6e0aa809 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:29.648522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:29.648522Z digest=sha256:84121973f90c98d7d112d1ba639cb503ece0eff09aba5d775821b495cb424947

Observation c4229669-f0ae-4e90-a50f-ac08f48c9807 · inbound

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding cites this paper.

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:47:10.369499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:43:26.736409Z digest=sha256:0258b2ef13409fdc584263502ceb36b13172294a83ca72a55146d3105a18fd60

Observation d25c8878-8950-4f13-a1ec-3ad356218ae2 · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:37.503819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:37.503819Z digest=sha256:88183b15e555f1f11f44fd8a24bee8e5044ef7d976b4dea0271daa518dbec059

Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.176899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.176899Z digest=sha256:e59b84d1e6eb2f449335d05e9aa59ae526ebb82a1d4f0a48cec5441457c84be4

Observation f21d3b11-8dd8-4f2a-8183-1df5bf4165df · inbound

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding cites this paper.

TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:35.652829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:35.652829Z digest=sha256:1bcb8854f778ed251fbf4eb1b186bdb24c6af2fa1d82dfab7ea7542bbf9f6fb4

Observation a6630497-dc85-42bc-83c5-efaa1c28942c · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.875262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.875262Z digest=sha256:44d1025c329e69fa5087cc67d12793641f336fe87c9c18e4f6442b3c8ace84fd

Observation 34b2886f-ccd6-447f-b1c7-55ef0e380930 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.248728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.248728Z digest=sha256:053cb8952b5e31eb09b189f605d68042029aa885f207fa03ace2412ac23ba732

Observation d5431206-c977-4ddb-9a17-869ded7abf2b · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:47.855446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:47.855446Z digest=sha256:e419f533fb06a214bb9417c367e846a10fddc5d5b86842339d330a5c1bcf9672

Observation 0536ee5a-5b72-4bb6-b5f3-d8f5ecdfcc87 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.906964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.906964Z digest=sha256:a286f55f73d6ae7878c7239187d5377677e31ee3219431962a8c7c7d3c5ad94b

Observation 3780dad6-ab2a-414b-8f9d-9e81fe75ccdc · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.556812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:750cb55fef475099d0f0c417da0bdff2273392620cf302228fe557085268be95

Observation 89f4997e-db72-495e-a46d-391c9c729afc · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.044127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:dfed842e699668b86e9b82576617d3e1e222dda4ff048a4f5772d03a5e8b78ed

Observation 053cc833-2807-415b-b9d3-2620cdc741e8 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.924293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:928316a761b60f2ef63065c9ae4450235bdd6a41d3866c96bee68c848b401b7a

Observation 6d2e3842-bbb3-4e92-84ad-800e932d1c89 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:f14c78eadd510a4ded7589e5a4b22a2fe2d82946781fc2799ba1010a43a4040a

Observation 6ef3b4ba-347d-4592-8f41-d38510568d4b · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.566195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:addb06d08d39e2dfeb8615906e957cf3c4c49bb17d58188b37c85c564fb6ac98

Observation 9a223249-f3da-42a3-af03-839e912aeace · inbound

ClimateVID -- Social Media Videos Analysis and Challenges Involved cites this paper.

ClimateVID -- Social Media Videos Analysis and Challenges Involved Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.448614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T05:25:05.760276Z digest=sha256:6865c4fb54b5e5f80e748707e5fcf7216c96ab1897a085a8d8d17b91356a2e98

Observation 696d60da-b03f-4d11-a459-80b296cf68c3 · inbound

Dynamic Model Merging Made Slim cites this paper.

Dynamic Model Merging Made Slim Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:18:25.408077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T15:16:24.651868Z digest=sha256:75c86196ef4a327de5e57e54913a057ddc3dbaf0a3edcf6a6bea0f1141c417a3

Observation 54e5d202-bf23-46d1-98fb-d3bd45514b1f · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.379995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:cb60a915f7c956e679bcdcb8bfcf92dea78985bfbc410c852dfbd2f9beb4fe67

Observation 3b7dfae4-d4b9-41b7-8a07-ff1bc762f762 · inbound

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning cites this paper.

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.879950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:53:53.520545Z digest=sha256:4fa7f5e58dd369ba5f983203dd04686f7e573251c8f8c76791f530d7306321c0

Observation c2a7f144-34cd-4dad-9773-a3100e589982 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.979640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:1719e1e0436e336d230c853c1e380dbf65418cdbcb00663a0712eb3bdaf806c7

Observation 0b1ceda6-c948-4963-b7a4-e89a65659738 · inbound

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning cites this paper.

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.125759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:13:49.266859Z digest=sha256:d92a8e19051a30e4e2937edc277a1868210316032c9b0aa67c8c4c3ea9d8a9c5

Observation 2bf24e0f-0851-454d-97e7-f1d6c87a3583 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:06.411854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:a3569b5239e6ce62706fad50a9698881c5d64144a034fc1d810353f5e58dc34b

Observation c1801018-9997-4953-8a27-fcc10b024224 · inbound

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models cites this paper.

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.685914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T13:44:48.250838Z digest=sha256:4b60c96ef4aaf73d123dabce25b952466efff71acfd3a95eb72142fe13dd28c7