Pith. sign in

Paper Citation Record · LEDGER

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.03916.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03916 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:04:14.254349Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c395a067-53ef-4ccb-b30a-dc08fa59f74d · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Docvqa: A dataset for vqa on document images,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:19.199956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:09.955672Z digest=sha256:8ee4ae9cbce6c46367c3c987880e3bca998c29205c6a9cdce3a0757b78cb7cbf

Observation 1b28438b-a3fe-4128-8c23-9f9353e9a0cb · outbound

This paper cites Infographicvqa,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Infographicvqa,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.920954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:10.072039Z digest=sha256:bc7ad6867f671c8960bd60b632952404a097cebd8c232f3f11693e850a37b0ee

Observation 2df9f805-09b6-4fb0-a073-b514c5a137e0 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Slidevqa: A dataset for document visual question answering on multiple images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.704656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:10.189362Z digest=sha256:d4d84cb288f52e4b28d2c09f93b61559fcd9305831dc21ed13a6c76310d43a8a

Observation f4563e43-4b88-4be2-8990-a3fceb74b5aa · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Msr-vtt: A large video description dataset for bridging video and language,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.495421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:10.342478Z digest=sha256:8e54ef4afe7d1a0a8a94caf07a165fb6b52e83c32ef867d34a6caaac42dbe33d

Observation 08d47dbe-8fcf-45b2-995e-772c54db6ee1 · outbound

This paper cites Dense-captioning events in videos,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Dense-captioning events in videos,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.280657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:10.475082Z digest=sha256:8195edbb3078d90fbab970eb7aea8b2973894c9f501505fa044bf57212c43f75

Observation 5e442f33-eac8-4dd1-8aa5-3fc047c20570 · outbound

This paper cites Towards vqa models that can read,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards vqa models that can read,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.037729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:10.636457Z digest=sha256:5034ffdd261ba3bc20a233c3b878a06da0f9665ed455e6049f79bc018ee28393

Observation 8f1a2db1-2a3a-4add-82ed-3acfec3fd233 · outbound

This paper cites Qwen2.5-VL Technical Report.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:10.751674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:10.751674Z digest=sha256:dd33319eb500cc133d3c237b6ffadb876ecbce715918d3a0ad8912fabf84621b

Observation afb972a4-28be-4dd1-8bba-bf8a6557453c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Lora: Low-rank adaptation of large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.727409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:10.850014Z digest=sha256:ea212c71933fe5ec574f6d4e0731ff03c9c38792159b64053087b94a71f210af

Observation b5a2dbb4-ea70-4edf-a9e3-95b4227a3f6e · outbound

This paper cites python-pptx: Create open xml powerpoint documents in python,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models python-pptx: Create open xml powerpoint documents in python,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.400977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:10.980238Z digest=sha256:bac27fceececceae02d585fc06054fdf058b998a757078d7d4991844c4ed880a

Observation 33776ba4-c58e-4dfe-86d9-32512bd56db7 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Bleu: a method for automatic evaluation of machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.119793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.119793Z digest=sha256:f1827d156741dd5fd947dd43e6575472431ca0c0ec3cd57112b905314264d93d

Observation 7e77378b-d7c8-4df0-bdf9-81e1c5430b33 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Rouge: A package for automatic evaluation of summaries,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.237427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.237427Z digest=sha256:ed3bf24ace21ffb4368990ce53b25bfa84b5c883833258007dfffcf22b31f5af

Observation 4fc245fc-f4d6-41aa-bc62-5fa3879789f2 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Spice: Semantic propositional image caption evaluation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.185503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:11.361447Z digest=sha256:98bec6db5854eea7d87e01b9a0e55d137f0c10a26a25ddba3344fb76d759caa1

Observation c8eb1cfb-6104-433e-9f23-ed5f6634ba94 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.505717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.505717Z digest=sha256:dec44f9296e08d5964e99447d67318a022216037e05da020d0160272afe3ed90

Observation f3d8a2c0-3142-4172-9299-1b0e655d37b1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.918071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:11.654830Z digest=sha256:ff87f2cb5fc5dba39aa66c9e3df4d8aa0622f28e39386a6fe2f1f1aca89a86d4

Observation 25579411-a4fa-4a53-8687-69c83adccef3 · outbound

This paper cites GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.787732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.787732Z digest=sha256:9ac167c9b53eb4bdfbf3fa2d73ef3f46fa18d14b56775f666065118583473c73

Observation ea5f9202-a5dc-4ec1-b97a-393df857d521 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.888949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.888949Z digest=sha256:63fdb791f889b4d689f143ee2cb4170d83633e295220e8e22d6949d3797a9020

Observation b39d791f-4869-43b5-95af-3875843b3abc · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Flamingo: a visual language model for few-shot learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.035982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.035982Z digest=sha256:595df134d08e434eca46a379b152cd64626ef04f845dfdd1662977c0d669a64e

Observation 32d36aad-f476-499a-bdc2-7e0f6b44d1e6 · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlm: Pre-training of text and layout for document image understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.638138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:12.156363Z digest=sha256:78efa22b55b5225f990c2ed16950a1495a346bf57ed70cdead224e9bce65b69f

Observation 3ceb7149-7d64-49fe-b295-13df1a98afe5 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.315928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.315928Z digest=sha256:dccfb5f81b2f2e49b94e17b08945950604f57beb419aaf367998ff27c178d5ec

Observation 3a9ecb96-d720-4380-80fc-6464a625a2d8 · outbound

This paper cites LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.416233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.416233Z digest=sha256:3c6a566b9317fa6857537c5468b45a14a79879aa648240bb382bb8bc83c21aa8

Observation 93069f4d-ff76-4b0a-a1f6-fc35d3299d62 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlmv3: Pre-training for document ai with unified text and image masking,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.381185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:12.517861Z digest=sha256:381628bec0deb19ac37e79e37ba1af5d3c9b9ae321cc8498c5f9a7f4667c7a27

Observation 5d641d5f-7ee8-4f4f-80cb-971c6b5ceb37 · outbound

This paper cites Unifying vision, text, and layout for universal document processing,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Unifying vision, text, and layout for universal document processing,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.189850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:12.637123Z digest=sha256:e496aad1a759880de5b659dd18b317c5094c81e28f18c769f441bce85d72cc36

Observation ab129cd5-8c04-4f1b-875a-cd566038aa19 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards automatic learning of procedures from web instructional videos,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.953547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:12.751843Z digest=sha256:469d4397030d88a523a070979ad6294794f90579544d2d4381493fde76c0e93e

Observation ce4ed28f-4ecf-4329-80cf-bd6c990a50dc · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.865362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.865362Z digest=sha256:b7c658e73041df52744219493c3590c454aa542b608342b2219ba202d2ecfaf0

Observation 28e60d0b-1127-4e64-9a6c-cc9b0317ddf2 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.002034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.002034Z digest=sha256:6dd75ef3421b403b348ffa2ba2fdbea50736ecdb3cc3654f69c16ab668c2d8c7

Observation 9d6141b0-04e1-40ac-91a3-b340e99843a4 · outbound

This paper cites X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.646879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:13.153091Z digest=sha256:2dd50cd08975ed658d8edbcd8e957cc032a2201f3ad78dc09fa172d5d0f3fafd

Observation 8f9c5f80-2681-4dc3-aed7-67647c00322f · outbound

This paper cites VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.314128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.314128Z digest=sha256:beadf6d6832b8677ef35131188db8769a17e339e4bdf83bd0f1f35ca0a761e6e

Observation b633c2c2-2d2e-4609-b6f5-f9cef1479bbb · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.425644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.425644Z digest=sha256:aea3487bc2be04da3cd741be3ce39a4b758d1ef21478b63daadcde2fd19dd58b

Observation e2191883-4d52-4f65-9ef9-e8cb41d415a1 · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.503234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.503234Z digest=sha256:69b7991df53a313f479625005ac37927acdeddcc1a725084c98175e23373ae0c

Observation a41c1694-56cf-4873-a39a-98715f7f79ea · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qlora: Efficient finetuning of quantized llms,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.548641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.548641Z digest=sha256:f215c1199979f6755b3f5f1d154641d5a24795526aa927c63a5928ce9f9fce94

Observation 241ac941-1d25-4a6f-9925-8e59ec464c91 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.653416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.653416Z digest=sha256:ea1db29775eac084cb055b8c282994741eed991981fb9cebc123f6256f2a688b

Observation 516f1b29-e1f9-4c7b-89f5-87d015388132 · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.722971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.722971Z digest=sha256:00e9d819a800d6758bab26fbe6a250c1a12152057f78e3697807702e91ed85cf

Observation ffec0415-ea13-4779-ae3b-7ebd73c1e1f9 · outbound

This paper cites Evaluation metrics for video captioning: A survey,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Evaluation metrics for video captioning: A survey,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.358915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:13.815372Z digest=sha256:f4a79444e1c613dfe8cbc922a0a0735e8811162206635e313089febb37cd8569

Observation 18ef3acb-5841-415a-b5a1-f8a70a19b3d7 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.061021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:13.893403Z digest=sha256:4da32e5ba46435700cb70b70578d68b436f1dd9eeece1ab8a4510d353957f51f

Observation 2f4cd553-01dc-4ddb-a5b0-af92db97c562 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:14.812717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:13.974323Z digest=sha256:205b68b2ff3849fc3adeb7e62a5449a32a992e57a39ac3b3e7d4d6b04ffe13fe

Observation 0a774904-43d6-4660-b4f7-18abba67c15c · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mmbench-video: A long-form multi-shot benchmark for holistic video understanding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:14.672111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:04:14.050908Z digest=sha256:9f973203f3c2ff1f50ee817e35b743b7c5bd14ed1c5739b9404725d69f5f95e0

Observation 618594b0-6d38-4c7e-a528-0ed42817d561 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models TempCompass: Do Video LLMs Really Understand Videos?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.128716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.128716Z digest=sha256:1987661040c472402146a549572c7884319cdb2de3a316154e9ddf69fff3d8ce

Observation 0b657244-8cae-4f8c-bbd8-ae81bfa0b2b9 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.186124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.186124Z digest=sha256:a92700ba86445d69661f67a470c2a353e50f05c35f03433527398ced9a60120f

Observation cbae45d3-4c04-40d1-9178-959b0eee45f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.254349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.254349Z digest=sha256:f35caeec2e156515682deae43349d98b54fa35cb189ccead8605b57145501b26

Pith citing papers

No inbound Pith citation observations are available.