Pith. sign in

Paper Citation Record · LEDGER

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.03916.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03916 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:04:14.254349Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c395a067-53ef-4ccb-b30a-dc08fa59f74d · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Docvqa: A dataset for vqa on document images,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:19.199956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:09.955672Z digest=sha256:627ba8d4d45183064f268501ab1f799c9749f1f75c680d25c13522ba85fdedfa

Observation 1b28438b-a3fe-4128-8c23-9f9353e9a0cb · outbound

This paper cites Infographicvqa,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Infographicvqa,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.920954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:10.072039Z digest=sha256:5f4fa92b9693144068a89f422e147e1e27add7d6c2f33bdc664add8f1c3919a9

Observation 2df9f805-09b6-4fb0-a073-b514c5a137e0 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Slidevqa: A dataset for document visual question answering on multiple images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.704656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:10.189362Z digest=sha256:08066337c244fc0032e2287683fbb97ba0156ce0ddfe37d0a01961ab773dab06

Observation f4563e43-4b88-4be2-8990-a3fceb74b5aa · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Msr-vtt: A large video description dataset for bridging video and language,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.495421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:10.342478Z digest=sha256:fdb99387d404fffc9243a7782afefe1010a03e67a1920c0eac6a1a639d2bbbe9

Observation 08d47dbe-8fcf-45b2-995e-772c54db6ee1 · outbound

This paper cites Dense-captioning events in videos,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Dense-captioning events in videos,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.280657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:10.475082Z digest=sha256:ed383aac5e5608e63b15b67f386bfcdec4e11c8f6bd8a3c743937fbf4c777d0f

Observation 5e442f33-eac8-4dd1-8aa5-3fc047c20570 · outbound

This paper cites Towards vqa models that can read,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards vqa models that can read,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:18.037729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:10.636457Z digest=sha256:3ccc5185116e2319cb763ef48a31ef43767150801173b1611c94a2129cec9248

Observation 8f1a2db1-2a3a-4add-82ed-3acfec3fd233 · outbound

This paper cites Qwen2.5-VL Technical Report.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qwen2.5-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:10.751674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:10.751674Z digest=sha256:818efbb478f0c876b7c73a3537f8bca131f0143d4ee1089d616f093364ebbf28

Observation afb972a4-28be-4dd1-8bba-bf8a6557453c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Lora: Low-rank adaptation of large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.727409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:10.850014Z digest=sha256:1aa30e07356536e26a1b760a3ca51169903bb4cff1ea1aefe181c1fccffafc23

Observation b5a2dbb4-ea70-4edf-a9e3-95b4227a3f6e · outbound

This paper cites python-pptx: Create open xml powerpoint documents in python,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models python-pptx: Create open xml powerpoint documents in python,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.400977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:10.980238Z digest=sha256:ebe4518c613e822fb0e1964e9c86114c1fcc819c961883f5bc19bc01e28bca7c

Observation 33776ba4-c58e-4dfe-86d9-32512bd56db7 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Bleu: a method for automatic evaluation of machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.119793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.119793Z digest=sha256:2dde9c544f6810c95b2107ae808092e9a9c19c99c7ece1acb90858ce002118cb

Observation 7e77378b-d7c8-4df0-bdf9-81e1c5430b33 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Rouge: A package for automatic evaluation of summaries,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.237427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.237427Z digest=sha256:d8951c584f47e820b1f435ec5912b51dfb0e10acf02a0038491c399879fdb40a

Observation 4fc245fc-f4d6-41aa-bc62-5fa3879789f2 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Spice: Semantic propositional image caption evaluation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:17.185503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:11.361447Z digest=sha256:5ee736f84b51ed3dad339727c43aaf7455fef29d620f2a7c6c4959a3c8b6350b

Observation c8eb1cfb-6104-433e-9f23-ed5f6634ba94 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.505717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.505717Z digest=sha256:d8dc537ef60827920508864ec7ac53a1bc763d602c4a4249f2ab0320a743f8db

Observation f3d8a2c0-3142-4172-9299-1b0e655d37b1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.918071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:11.654830Z digest=sha256:d45e550b73c20a673b7555b2f4355d191ef4f81e989e7eec0b877af78f702cdb

Observation 25579411-a4fa-4a53-8687-69c83adccef3 · outbound

This paper cites GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.787732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.787732Z digest=sha256:d6123a184c91c1a9c9a95353cdf45519aeeb1ea9194a1791cac1cc2aca28b162

Observation ea5f9202-a5dc-4ec1-b97a-393df857d521 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:11.888949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:11.888949Z digest=sha256:b31bbc9ab3eb3762f238bfc3d7773b41377c0a76988126fdeb81be3059f9b324

Observation b39d791f-4869-43b5-95af-3875843b3abc · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Flamingo: a visual language model for few-shot learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.035982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.035982Z digest=sha256:0dff48652127714951b4478c61002504b61b32d7df82694a348ae33cf259bfec

Observation 32d36aad-f476-499a-bdc2-7e0f6b44d1e6 · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlm: Pre-training of text and layout for document image understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.638138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:12.156363Z digest=sha256:536a9f42c7e78b7f203b303ca702eff1ab701642724e37bd37c12101ad9c5863

Observation 3ceb7149-7d64-49fe-b295-13df1a98afe5 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.315928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.315928Z digest=sha256:661659b5d0c77dc9ce1831115d647154971fffa53a226b368c3797aceb559ed0

Observation 3a9ecb96-d720-4380-80fc-6464a625a2d8 · outbound

This paper cites LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.416233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.416233Z digest=sha256:952c7a83d96f3d02688dcc8da1897da6b7e797c718b25cee7db2a607c1520c52

Observation 93069f4d-ff76-4b0a-a1f6-fc35d3299d62 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Layoutlmv3: Pre-training for document ai with unified text and image masking,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.381185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:12.517861Z digest=sha256:54c68cb0cace34430dd82c787c3b3b46961a49289c590d277843af128c12fc25

Observation 5d641d5f-7ee8-4f4f-80cb-971c6b5ceb37 · outbound

This paper cites Unifying vision, text, and layout for universal document processing,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Unifying vision, text, and layout for universal document processing,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:16.189850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:12.637123Z digest=sha256:6a5bdc49a0f64c2d9167fac0f3958b065fbaf1332263476435eb24fb1bffd55e

Observation ab129cd5-8c04-4f1b-875a-cd566038aa19 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Towards automatic learning of procedures from web instructional videos,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.953547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:12.751843Z digest=sha256:beb30d5f675b5f996830ced230365e2d68132773b36e4fa4fbdce09584658aef

Observation ce4ed28f-4ecf-4329-80cf-bd6c990a50dc · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.865362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.865362Z digest=sha256:7418a4969707ad17fbfff8bf3b4545b47f4e1d63842a11301f01f2e8e9a59e81

Observation 28e60d0b-1127-4e64-9a6c-cc9b0317ddf2 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.002034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.002034Z digest=sha256:ce61a1c308f79936968d8b44cbc715d3c014f355271e69f080d77071896df3f0

Observation 9d6141b0-04e1-40ac-91a3-b340e99843a4 · outbound

This paper cites X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.646879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:13.153091Z digest=sha256:cd1887d4937db79dd20923f0af7c9fed978d27a33620070b3f9ef80413372cc4

Observation 8f9c5f80-2681-4dc3-aed7-67647c00322f · outbound

This paper cites VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.314128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.314128Z digest=sha256:26831357d8697380e01609aba36261eca581e1475f31df58a7fb155526b00b83

Observation b633c2c2-2d2e-4609-b6f5-f9cef1479bbb · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.425644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.425644Z digest=sha256:ad2f78d290d15ac9d6d3ce8ea9d33d33a471153dd88e921c91458ced96fd297c

Observation e2191883-4d52-4f65-9ef9-e8cb41d415a1 · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.503234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.503234Z digest=sha256:905a335ba149b8609b3ca43cfd4db0a200b21f844f4b7d4fe2a12df6c2e334c6

Observation a41c1694-56cf-4873-a39a-98715f7f79ea · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Qlora: Efficient finetuning of quantized llms,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.548641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.548641Z digest=sha256:cd9218b860f842c53cd1d5dd6abc48124c9053133a40ad1069fabe1f3ceb0981

Observation 241ac941-1d25-4a6f-9925-8e59ec464c91 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.653416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.653416Z digest=sha256:34b45a33a636982444286e38e7a1fb0db98ee916cde5a653d89c121c094a309c

Observation 516f1b29-e1f9-4c7b-89f5-87d015388132 · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.722971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.722971Z digest=sha256:977969e7c5af1f38c7680844dfc549fdf57343ba572e524f7ff400c59476717a

Observation ffec0415-ea13-4779-ae3b-7ebd73c1e1f9 · outbound

This paper cites Evaluation metrics for video captioning: A survey,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Evaluation metrics for video captioning: A survey,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.358915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:13.815372Z digest=sha256:418515b70768189f085b8d7fba5fa2d051b9f3e8c72aba03cc95234ef2a0f6d3

Observation 18ef3acb-5841-415a-b5a1-f8a70a19b3d7 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:15.061021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:13.893403Z digest=sha256:489d190b6038a724aadcd41a85847f27315d53ab0b6ce88a7de9fd8bc807b610

Observation 2f4cd553-01dc-4ddb-a5b0-af92db97c562 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:14.812717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:13.974323Z digest=sha256:d5e1a1920696fd630d8e632b79e434b2c4bcea8621906910a0e34cde206de093

Observation 0a774904-43d6-4660-b4f7-18abba67c15c · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding,.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models Mmbench-video: A long-form multi-shot benchmark for holistic video understanding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:04:14.672111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:04:14.050908Z digest=sha256:0adff43c27f35f896d3d345b0af6705b5e68adcf6d86c4b7cb074e57167e8f01

Observation 618594b0-6d38-4c7e-a528-0ed42817d561 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models TempCompass: Do Video LLMs Really Understand Videos?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.128716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.128716Z digest=sha256:55baf2b5a9fc8b038bc7886581272d0df40bde54861397bba3b7c0aa887d9f32

Observation 0b657244-8cae-4f8c-bbd8-ae81bfa0b2b9 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.186124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.186124Z digest=sha256:b17af49aaecdcce7924cef0c3c892e898b6a971f90dc791de5c1a3a3b95e02c9

Observation cbae45d3-4c04-40d1-9178-959b0eee45f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:14.254349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:14.254349Z digest=sha256:8b4e7b8eea4363a75aa6c5a71b7da5bd9b39f9fc841b7442b31171a0632f36b5

Pith citing papers

No inbound Pith citation observations are available.