Pith. sign in

Paper Citation Record · LEDGER

PresentAgent: Multimodal Agent for Presentation Video Generation

As of 13 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2507.04036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04036 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:01:29.172482Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:41:58.268576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c888856-2c80-426c-9a26-0aae806c6ef8 · outbound

This paper cites GPT-4 Technical Report.

PresentAgent: Multimodal Agent for Presentation Video Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:28.433094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:28.433094Z digest=sha256:97ed637e89b4d97d130ceab71e4fb6dea7f68cddea2a6c9962c7c09e78197fe3

Observation 1e7c8d37-d7df-4fcd-8a75-838a4cb2c463 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.670856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:28.569354Z digest=sha256:3c4fd6ed6c581ef77df8f92701d80201e4e84e20fcdb3e0f1ec2c104df56274b

Observation 1983defb-0945-4889-b021-e1d98d195e20 · outbound

This paper cites Qwen2.5-VL Technical Report.

PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:28.715034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:28.715034Z digest=sha256:b3af3a5ca4741c0e73c95e20a1f78ca677c331ad3758ceaebdbf12dd32015d95

Observation 88a157cb-a991-44fe-a0be-9b921b80eafc · outbound

This paper cites Longformer: The Long-Document Transformer.

PresentAgent: Multimodal Agent for Presentation Video Generation Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:28.793825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:28.793825Z digest=sha256:308fe03b20fc59b84279c2e4b8783f8123b6b6553556c70f6b09b9b645d6eba8

Observation b8208639-f3d4-418f-b57b-3510eb76ba11 · outbound

This paper cites Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs.

PresentAgent: Multimodal Agent for Presentation Video Generation Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:01:29.561919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:28.960148Z digest=sha256:556358ef75c431bf04629f8e877f4195a05cd6b970f3e72b37555609b0731866

Observation 0ed332b6-4134-4a7a-a66a-3c1564009960 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.664601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:28.993820Z digest=sha256:21fb384fabe31cada67bc2107407b081ed37db7c4a73ab28d942f2606133cca3

Observation 224f5bd7-7d46-4e52-9b9a-53f049291259 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

PresentAgent: Multimodal Agent for Presentation Video Generation Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.043152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.043152Z digest=sha256:f51b717311aa8a4197a5268ca80e7309ea9a52a8d68a4583e89e47fc96b09836

Observation 1a67b510-e71b-4042-ab2b-505d9b1d4613 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.658073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.047123Z digest=sha256:c648a71318124416427ba8e65c91e38cf6a1aedf1afa2438d110a575ff04bafe

Observation 94db5d95-d83a-4d52-920d-7c92e15c56fa · outbound

This paper cites AutoPresent: Designing Structured Visuals from Scratch.

PresentAgent: Multimodal Agent for Presentation Video Generation AutoPresent: Designing Structured Visuals from Scratch

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.060673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.060673Z digest=sha256:8a9eb8369e74f76f1dae823de83e186a09a4688e6200699b9b23f0c846c7219e

Observation 26aa9c68-8bfe-4f3d-8223-5b4d8b68b925 · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.063984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.063984Z digest=sha256:b735eb6b5967ef5d1abc5f8e6efc5d4698e53f3cffe5f2184efea9bc3f53eee0

Observation 9e819c4f-dea6-4ea6-938f-68a682b9c7a7 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.651742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.066441Z digest=sha256:7bf5b8c2c3424ba0b7a754b8dee992047591b05464a0eb7ea932dee8372e2a2c

Observation 7432ab6a-fd75-46f2-abbb-4d672fb2d998 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

PresentAgent: Multimodal Agent for Presentation Video Generation BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.068517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.068517Z digest=sha256:18092cb3ddf1ef321c862f2971fb1070a9ebd19400a8a14280e821fa6c0cff43

Observation 539fab81-7e54-484a-800f-50c400b1b6c9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

PresentAgent: Multimodal Agent for Presentation Video Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.070705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.070705Z digest=sha256:1c59a6a963f2333d472f48ce71a53116d970f208f7aba0b347aaa00d3a2caabd

Observation d594df54-74ca-4f9d-a7c0-8ed5ff2fc62f · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.072774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.072774Z digest=sha256:00113a2d2ee04dc90cf12d53a68939fa6497233584849eb0ff0333cbb6f0e58c

Observation 2eee608d-98fc-44ba-9e0e-63a0d028d367 · outbound

This paper cites VideoGUI: A Benchmark for GUI Automation from Instructional Videos.

PresentAgent: Multimodal Agent for Presentation Video Generation VideoGUI: A Benchmark for GUI Automation from Instructional Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:01:29.519212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.074771Z digest=sha256:050e4f733f04e7280749a01ef3cc3c161386cfd708b42b4ee84cfa1323e38f0c

Observation 2fd21acb-63aa-437f-97d1-f4854ec9f7cb · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

PresentAgent: Multimodal Agent for Presentation Video Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.076779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.076779Z digest=sha256:027a8d7b18934811eb6ef8b76677668c9452265407750ac6b92aa52633764a66

Observation b5e48018-5dff-4b98-a175-551b82a549ee · outbound

This paper cites FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion.

PresentAgent: Multimodal Agent for Presentation Video Generation FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.078845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.078845Z digest=sha256:cef4f102f2870b044ff849e471799bdf8cc94d88dc32cf010516a32ebd9d1b08

Observation 91da6878-55ea-4bd4-95f4-0c90c856511a · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

PresentAgent: Multimodal Agent for Presentation Video Generation OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.081498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.081498Z digest=sha256:2ebbd0ac292077eb6514de2e6d8b0ed5c50a73f091ca11096a2de5eb310aa6d1

Observation 7d0b4b6f-1e3b-4012-9595-31c2f688a340 · outbound

This paper cites UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction.

PresentAgent: Multimodal Agent for Presentation Video Generation UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.083800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.083800Z digest=sha256:6d6a5ec1e3bb445dd32dd3f9e0dbe00693edd88c1082dbf9592b0a20b832e6ac

Observation 376c5d22-5ea4-4e77-836e-2d90a3d977d4 · outbound

This paper cites Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition.

PresentAgent: Multimodal Agent for Presentation Video Generation Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:01:29.486456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.086266Z digest=sha256:72ee9013d2b60062f0abc89962a65ddaaf5030c1b4ae17f77af5f7d2ce1c7aaa

Observation 211f0bf1-80de-4007-b261-e93d1d319f58 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.088530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.088530Z digest=sha256:72dc022871e09d7cfd1fb02f26d71a21fce9b52bcb7c5e31699e46d14ebc5e82

Observation 2fe2d6b6-f18a-47d0-80e3-05fa14eef9b6 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.645601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.090591Z digest=sha256:84194c95f0a9ae98ffd24561b5e0e0b53d0a837126660a971fc4e9096d8ce539

Observation 7d91f3dc-3361-40e7-9c90-5fae0056b1a9 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

PresentAgent: Multimodal Agent for Presentation Video Generation UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.092791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.092791Z digest=sha256:53b2473995bc20cad9ec7569e4970afa832229d6b3803f71a84dcbe6fd41e24e

Observation 45e9d53e-ec2e-48c7-af64-a1abd50481d2 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.095189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.095189Z digest=sha256:5d5b6ab2568a97f35b2b1c85bb8a5bdd30e2579da2ce08f8e3426e533406a070

Observation ea3df4ac-9fa5-4d50-b165-978d2dfe9784 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

PresentAgent: Multimodal Agent for Presentation Video Generation Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.097148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.097148Z digest=sha256:29d4f686eb9fa338afe65bc2fff1d7e515f7376751022e49d698e89c2d46771d

Observation 3d92604c-22f3-4e69-ab1b-a0ab736b2cec · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.635684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.099566Z digest=sha256:c87d79000d167180200ecaac9bc93b044cf791921200c3504889fc2e8283dee9

Observation 3c24cbcf-69cf-4df7-b40a-5d88ccaa7dba · outbound

This paper cites Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models.

PresentAgent: Multimodal Agent for Presentation Video Generation Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.101657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.101657Z digest=sha256:554c448217c175cce08f1b64940737b754ba658e9cc195bb068a5ef027833145

Observation 86bda43d-60f2-4d9a-8427-55fae60e7051 · outbound

This paper cites Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies.

PresentAgent: Multimodal Agent for Presentation Video Generation Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.103900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.103900Z digest=sha256:8b75462313e20124bd5cc05d1d5aa1994317ca32d26cbc31d3bf5dd81992fdf7

Observation 21327823-47e2-4976-a1c2-77615640adec · outbound

This paper cites ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models.

PresentAgent: Multimodal Agent for Presentation Video Generation ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.106548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.106548Z digest=sha256:d3f7d10d6281c126d731c96edd5da2e976fdf913d915d5cbf85a435cc65ae28d

Observation ce08266a-9130-4271-9094-73e58b3c01e7 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.108567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.108567Z digest=sha256:747e964651b9710835e11609f9143a5f210c942de84c50f571c8fc969ab58436

Observation 16049078-62be-42d1-a50e-9ccac96c5712 · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

PresentAgent: Multimodal Agent for Presentation Video Generation OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.110892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.110892Z digest=sha256:9c11552a56492124fe37201df9ccaae0230e695384cf56b75978707a4df2251a

Observation 01355dde-d6a2-459d-85e7-52f3ae7a23d3 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.629435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.113166Z digest=sha256:81aeebff8750e38356c8be01b99332ea52cb655a5cb28d864f9702183a5e24ee

Observation 67b53045-bf75-4f75-8880-745f237be6e1 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.115335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.115335Z digest=sha256:6b9e9bcbc3cc7e0ad2119511373d7e160881f762876351e5290480525d6db84d

Observation ae5a6918-b9a4-4818-842d-42d62a5bbd7c · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.623264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.118180Z digest=sha256:99c89f0640a2ff9e0a3f7559a3b33ba1ee5e68387ffb2a9c331534a210dd1708

Observation d39aec71-f6ff-46c5-b807-f5cc87ce74b1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.120314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.120314Z digest=sha256:8fa3c3b9142f8cfb12819bb5ff7a8f6de9515d0650ae19f4b4a784b7aa562a20

Observation 49e8f1e3-eb15-4360-b7c2-2deda2d51ef4 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

PresentAgent: Multimodal Agent for Presentation Video Generation OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.122704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.122704Z digest=sha256:fe4faaf1bf3c60f084fb16e0876ae9b3dfa1508066ed005231134762b98e6eb4

Observation 12b6aed6-28ce-4f33-9182-757289dc56b1 · outbound

This paper cites Foundations and Recent Trends in Multimodal Mobile Agents: A Survey.

PresentAgent: Multimodal Agent for Presentation Video Generation Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.125527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.125527Z digest=sha256:82cd5774fc0e8de7a3a72567af75389972ffe4de5c3a38434b3aceec364e0d33

Observation a53afc44-f45d-41ad-9a2e-36e6e8c9b45a · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.127674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.127674Z digest=sha256:803ede69d656c178782c24f4a5179b8596fc89333b263cf703954fc06fd62c8f

Observation cfe0fcee-d956-44f2-9803-31d2e0e4053d · outbound

This paper cites Qwen2.5-Omni Technical Report.

PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2.5-Omni Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.129735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.129735Z digest=sha256:32083fc7893083d0fa2e6472b53fe7966eb15c02cc4ad289b3e2bc7d7011fcda

Observation a46645cd-3d1c-4db9-b8a3-aff88a0a87b4 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.616942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.133015Z digest=sha256:543e3b4e58c98d0f68c52b7d36752ad030e68184185547f5c5adbe426e1e340c

Observation 3ba96308-d4b5-40ef-8639-ee575a2b19da · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.610446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.135592Z digest=sha256:d24fe7a0c5b7c11f8f0ef0e68021050b12db1de8d191d0a81d45edc2e076b53e

Observation 3c23a90b-6986-430d-86d7-7b7ac803d0b6 · outbound

This paper cites If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents.

PresentAgent: Multimodal Agent for Presentation Video Generation If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.138233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.138233Z digest=sha256:e89c9e997faa2bdca05cac3d339e05397df6b15f91dd06cf1aceba6ac2713d72

Observation 093e1986-e296-40c3-a1b8-2eba82a8abc1 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.604000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.140908Z digest=sha256:1888b07672827da7c92e6601c52d0437a195e62f79e998c37d1082766ad6bf62

Observation cb17e108-8c89-4af7-bf13-994577aa12d7 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

PresentAgent: Multimodal Agent for Presentation Video Generation MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.143169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.143169Z digest=sha256:15da7ea2c6a3a3995c7d8cfb7da401415a50ffd98baa933e78168e708a9facec

Observation 4529811b-4832-47df-9ae5-e4180b04240e · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

PresentAgent: Multimodal Agent for Presentation Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.145529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.145529Z digest=sha256:74132c7105186d8c31779a0acd6952413645eea9dd9e07f9c804aba7d0a8719c

Observation a0e4689e-5467-46cc-959d-1fc467deadbe · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.147965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.147965Z digest=sha256:6843914b29b7b0a02b2842cfda3620ff0aede132ce079dfafbc0866a52203e79

Observation 88ac403f-f968-4ebe-a156-1159ee29b95b · outbound

This paper cites DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search.

PresentAgent: Multimodal Agent for Presentation Video Generation DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.149835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.149835Z digest=sha256:a49381bf17b2af31660ee4522530fb1a6eb50d2ab3ebb9258aff3d60b9f7c2d1

Observation d525c7a2-99f6-4cc1-9d96-938a6d869311 · outbound

This paper cites KMM: Key Frame Mask Mamba for Extended Motion Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation KMM: Key Frame Mask Mamba for Extended Motion Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.152345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.152345Z digest=sha256:0ca5952940b65cb7bc67f67b95ee52d6cb3cd5fd970d73f8ca6925067421a944

Observation cbe9c6da-fb73-4065-b58c-08377835204f · outbound

This paper cites InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.154828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.154828Z digest=sha256:a2cccd13381df9d94309d0d159276f374db47c20f9223e44fd4594dc0c03517c

Observation 7cdf0a49-32a4-45c5-a7c1-f182fa2ceabb · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.593307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.157288Z digest=sha256:8f1c58ffb1aa0df4c33ab89f636f25b5a640f7f86c445e1965c2053e6ace4cc3

Observation ef6585ef-abd2-45ff-a142-5324e57784e0 · outbound

This paper cites Motion Anything: Any to Motion Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation Motion Anything: Any to Motion Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.159405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.159405Z digest=sha256:ae4f73149e3bb093ac0ebaf36f11ba70a99772a75b95a7ca9123ba351cb3011f

Observation f1da34a8-069d-4838-b864-111ed8088c27 · outbound

This paper cites Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion.

PresentAgent: Multimodal Agent for Presentation Video Generation Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.161584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.161584Z digest=sha256:df3e8644368747422b0daf96a9f2e4b36ba2e2035adec0a32696beba50b68310

Observation e687567a-663b-4744-babe-7ccae67079fb · outbound

This paper cites PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides.

PresentAgent: Multimodal Agent for Presentation Video Generation PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.166158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.166158Z digest=sha256:879a6ae3e4b43e5f85b722e979cc4c990b3fbe04b315cf44772cc1bdd9fa607e

Observation fdaa49b5-5483-487c-a631-1e9abc9c2c40 · outbound

This paper cites online" 'onlinestring :=.

PresentAgent: Multimodal Agent for Presentation Video Generation online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.168686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.168686Z digest=sha256:41ceae0270dd3a14a89e4bf350f0de723b8ec781f50aa7771c47722cfb9ffbbc

Observation e030cbe4-6ff5-4983-8a81-176fb28365b1 · outbound

This paper cites write newline.

PresentAgent: Multimodal Agent for Presentation Video Generation write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.172482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.172482Z digest=sha256:119c42e1a03eca6696abd2b479f4894c17ae8534669d69bb74dd777a4303e277

Pith citing papers

Observation 7111bdd9-9630-4eaa-b5a4-aa5db3658a83 · inbound

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation cites this paper.

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation PresentAgent: Multimodal Agent for Presentation Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T19:41:58.268576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:41:58.268576Z digest=sha256:ebf7423e684e6bc6dc9a321d88b26a6fecf03482553f11e65369ba1f3819f8f7

Observation 95488604-1410-474a-8483-f99895a38273 · inbound

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers cites this paper.

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers PresentAgent: Multimodal Agent for Presentation Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T09:30:41.159870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:30:41.159870Z digest=sha256:efaef73ac91d540d6267cda9b6654a77280b96690b2e4514eb45da1ec5524254