Pith. sign in

Paper Citation Record · LEDGER

Effectively obtaining acoustic, visual and textual data from videos

As of 8 August 2026, this Paper Citation Record lists 100 of 139 outbound references and 1 inbound Pith citation observation for arXiv:2509.05786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.05786 v1

Coverage vector

measured 100 of 139 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:01:36.985927Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:25:26.928879Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T21:25:28.232250Z

Reference resolution

100 of 139 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec8d8242-3b6f-423d-923a-7ebb9df55834 · outbound

This paper cites MusicLM: Generating Music From Text.

Effectively obtaining acoustic, visual and textual data from videos MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.507572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.507572Z digest=sha256:9ce6f2373eeb9beafb38692f3c60a0d1fe886e6dd805c39c2590d6cf65b46b61

Observation 3cd07964-a3bb-4358-b095-78de3fdb81ea · outbound

This paper cites Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering.

Effectively obtaining acoustic, visual and textual data from videos Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.513871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.513871Z digest=sha256:f33bf8f62d0a5ffe5d0f993c6ac52d4a395d899aaf51e5b78615a3f36ac699ef

Observation 99c6f090-968f-400b-a719-1ec22b725617 · outbound

This paper cites Mistral Models, 2024.

Effectively obtaining acoustic, visual and textual data from videos Mistral Models, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.518771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.518771Z digest=sha256:ec831d97b7b60d994e88abac9e66e39a8284237263e19ce5e7740725ba9c812b

Observation 372ebeb3-17cf-44e7-a93e-0f4d91613e8b · outbound

This paper cites Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019.

Effectively obtaining acoustic, visual and textual data from videos Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.523453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.523453Z digest=sha256:1d22d786224c1d95fbe9b8130d3a9e23fcecba5a2f6e1464a9dc7f1b2ab0a0d6

Observation fc2a9b67-4ffb-489d-8878-eb742ee8314b · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024.

Effectively obtaining acoustic, visual and textual data from videos The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.528094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.528094Z digest=sha256:c78ce800d11313f501dcbfeb0c9b0dee97c1fe5236e565bf2f1978a145a8239e

Observation e70fe5ab-2960-40f9-85db-628bf2a1f1c6 · outbound

This paper cites SoundNet: Learning Sound Repre- sentations from Unlabeled Video.

Effectively obtaining acoustic, visual and textual data from videos SoundNet: Learning Sound Repre- sentations from Unlabeled Video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.532854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.532854Z digest=sha256:b785ac8767ef61dba79a0ba660ff56043eebc02a662ec044397dd1eb7c07e8fc

Observation defe6cc8-1aa3-440e-b120-7fc94b5f0988 · outbound

This paper cites Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph.

Effectively obtaining acoustic, visual and textual data from videos Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.538269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.538269Z digest=sha256:3ab25d155c76074ec8762677c21eec00111ee308cd6d8c44ed4beb2aa609ab66

Observation 316c0bbb-e9d0-480e-b692-35af49aa94d9 · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models.

Effectively obtaining acoustic, visual and textual data from videos AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.544163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.544163Z digest=sha256:977d81ca60bd621c12ba12ca336933aea699f96ee2b925a87b1e3bfe12325ef9

Observation f4562230-9781-45e6-8bc2-4ce9ac26b683 · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval.

Effectively obtaining acoustic, visual and textual data from videos Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.548911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.548911Z digest=sha256:50a2af7fba9495d962decbaeddd66cebcb6e2cf31725cf262338cbdb26ff906e

Observation 6d290b8a-a6a8-458e-90a5-6fc08362aa87 · outbound

This paper cites Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024.

Effectively obtaining acoustic, visual and textual data from videos Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.554025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.554025Z digest=sha256:1c47f84094e03717776df14291e8306aa9694aa29bf6387bf286c6d44af8e702

Observation 03119dd8-0b6d-496a-b321-f8fee36243ba · outbound

This paper cites Ballester.

Effectively obtaining acoustic, visual and textual data from videos Ballester

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.558715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.558715Z digest=sha256:bdedd847dc3e0dd9c390e2479776576dce575fd8c21c6843cfb8580090fd5e86

Observation b28a154a-047f-41e8-bfa6-deb51104220b · outbound

This paper cites Improving Image Generation with Better Captions.

Effectively obtaining acoustic, visual and textual data from videos Improving Image Generation with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.563846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.563846Z digest=sha256:ca06d50f286d102c4161c07d226b4deb8b70aa7034ee8589b8a8e7c4a01752af

Observation 2ec45e4e-f960-4ff1-b3ea-6aada080ed54 · outbound

This paper cites RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model.

Effectively obtaining acoustic, visual and textual data from videos RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.568575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.568575Z digest=sha256:a7a470a3363c700ec99cb07c21c9b8d6438dc60ee078ba1bfadd992d9f0aa370

Observation 02198a59-9a78-4616-90b4-f1440f833053 · outbound

This paper cites Birhane and V.

Effectively obtaining acoustic, visual and textual data from videos Birhane and V

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.573366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.573366Z digest=sha256:9b34e9743ca17fa96f326697b7b7a60b070b261974f19ab5777fb7b9c1ef9e64

Observation 73c69020-6c12-4726-8872-7bfcaacb2a72 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.577741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.577741Z digest=sha256:356053fdc16114db64ef8c303ad984c9d06dd7d81804be8cf8a8c34105d69d62

Observation ad5ee963-7aeb-41d9-9f81-57849769cf6a · outbound

This paper cites Using acoustic indices in ecology: Guidance on study design, analyses and interpretation.Methods in Ecology and Evolution, 14(9):2192–2204, 2023.

Effectively obtaining acoustic, visual and textual data from videos Using acoustic indices in ecology: Guidance on study design, analyses and interpretation.Methods in Ecology and Evolution, 14(9):2192–2204, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.582219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.582219Z digest=sha256:d8f7d02ae77d940bf0a592ac0ece1f03fe205bad07f0b95a5fa11913b4e7cfed

Observation bd02f02c-f045-4923-9bf2-a8c0dcc8a0b3 · outbound

This paper cites Bar- nett, Amy Beeston, Jennifer Darby, Benedict Dell, Nick Gardner, Amandine Gasc, Becky Heath, Nia Howells, Magnus Janson, Maria-Viktoria Kyoseva, Thomas Luy- paert, Oliver C.

Effectively obtaining acoustic, visual and textual data from videos Bar- nett, Amy Beeston, Jennifer Darby, Benedict Dell, Nick Gardner, Amandine Gasc, Becky Heath, Nia Howells, Magnus Janson, Maria-Viktoria Kyoseva, Thomas Luy- paert, Oliver C

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.586717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.586717Z digest=sha256:d377614f40f8c0ad9637d0d05a998cd713816e8d104f2c71983494a9d6b3d081

Observation d509692d-bf29-4afd-9c5a-4e0081b145cd · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.591285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.591285Z digest=sha256:2e7d68503466026d82819e0c6e7175395b34b762ebb9e7c5d8cc22de2292b20d

Observation c42c01c6-2259-40b2-8bf6-054b689513a7 · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Con- cepts.

Effectively obtaining acoustic, visual and textual data from videos Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Con- cepts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.595658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.595658Z digest=sha256:d546206d3c3867a6c5c24cbaa21ac7c9bd588b227ad5de0ba7c5d8a0944e5a5e

Observation a1f20d76-310a-4586-a321-85045de90fa4 · outbound

This paper cites Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset.

Effectively obtaining acoustic, visual and textual data from videos Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.599895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.599895Z digest=sha256:2fefd08725e09efdd19f31d4d3ff3475585a04b4d86b2ec619cc34affbc33d17

Observation f5cc86ea-68c2-4fc8-be17-1c99e942cdf4 · outbound

This paper cites Veo, 2024.

Effectively obtaining acoustic, visual and textual data from videos Veo, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.604495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.604495Z digest=sha256:e5fd08765dd2795e9812439a34e435ba435103ef007327938747d628a903a71c

Observation 6d80e0b5-c3f6-4668-94e8-56131c89436a · outbound

This paper cites Central Limit Theorem in the Functional Approach.IEEE Transactions on Signal Processing, 61(16):4025–4037, 2013.

Effectively obtaining acoustic, visual and textual data from videos Central Limit Theorem in the Functional Approach.IEEE Transactions on Signal Processing, 61(16):4025–4037, 2013

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.608792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.608792Z digest=sha256:0dad775d0cfab788fbeb9fa155b272e2b323dd587358f73577f1588a13700616

Observation b8407a28-33a6-4990-8995-1dc7dbda6f68 · outbound

This paper cites A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021.

Effectively obtaining acoustic, visual and textual data from videos A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.613359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.613359Z digest=sha256:06c5ff0cd92e9fd9a9f364e364e1cc15ec08533e9ca09c1b083ebe7a4b2a6589

Observation 8963eb35-2f12-492e-8765-595ec69aed77 · outbound

This paper cites Jukebox: A Generative Model for Music.

Effectively obtaining acoustic, visual and textual data from videos Jukebox: A Generative Model for Music

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.617814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.617814Z digest=sha256:1229508c1a7d5ee9f529bd04401b2a185ad4a8bedfb6d2d10c12056bac26d1c9

Observation 1767255c-ec5e-4753-80da-4b32e29684a4 · outbound

This paper cites The Llama 3 Herd of Models.

Effectively obtaining acoustic, visual and textual data from videos The Llama 3 Herd of Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.622952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.622952Z digest=sha256:34772e26a73cd1a259d9bfa57a0d27071e890b603465653aa4ad26195f717cc9

Observation 31693e76-b8fc-4ca9-a1f6-a460b5f4697e · outbound

This paper cites Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022.

Effectively obtaining acoustic, visual and textual data from videos Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.628250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.628250Z digest=sha256:10cb5b5ceea4eed19c144bab018e866620636075bf88b984731c173273943d69

Observation ca63f4b9-2338-4a79-9b2f-30cb1031323f · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Effectively obtaining acoustic, visual and textual data from videos Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.632876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.632876Z digest=sha256:ead64e7fbe65718bd1de92a28b18a60a4f7d86a801cbe56fcd3f076f5763e582

Observation 38f1b52f-fbff-4cd1-958c-05edc33f38c5 · outbound

This paper cites Creativity and Machine Learning: A Survey.

Effectively obtaining acoustic, visual and textual data from videos Creativity and Machine Learning: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.637916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.637916Z digest=sha256:565a06279c625b50e55ab256557fa1adaa1c8a00b478e2b4b5a99a7d10829f1b

Observation b89344f3-fc01-4770-b056-cf9382f410a9 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Effectively obtaining acoustic, visual and textual data from videos The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.642585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.642585Z digest=sha256:5c4302db3e6ecf5e763207dc4b05bd9f1c2ec2fc2bc6291e47adeec88c0c2af5

Observation 2c00d998-a181-45f2-900c-f5f5773eb485 · outbound

This paper cites Listen to Look: Action Recognition by Previewing Audio.

Effectively obtaining acoustic, visual and textual data from videos Listen to Look: Action Recognition by Previewing Audio

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:01:38.060083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.647072Z digest=sha256:5eb88d4c8db8cd76c4afcc7de78fde5930c30285101e35e479fc02e963f788a6

Observation 6b0bc007-5648-4acb-a8c1-f6965de53a81 · outbound

This paper cites Gemmeke, Daniel P.

Effectively obtaining acoustic, visual and textual data from videos Gemmeke, Daniel P

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.651713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.651713Z digest=sha256:bfa5e5a6f003c6e07a60148121be968478c1830b4ad4f23da3c68ca5febe329d

Observation cc9105c9-c3f9-462e-85db-995cc7d6b868 · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

Effectively obtaining acoustic, visual and textual data from videos ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.656278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.656278Z digest=sha256:b281154b3a81aea2fcb1bc432cef11f126b8a424eb25ff3fef4539ec07782128

Observation 1d6d9822-6e67-4852-adcb-f8fbcddbb45e · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.661256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.661256Z digest=sha256:d7487195eaf87c95595563644dbcd4fe85e79c866b2e2ce23cf5223823a9b01e

Observation 324a05e5-c1e6-4f91-a830-2717aba335bf · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Effectively obtaining acoustic, visual and textual data from videos Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.666271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.666271Z digest=sha256:35f3cfb05f2b0be2fe523ff5c0a8be6f9430f55f1b6b6ef6f4e4afef0659d974

Observation 3ebed149-c399-4210-83ca-584ff496e8f1 · outbound

This paper cites Temporal Alignment Networks for Long-term Video.

Effectively obtaining acoustic, visual and textual data from videos Temporal Alignment Networks for Long-term Video

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:01:38.001715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.671184Z digest=sha256:6218210561cb753eda1e58496868e65a9efb53cf5394cdc6c3cf2d26bab42d3b

Observation f68d5808-7bb5-4605-a3ca-809bfd1eb156 · outbound

This paper cites Rae, and Laurent Sifre.

Effectively obtaining acoustic, visual and textual data from videos Rae, and Laurent Sifre

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.677051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.677051Z digest=sha256:b74cf82a422e0b90f82151b472378ef7e7f6b600d38870ce556f006bc0d2b1f2

Observation c08dc387-7c57-484b-95b9-115f89fbfe98 · outbound

This paper cites Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model.

Effectively obtaining acoustic, visual and textual data from videos Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.681822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.681822Z digest=sha256:cf04f19b41db5c95e18da1864056ebb588b1575544958afa92d7be0ff14b559d

Observation 1a400fbb-7282-4dcf-9cc2-f6a894c37edf · outbound

This paper cites Key Frame Selection for Temporal Graph Opti- mization of Skeleton-Based Action Recognition.Applied Sciences, 14(21), 2024.

Effectively obtaining acoustic, visual and textual data from videos Key Frame Selection for Temporal Graph Opti- mization of Skeleton-Based Action Recognition.Applied Sciences, 14(21), 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.686461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.686461Z digest=sha256:3b7133cb78f65ac387f87196b72d39e54e3ffa1bcd742ae9dbcf350c9a1394cc

Observation e2619ad0-e7db-43b4-b75d-10b811b3eb7e · outbound

This paper cites NLIP: Noise-Robust Language-Image Pre-training.

Effectively obtaining acoustic, visual and textual data from videos NLIP: Noise-Robust Language-Image Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.691096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.691096Z digest=sha256:bb77b39e3325da40da666e223076558aa2cdc0bb740eaa592c597a74fc677206

Observation 8030a54d-4642-4e07-ac9d-8312ffd4ec29 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.695555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.695555Z digest=sha256:b6746f0c8a7eeb78aceb58da6e388301ed8a78187498735996443894cb8050f8

Observation d5b226b3-6510-48a0-9811-c9febbe51052 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.700455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.700455Z digest=sha256:04a9222ad266e0538474f912fe72678b3b2d84fae4e49834d89bf67a8c3bf385

Observation 947368cd-25dd-45e2-8879-1fd29ceb871a · outbound

This paper cites LL VIP: A Visible- infrared Paired Dataset for Low-light Vision.

Effectively obtaining acoustic, visual and textual data from videos LL VIP: A Visible- infrared Paired Dataset for Low-light Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.705388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.705388Z digest=sha256:e202be697bf4b99d40e54225058c6b0463ad075f9618c089ed815d2c51e75804

Observation 93caa914-ba11-49d9-8659-e4964ef9e460 · outbound

This paper cites TimbreCLIP: Connecting Timbre to Text and Images.

Effectively obtaining acoustic, visual and textual data from videos TimbreCLIP: Connecting Timbre to Text and Images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.709931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.709931Z digest=sha256:108a0d8ed79d829b95ac8366f5f0da0128f9dc28186836216e0c25082b1ec77a

Observation 1f7baff1-df83-4c62-853a-066e913ca8a9 · outbound

This paper cites Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning.

Effectively obtaining acoustic, visual and textual data from videos Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.714574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.714574Z digest=sha256:0e308363c3ea7d2dbfca438f0791dcbcc6c23aa3f30d4b98e82bfca896af0efe

Observation 78a115d0-1f88-48b5-ae44-df72cada4263 · outbound

This paper cites MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition.

Effectively obtaining acoustic, visual and textual data from videos MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:01:37.742120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.718947Z digest=sha256:49fd75d1075c7393764a339953ad5ba976590d208bebc0dbed2a063db29cebad

Observation 5a4026dd-2113-40a0-9e48-6acf31061119 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.723724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.723724Z digest=sha256:626186fe717f650faa0e0264cd9906046bef4bcfe65cb30f0a85e46ea49f1b67

Observation bc087ea3-9c7c-48f1-bec4-ac70b048557b · outbound

This paper cites AudioCaps: Generating Captions for Audios in The Wild.

Effectively obtaining acoustic, visual and textual data from videos AudioCaps: Generating Captions for Audios in The Wild

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.728086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.728086Z digest=sha256:60e08c532ef7b03e75c634f3377ba749faed8c3bd50b5ac7119905eebde1e122

Observation 4e1e2dec-7704-4b8c-935e-200e8f2c817e · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Effectively obtaining acoustic, visual and textual data from videos Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.732735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.732735Z digest=sha256:ccc06094c675658535a2ebacdc5c9945a48d3f76936a87f00eb2cd245e2b33d9

Observation a1260ad0-06c3-4bd2-87ee-c4e0c0a3b362 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Effectively obtaining acoustic, visual and textual data from videos AudioGen: Textually Guided Audio Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.737476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.737476Z digest=sha256:e39bd7b4dd4711fca84aaa8840a75747eb223bc2f3da10b8e1b5420434e847a3

Observation 40ea535a-3d4a-45f8-96d2-14b38f8a8f55 · outbound

This paper cites BindDiffusion: One Diffusion Model to Bind Them All, 2024.

Effectively obtaining acoustic, visual and textual data from videos BindDiffusion: One Diffusion Model to Bind Them All, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.742843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.742843Z digest=sha256:af1ab1b5846498e2d19b90fc972ba79e2af842068cc39a7c03ca422157583a36

Observation e99ea385-2b6e-4480-b21e-ab166308de09 · outbound

This paper cites FLUX, 2024.

Effectively obtaining acoustic, visual and textual data from videos FLUX, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.748006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.748006Z digest=sha256:0a3585fcbec913a728288f6bf379ee27cad27fd94a9ad8cef348a19651b00d2b

Observation 0faaade5-8f3f-4af3-a586-0d186385b1ab · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.752587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.752587Z digest=sha256:20e1cc9fdd4a6f70a59dd83539c6616816501c1d883477e34cd19847b364fe93

Observation a61c7e02-bda4-4af6-9a4d-542b7246ea2c · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.756811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.756811Z digest=sha256:1753d4f100bdef67bc7ab723f9b0e775ddb2276307f2f577b2eec09197ff9ed1

Observation feb64b77-e522-4dca-b0aa-798fe3fb0323 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.762891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.762891Z digest=sha256:fc809c62139d60c4ecab1eb838d0aadc3f00c6f9a36fc1979461ff3662d17449

Observation 23acdd7b-6340-44a0-a435-65b815bff5fc · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.767567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.767567Z digest=sha256:9771801d9702222c49294a3fd86c876e8c970aa02c89722061993b654d2d1fba

Observation 566fad45-6b69-461c-9233-29cfb1f62629 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.772256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.772256Z digest=sha256:6a984af70b0d6260a1ff9a9328218de2d518ed4e4f7a936922bfa3298dc53ba0

Observation 61148e1b-803c-4613-b3ff-dd1dd8afa080 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.777288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.777288Z digest=sha256:2466c306e7eb7a3d37c02d2c6e501d7ca8bc630649543d6fb0f962539898daa6

Observation 5074745a-3fc1-4556-95af-4a061d8d6e5e · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.781925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.781925Z digest=sha256:895dd005ca90020a9ae1676e1545dff77971e0d36590f01ee85e13a783d177ad

Observation dd476d19-c547-42a2-b863-03bd1bda0ba6 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.786535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.786535Z digest=sha256:7790de4b4fb3d05fe580194dedb9f042bad3936b796f25b5b7e83049bc886e49

Observation 77b59946-7569-4cfe-8d9f-76164e9c046d · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

Effectively obtaining acoustic, visual and textual data from videos OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.791275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.791275Z digest=sha256:b55633c5e9e83d9d574041f22479b3d6747272170420904aa37ca59d823aa00c

Observation fc848145-e1df-42a1-9291-03dd77cb3ef1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Effectively obtaining acoustic, visual and textual data from videos BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.796174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.796174Z digest=sha256:23eaa58a89c03aa91f5141aa2b0cea94fc8d1923b44185e48c7064cd4b56096c

Observation 655e029f-4504-4828-b44b-460328908220 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Effectively obtaining acoustic, visual and textual data from videos BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.801004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.801004Z digest=sha256:aca49792692648f456bba478e58330084597c2a99cfcb0930cacf280141c4d48

Observation 02170d03-b9a9-48b0-9946-3de063dc4f0b · outbound

This paper cites Word-Level Explanations for Analyzing Bias in Text-to-Image Models.

Effectively obtaining acoustic, visual and textual data from videos Word-Level Explanations for Analyzing Bias in Text-to-Image Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.805782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.805782Z digest=sha256:4dd0039634eaaf56f5ba65349416362171e32a42cfc78cfadfcdf5e8d24106ee

Observation bac5b7e1-71c3-425a-ad31-eb94a85b5459 · outbound

This paper cites Lawrence Zitnick.

Effectively obtaining acoustic, visual and textual data from videos Lawrence Zitnick

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.810527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.810527Z digest=sha256:aa4505d96353292355ed6e0cebf63a24c3f029f32207c17e213521f64a365391

Observation 724e1ee9-d848-4c4b-99ea-2e5f3deec07c · outbound

This paper cites A Comparison Between KeyFrame Extraction Methods for Clothing Recognition, 2023.

Effectively obtaining acoustic, visual and textual data from videos A Comparison Between KeyFrame Extraction Methods for Clothing Recognition, 2023

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.815034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.815034Z digest=sha256:025669e7ac4c3dbd32e08169fa5716ce7f086350557befb1ecc0d5b8d0bc22be

Observation 3644ee48-de0c-47c8-bb57-9eed6a07f0d3 · outbound

This paper cites Plumbley.

Effectively obtaining acoustic, visual and textual data from videos Plumbley

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.819846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.819846Z digest=sha256:c8a57c4f2698d7187e70c7cc40daf498e146b008dd777ec30b8caa57695841b3

Observation 734295a9-b47e-49ef-a458-dffdfe7d8398 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Effectively obtaining acoustic, visual and textual data from videos Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.824327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.824327Z digest=sha256:c4e2b50cd03eb20e25d67beb13836d68b92d7f95d2499f68ac9c729a5fa82841

Observation ba28e53f-7947-48ae-8690-28f0e7e5beb0 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.829712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.829712Z digest=sha256:bcfe58be6b82894b4340d12038a78ac723763dcdce2be27c06485d199d8ca497

Observation 94c9feb5-863e-4d14-ba60-ec85cecbc01e · outbound

This paper cites BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning.

Effectively obtaining acoustic, visual and textual data from videos BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.835254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.835254Z digest=sha256:38705d43c9f18fb18d70ef67725303215c4556c2bd48e952ccb97a1bb7bac76f

Observation e4fd3a40-1031-4c09-8502-3071e09ade49 · outbound

This paper cites Stable Diffusion Akashic Records, 2023.

Effectively obtaining acoustic, visual and textual data from videos Stable Diffusion Akashic Records, 2023

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.840392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.840392Z digest=sha256:e24dcdb04ddbbe716fadab76b6013da9f927726cf7c02da41c549eb340406085

Observation 9af69ee2-41fc-4bb7-9a40-bea10e6dbff3 · outbound

This paper cites GenRL: Multimodal-foundation world models for generalization in embodied agents.

Effectively obtaining acoustic, visual and textual data from videos GenRL: Multimodal-foundation world models for generalization in embodied agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.845122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.845122Z digest=sha256:9d1a4bfa8ba34598af3cc1dc09b71819a279b8ba3bb456393b0cb2f7d887052f

Observation dce89b1f-caf0-430f-a02a-6568fa5b4c7e · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Effectively obtaining acoustic, visual and textual data from videos Mustango: Toward Controllable Text-to-Music Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.849835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.849835Z digest=sha256:6876747ad4325aedd74a7aa682f413675ba257e4edb6e5fd4a0e58c27597b2b9

Observation e2bddc2e-62e9-4019-8434-618c75b22c38 · outbound

This paper cites Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis.

Effectively obtaining acoustic, visual and textual data from videos Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.854428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.854428Z digest=sha256:c307b7f3194181f51bca5409421135ccdeb264627c0562daecc25c9dbb8d6c46

Observation 95916d8e-de34-48b0-ba78-b829705542a5 · outbound

This paper cites DALL·E 3 System Card, 2023.

Effectively obtaining acoustic, visual and textual data from videos DALL·E 3 System Card, 2023

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.860309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.860309Z digest=sha256:11493c400cfa6dbbba11f6e83ce8880d7ac74a2185415f5ef4d3f30ba13447cd

Observation f1955f34-dc82-4f77-897c-572158524ffe · outbound

This paper cites Video generation models as world simulators, 2024.

Effectively obtaining acoustic, visual and textual data from videos Video generation models as world simulators, 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.864946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.864946Z digest=sha256:539594fcaba2de4066a0c2112780f569836065f6d21c66e4a2300cbd5e363d69

Observation 232e3c2f-2a47-448c-ae2f-a304e18c9c82 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.869533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.869533Z digest=sha256:78ff2ded1abc4cff3bc7b39dfe89738d5210ebb163ba9e1f0838d9d78852ee0b

Observation 42c840a9-50bc-43b6-9010-7f8bb454a298 · outbound

This paper cites Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022.

Effectively obtaining acoustic, visual and textual data from videos Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.874061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.874061Z digest=sha256:d2c791ee95368b8a4e1fe2b95b70198d267c785ad3046ba7cd164e85fc6289ed

Observation 071a64e6-495c-4522-bfe9-fbbd871374f7 · outbound

This paper cites AudioSetZSL, 2019.

Effectively obtaining acoustic, visual and textual data from videos AudioSetZSL, 2019

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.879208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.879208Z digest=sha256:ebcbbdbc0e7b2d3c8c10bdde9d89e08b2e820c955f94b43b1dcdf14ecab135e2

Observation 2223b9a0-d786-4723-b199-03d5d455ed09 · outbound

This paper cites Coordi- nated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classifi- cation and Retrieval of Videos.

Effectively obtaining acoustic, visual and textual data from videos Coordi- nated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classifi- cation and Retrieval of Videos

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.884006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.884006Z digest=sha256:56ae0ef17e5fb4b3d9d66fb429cbd4dccc5f5f5fd73b62f23afe197147e21593

Observation 4ccd66b9-3ce4-4bb6-8ba1-378cccafe673 · outbound

This paper cites Pijanowski, Luis J.

Effectively obtaining acoustic, visual and textual data from videos Pijanowski, Luis J

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.888564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.888564Z digest=sha256:e376bb4ea00c51596411dc38765149eaf69ee78158754912db2767a3c2ba869a

Observation cbee204a-803b-4c13-bc2c-e870549cd3b2 · outbound

This paper cites Plummer, Liwei Wang, Chris M.

Effectively obtaining acoustic, visual and textual data from videos Plummer, Liwei Wang, Chris M

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.893458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.893458Z digest=sha256:1e141a59c21cd6d55e86f90d47f0890549ff5437c6091452768ef49b167af8ad

Observation df676aed-bb72-4ac0-b295-29901731b91d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Effectively obtaining acoustic, visual and textual data from videos SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.898253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.898253Z digest=sha256:77922bc70ce2f48feaccacf6b53d5f27915204a059781c999f021cf1f55bdf90

Observation d13fe9b0-f7e7-44bd-b383-7d52c3a1a9aa · outbound

This paper cites Does mixing of speech signals comply with central limit theorem? International Journal of Electronics and Communications, 62(10):782–785, 2008.

Effectively obtaining acoustic, visual and textual data from videos Does mixing of speech signals comply with central limit theorem? International Journal of Electronics and Communications, 62(10):782–785, 2008

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.722935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.903031Z digest=sha256:6d267e86439580a882cd95ee47cab467c03c9bbe4d8cc6b9c4ff866f0d71b753

Observation 27f08d89-5358-4c74-84b1-f2cb26e6d60a · outbound

This paper cites MirrorGAN: Learning Text-To-Image Generation by Redescription.

Effectively obtaining acoustic, visual and textual data from videos MirrorGAN: Learning Text-To-Image Generation by Redescription

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.907592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.907592Z digest=sha256:47a77c205a680a35c8d7beb4e0782811abefbab52ee4fe5bcd05fe4d09de8b2c

Observation e3c07864-0d86-489c-95ab-81798b370b51 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Effectively obtaining acoustic, visual and textual data from videos Learning Transferable Visual Models From Natural Language Supervision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.912294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.912294Z digest=sha256:b1321a47dae383e9044551129fe70f9c4de93e38b8f7eba66908432971beae2a

Observation aedb676f-01de-41a3-996e-a5061d3e48ee · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Effectively obtaining acoustic, visual and textual data from videos Robust Speech Recognition via Large-Scale Weak Supervision

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.917064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.917064Z digest=sha256:6e00c675e45198a774dd390cb7fe237ffe75db83ecea41efcbf22c757df75381

Observation 87d9060f-42ac-4c74-b9f8-de95ef1ef747 · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Effectively obtaining acoustic, visual and textual data from videos Zero-Shot Text-to-Image Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.921533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.921533Z digest=sha256:cf2aef9b81d2640db69e9f8623c32db57ebffe1c15fa2975739c56e7449ec8ed

Observation f0de62db-f4ac-4b36-87d6-36ac83a0c53d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Effectively obtaining acoustic, visual and textual data from videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.926838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.926838Z digest=sha256:11aea819063d5ba21bc07901178b07f76f84e66409bfd0ece9ada90f2ab299cc

Observation becfe866-8825-413b-8892-eb51594631fe · outbound

This paper cites Stable Diffusion, 2021.

Effectively obtaining acoustic, visual and textual data from videos Stable Diffusion, 2021

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.683239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.932052Z digest=sha256:053c1ef61a7518ce199620f2d5cb8905f2f07df398d8212e6b1c25948ea6c039

Observation 0465db97-e607-4ada-a9ba-331eec2aa695 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Effectively obtaining acoustic, visual and textual data from videos High-Resolution Image Synthesis with Latent Diffusion Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.937443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.937443Z digest=sha256:cc7706f8adc82d54451dd391f205768087d701e50c67f9544565c55139c80819

Observation 8e505e95-660f-4781-bda9-a9e2b108ebb9 · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024.

Effectively obtaining acoustic, visual and textual data from videos Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.942534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.942534Z digest=sha256:82a7a7bfb67c419575d9aa741ababfe346ed666cb9d69c5971550f61268d315f

Observation 74544f37-caf6-4c0a-ba80-96e6065ba4e1 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Effectively obtaining acoustic, visual and textual data from videos Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.947834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.947834Z digest=sha256:cf034c1f5a9d26d9769bce1b4c583435a26c335cde7bc9d87a0e6bce767ceef6

Observation 3fd18f94-87ce-447f-86a3-1f94f9743bd8 · outbound

This paper cites ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition.

Effectively obtaining acoustic, visual and textual data from videos ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.644398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.953007Z digest=sha256:7a5693de35d8eea9d4d4f03220a745f2da734243e105e220796818feb46d6453

Observation 6c2abd28-5c1e-43b3-abfb-6a13cf33016d · outbound

This paper cites Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey.

Effectively obtaining acoustic, visual and textual data from videos Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.957524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.957524Z digest=sha256:b7f540af717c2a2741287e795f8dedc05c1e1115d47e1d8ef2b675920f288680

Observation 1e99fdc3-62d6-4ba8-beb5-859e76015007 · outbound

This paper cites What is noise?Geophysics, 63(4):1122–1124, 1998.

Effectively obtaining acoustic, visual and textual data from videos What is noise?Geophysics, 63(4):1122–1124, 1998

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.628464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.962343Z digest=sha256:48f647effebf2565ec9c1edf3c8e969a32ea119e39f5472c70d9d478d30c3cee

Observation f86c7b08-e085-4929-a849-b350da166928 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Effectively obtaining acoustic, visual and textual data from videos LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.612113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:01:36.967095Z digest=sha256:9ce07455c79f1b3d3802fc8d7b6af7287bbc548334a6302d259fb5843354138c

Observation 1141ad1b-87bf-4cf4-8205-39e917d797f9 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Effectively obtaining acoustic, visual and textual data from videos LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.971628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.971628Z digest=sha256:38dc04e8b6d95be1f97e9feda8f9f02072cd80c7f229e01a2a02dfa8ad028d32

Observation 181fc8f0-d346-40ef-867e-3e508f577b82 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.976501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.976501Z digest=sha256:c961d682cf1d15216a9800c73e50c2dc20d4d156cd9c7dd307c047bf35c8f00b

Observation 7bae9aa2-c149-499e-a4ec-8a463198ecb6 · outbound

This paper cites I Hear Your True Colors: Image Guided Audio Generation.

Effectively obtaining acoustic, visual and textual data from videos I Hear Your True Colors: Image Guided Audio Generation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.981088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.981088Z digest=sha256:07aeefd53d26e3ad393810f9b3b3107f7ba20c7294d3e05ab0a291caa2f6c7de

Observation b956306a-aaff-4ab1-bd9c-34721c77bd74 · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Effectively obtaining acoustic, visual and textual data from videos A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.985927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.985927Z digest=sha256:1387ce53b0f862c8ba1d9ebca40d7e11deb51fa2fbc8b6c48558cc8384af85c9

Pith citing papers

Observation 7f772987-5f54-46cf-9866-ec6efe4b5788 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Effectively obtaining acoustic, visual and textual data from videos

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:25:28.242608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T21:25:26.928879Z digest=sha256:ad80beda56b6a4a7c9b6d0f6c5f286765048b8e4a3ad7f9229ebf88050bed732