Pith. sign in

Paper Citation Record · LEDGER

Effectively obtaining acoustic, visual and textual data from videos

As of 13 August 2026, this Paper Citation Record lists 100 of 139 outbound references and 1 inbound Pith citation observation for arXiv:2509.05786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.05786 v1

Coverage vector

measured 100 of 139 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:01:36.985927Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:25:26.928879Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T21:25:28.232250Z

Reference resolution

100 of 139 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec8d8242-3b6f-423d-923a-7ebb9df55834 · outbound

This paper cites MusicLM: Generating Music From Text.

Effectively obtaining acoustic, visual and textual data from videos MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.507572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.507572Z digest=sha256:6760bd04332e1307a5c3892aa508d10a9ad77c05020a0abf04930b5dae8c7d1d

Observation 3cd07964-a3bb-4358-b095-78de3fdb81ea · outbound

This paper cites Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering.

Effectively obtaining acoustic, visual and textual data from videos Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.513871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.513871Z digest=sha256:cc71afabb6b9abdff098c7265192d1d3ebb416ff66cb453dbd1c339104f40d88

Observation 99c6f090-968f-400b-a719-1ec22b725617 · outbound

This paper cites Mistral Models, 2024.

Effectively obtaining acoustic, visual and textual data from videos Mistral Models, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.518771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.518771Z digest=sha256:ac36c291f43411cade7df160797cfaa4ab153295eca0e0f48594cd3945c48e56

Observation 372ebeb3-17cf-44e7-a93e-0f4d91613e8b · outbound

This paper cites Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019.

Effectively obtaining acoustic, visual and textual data from videos Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.523453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.523453Z digest=sha256:0592aa2d12bd4d74de3fd57c043133fe74e1bd379adc89820bc518a00dfe9c06

Observation fc2a9b67-4ffb-489d-8878-eb742ee8314b · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024.

Effectively obtaining acoustic, visual and textual data from videos The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.528094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.528094Z digest=sha256:fdb5f17992dd53366187b84509b0e75f0ca70341220e4ba4c62033962fc150f4

Observation e70fe5ab-2960-40f9-85db-628bf2a1f1c6 · outbound

This paper cites SoundNet: Learning Sound Repre- sentations from Unlabeled Video.

Effectively obtaining acoustic, visual and textual data from videos SoundNet: Learning Sound Repre- sentations from Unlabeled Video

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.532854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.532854Z digest=sha256:0bbc5fda4c0ead9328ea6e7e160f59a8a3c357b2cf5c9d9fd29639df7a6817fc

Observation defe6cc8-1aa3-440e-b120-7fc94b5f0988 · outbound

This paper cites Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph.

Effectively obtaining acoustic, visual and textual data from videos Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.538269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.538269Z digest=sha256:d33907aef2e2d64ce18a66661829089878a21c1418c1f94e062349985917effd

Observation 316c0bbb-e9d0-480e-b692-35af49aa94d9 · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models.

Effectively obtaining acoustic, visual and textual data from videos AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.544163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.544163Z digest=sha256:f593cdc0ec02ee09ecca1ba76dcbbf55f02e457ddb42bf392d87820a45f825a5

Observation f4562230-9781-45e6-8bc2-4ce9ac26b683 · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval.

Effectively obtaining acoustic, visual and textual data from videos Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.548911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.548911Z digest=sha256:f65d5d4aac611ba0e481a8be5eecd320b04c424098e1e6be6f7b9ac981144796

Observation 6d290b8a-a6a8-458e-90a5-6fc08362aa87 · outbound

This paper cites Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024.

Effectively obtaining acoustic, visual and textual data from videos Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.554025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.554025Z digest=sha256:3a3f1fe28cc8200ecc0ca09a04b56bffa5a312d6f0bf524685776043cf4f0a9b

Observation 03119dd8-0b6d-496a-b321-f8fee36243ba · outbound

This paper cites Ballester.

Effectively obtaining acoustic, visual and textual data from videos Ballester

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.558715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.558715Z digest=sha256:8841e2c3cc1e9be10830779de44788b27a977e7c104d97552352fc51abef2cdf

Observation b28a154a-047f-41e8-bfa6-deb51104220b · outbound

This paper cites Improving Image Generation with Better Captions.

Effectively obtaining acoustic, visual and textual data from videos Improving Image Generation with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.563846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.563846Z digest=sha256:3705e3d16b310ccb9f5d01dac15ff1ffc7d607a349168b31202dfff3359863dd

Observation 2ec45e4e-f960-4ff1-b3ea-6aada080ed54 · outbound

This paper cites RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model.

Effectively obtaining acoustic, visual and textual data from videos RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.568575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.568575Z digest=sha256:6eb8f15660018fa16efaad6b20e7b77c60b49aaa1a52b876e331b989b5771db3

Observation 02198a59-9a78-4616-90b4-f1440f833053 · outbound

This paper cites Birhane and V.

Effectively obtaining acoustic, visual and textual data from videos Birhane and V

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.573366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.573366Z digest=sha256:bbb65c59c9c0c6a35407a6988164e9714f4884917590051a906b7dbeb24f0bad

Observation 73c69020-6c12-4726-8872-7bfcaacb2a72 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.577741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.577741Z digest=sha256:581372de65493fd010e990c0d990beb65d28905156214593164405525d0b708a

Observation ad5ee963-7aeb-41d9-9f81-57849769cf6a · outbound

This paper cites Using acoustic indices in ecology: Guidance on study design, analyses and interpretation.Methods in Ecology and Evolution, 14(9):2192–2204, 2023.

Effectively obtaining acoustic, visual and textual data from videos Using acoustic indices in ecology: Guidance on study design, analyses and interpretation.Methods in Ecology and Evolution, 14(9):2192–2204, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.582219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.582219Z digest=sha256:2181a3e6c82f766fba685b26a54ef704ba81c31c14894a7a35b0961586e01a99

Observation bd02f02c-f045-4923-9bf2-a8c0dcc8a0b3 · outbound

This paper cites Bar- nett, Amy Beeston, Jennifer Darby, Benedict Dell, Nick Gardner, Amandine Gasc, Becky Heath, Nia Howells, Magnus Janson, Maria-Viktoria Kyoseva, Thomas Luy- paert, Oliver C.

Effectively obtaining acoustic, visual and textual data from videos Bar- nett, Amy Beeston, Jennifer Darby, Benedict Dell, Nick Gardner, Amandine Gasc, Becky Heath, Nia Howells, Magnus Janson, Maria-Viktoria Kyoseva, Thomas Luy- paert, Oliver C

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.586717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.586717Z digest=sha256:45498d01e348ab49e38c6a82e5ccc2540c25c70bccd4f9d41d8408ee7ed8726f

Observation d509692d-bf29-4afd-9c5a-4e0081b145cd · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.591285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.591285Z digest=sha256:eb8fce224edfa093fc4114740b3443923809870c61f456ed55ad3ac8eec26b09

Observation c42c01c6-2259-40b2-8bf6-054b689513a7 · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Con- cepts.

Effectively obtaining acoustic, visual and textual data from videos Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Con- cepts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.595658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.595658Z digest=sha256:cedc5a6497af1234d2b0c18830d61c15c106574260e73f25a51aa3a0c085a470

Observation a1f20d76-310a-4586-a321-85045de90fa4 · outbound

This paper cites Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset.

Effectively obtaining acoustic, visual and textual data from videos Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.599895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.599895Z digest=sha256:378eacc31aecc3c47715de35b01d217ececc3a33f22e861f523853c736623cd7

Observation f5cc86ea-68c2-4fc8-be17-1c99e942cdf4 · outbound

This paper cites Veo, 2024.

Effectively obtaining acoustic, visual and textual data from videos Veo, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.604495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.604495Z digest=sha256:c1a69aa9c0b12930c4a13e3d989046ce26af840128ef2b71c7509e0bff562c1f

Observation 6d80e0b5-c3f6-4668-94e8-56131c89436a · outbound

This paper cites Central Limit Theorem in the Functional Approach.IEEE Transactions on Signal Processing, 61(16):4025–4037, 2013.

Effectively obtaining acoustic, visual and textual data from videos Central Limit Theorem in the Functional Approach.IEEE Transactions on Signal Processing, 61(16):4025–4037, 2013

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.608792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.608792Z digest=sha256:3b8e2f918b1ac94e0397cad7e109d2346d8bde08993803ae604b0036d878d723

Observation b8407a28-33a6-4990-8995-1dc7dbda6f68 · outbound

This paper cites A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021.

Effectively obtaining acoustic, visual and textual data from videos A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.613359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.613359Z digest=sha256:9769f6bded7492218fef5054b93bca0b5d93ff5bcf5f37f90ee51c4c23c9e8ea

Observation 8963eb35-2f12-492e-8765-595ec69aed77 · outbound

This paper cites Jukebox: A Generative Model for Music.

Effectively obtaining acoustic, visual and textual data from videos Jukebox: A Generative Model for Music

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.617814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.617814Z digest=sha256:9e28e7769bb3030dbb5277486a36ba247714bfa28572fe5d9a16bfcc055cfb13

Observation 1767255c-ec5e-4753-80da-4b32e29684a4 · outbound

This paper cites The Llama 3 Herd of Models.

Effectively obtaining acoustic, visual and textual data from videos The Llama 3 Herd of Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.622952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.622952Z digest=sha256:daf6f1d0c6a433fd90b3e45d3e28eeececa5edff96309b5650e782edbe37e0f4

Observation 31693e76-b8fc-4ca9-a1f6-a460b5f4697e · outbound

This paper cites Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022.

Effectively obtaining acoustic, visual and textual data from videos Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.628250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.628250Z digest=sha256:78eccc4bd3c382ed25ef12b62b3692b895a1e1b1605b93c617a93e74a2744e43

Observation ca63f4b9-2338-4a79-9b2f-30cb1031323f · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Effectively obtaining acoustic, visual and textual data from videos Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.632876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.632876Z digest=sha256:252da8f278a231787ba695e9a2b21d0f0bdbf4fce6d0c0cfbc3fd3974289be18

Observation 38f1b52f-fbff-4cd1-958c-05edc33f38c5 · outbound

This paper cites Creativity and Machine Learning: A Survey.

Effectively obtaining acoustic, visual and textual data from videos Creativity and Machine Learning: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.637916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.637916Z digest=sha256:453ee8ec4ebf29832032eb6ca56c04175215c8c3fcda135fc5e933195a4fa3fe

Observation b89344f3-fc01-4770-b056-cf9382f410a9 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Effectively obtaining acoustic, visual and textual data from videos The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.642585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.642585Z digest=sha256:359c64b51788a3e60f74666dddd9b95fdab1f2d9291f52980e5948273197e852

Observation 2c00d998-a181-45f2-900c-f5f5773eb485 · outbound

This paper cites Listen to Look: Action Recognition by Previewing Audio.

Effectively obtaining acoustic, visual and textual data from videos Listen to Look: Action Recognition by Previewing Audio

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:01:38.060083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.647072Z digest=sha256:3d25345718bb282b167a6bf0b498c1f41830a998ace508cd312bcf7d7a9a2a67

Observation 6b0bc007-5648-4acb-a8c1-f6965de53a81 · outbound

This paper cites Gemmeke, Daniel P.

Effectively obtaining acoustic, visual and textual data from videos Gemmeke, Daniel P

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.651713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.651713Z digest=sha256:c7ebcafc73e49a7dcb3e4f93e549909fed8972e744da38007b23ae8721c56d2f

Observation cc9105c9-c3f9-462e-85db-995cc7d6b868 · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

Effectively obtaining acoustic, visual and textual data from videos ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.656278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.656278Z digest=sha256:df8faed6b5ddcc064300313cacfb75f1df7374f36e26e0a887ba42b022f995bb

Observation 1d6d9822-6e67-4852-adcb-f8fbcddbb45e · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.661256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.661256Z digest=sha256:95bb0f1baddd2cd459eb883b0135a524a57ee75c1d3b825abb07b33983500852

Observation 324a05e5-c1e6-4f91-a830-2717aba335bf · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Effectively obtaining acoustic, visual and textual data from videos Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.666271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.666271Z digest=sha256:9cdbc5edc17d2516eb2197ab8723fabae2340eb603b19cf4010995a43e33a51f

Observation 3ebed149-c399-4210-83ca-584ff496e8f1 · outbound

This paper cites Temporal Alignment Networks for Long-term Video.

Effectively obtaining acoustic, visual and textual data from videos Temporal Alignment Networks for Long-term Video

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:01:38.001715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.671184Z digest=sha256:8e86faece07bab6b2c530ecdf03b4baa33e53751d80d3b5ebb3669edf2ef0d4c

Observation f68d5808-7bb5-4605-a3ca-809bfd1eb156 · outbound

This paper cites Rae, and Laurent Sifre.

Effectively obtaining acoustic, visual and textual data from videos Rae, and Laurent Sifre

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.677051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.677051Z digest=sha256:13ef4d16d1f10297d99cd69744d2b0841ed1b95f1243d8c69abf1ef0c4abfe0c

Observation c08dc387-7c57-484b-95b9-115f89fbfe98 · outbound

This paper cites Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model.

Effectively obtaining acoustic, visual and textual data from videos Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.681822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.681822Z digest=sha256:543f705da4f79581ceb99133c7c2a933c9decaa094c0962a63a2551a1aeba9d3

Observation 1a400fbb-7282-4dcf-9cc2-f6a894c37edf · outbound

This paper cites Key Frame Selection for Temporal Graph Opti- mization of Skeleton-Based Action Recognition.Applied Sciences, 14(21), 2024.

Effectively obtaining acoustic, visual and textual data from videos Key Frame Selection for Temporal Graph Opti- mization of Skeleton-Based Action Recognition.Applied Sciences, 14(21), 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.686461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.686461Z digest=sha256:4de354667e5bca31ab40d14dc9e6d3ee59711fbbde486033720808603a461a5f

Observation e2619ad0-e7db-43b4-b75d-10b811b3eb7e · outbound

This paper cites NLIP: Noise-Robust Language-Image Pre-training.

Effectively obtaining acoustic, visual and textual data from videos NLIP: Noise-Robust Language-Image Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.691096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.691096Z digest=sha256:50dab691f6f3b6771c28b041a8ba7d17473217a74b0a1944efcfce3e65e71048

Observation 8030a54d-4642-4e07-ac9d-8312ffd4ec29 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.695555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.695555Z digest=sha256:ff85f3c121409280ed20a95b3c39013850002719511cf372c433811e6a64c1f7

Observation d5b226b3-6510-48a0-9811-c9febbe51052 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.700455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.700455Z digest=sha256:312b9c29c4a676228c4dd0b7dd123f8a587039a04a39c6fa493112c89491d7e7

Observation 947368cd-25dd-45e2-8879-1fd29ceb871a · outbound

This paper cites LL VIP: A Visible- infrared Paired Dataset for Low-light Vision.

Effectively obtaining acoustic, visual and textual data from videos LL VIP: A Visible- infrared Paired Dataset for Low-light Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.705388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.705388Z digest=sha256:5207bc923bdbcb8f01cb061964637772b2086377742544eb0dca68302226b630

Observation 93caa914-ba11-49d9-8659-e4964ef9e460 · outbound

This paper cites TimbreCLIP: Connecting Timbre to Text and Images.

Effectively obtaining acoustic, visual and textual data from videos TimbreCLIP: Connecting Timbre to Text and Images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.709931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.709931Z digest=sha256:d6d818b2e99016f9c89f3cae0a0871e4e78b19ffdece6a9c2262c526c9f8f62c

Observation 1f7baff1-df83-4c62-853a-066e913ca8a9 · outbound

This paper cites Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning.

Effectively obtaining acoustic, visual and textual data from videos Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.714574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.714574Z digest=sha256:aa633525e2ce78e53f1d0e056f5bb8e657e25d87537f91b8956b3a8ef0a71f49

Observation 78a115d0-1f88-48b5-ae44-df72cada4263 · outbound

This paper cites MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition.

Effectively obtaining acoustic, visual and textual data from videos MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:01:37.742120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.718947Z digest=sha256:e994f9379b486114924c777cf4917c4214770ef33a5d7bcced1fb7be45d514e0

Observation 5a4026dd-2113-40a0-9e48-6acf31061119 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.723724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.723724Z digest=sha256:27e16c22f2bfb0d0c3e67d689f2447350645eae2d3e37343570236ab71835d8c

Observation bc087ea3-9c7c-48f1-bec4-ac70b048557b · outbound

This paper cites AudioCaps: Generating Captions for Audios in The Wild.

Effectively obtaining acoustic, visual and textual data from videos AudioCaps: Generating Captions for Audios in The Wild

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.728086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.728086Z digest=sha256:899dd4bb371708fe42cc94e9fe80e798bf92bac133babcc9ae632694cc51cc1c

Observation 4e1e2dec-7704-4b8c-935e-200e8f2c817e · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Effectively obtaining acoustic, visual and textual data from videos Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.732735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.732735Z digest=sha256:bc5dcf554df575498e7e7ffdedf29c794bd7ae17d74bc5fb34f38a66fae2a47c

Observation a1260ad0-06c3-4bd2-87ee-c4e0c0a3b362 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Effectively obtaining acoustic, visual and textual data from videos AudioGen: Textually Guided Audio Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.737476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.737476Z digest=sha256:6f2dc4a928518f634d2fdef1b687309f8eb6f635e0e57a5cd9866e627b410d6d

Observation 40ea535a-3d4a-45f8-96d2-14b38f8a8f55 · outbound

This paper cites BindDiffusion: One Diffusion Model to Bind Them All, 2024.

Effectively obtaining acoustic, visual and textual data from videos BindDiffusion: One Diffusion Model to Bind Them All, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.742843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.742843Z digest=sha256:afc080af2d0ddc692db43e0467e361d0fc443784c4d2389fdd026ae97b4430d0

Observation e99ea385-2b6e-4480-b21e-ab166308de09 · outbound

This paper cites FLUX, 2024.

Effectively obtaining acoustic, visual and textual data from videos FLUX, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.748006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.748006Z digest=sha256:031671c23b64e2a2cc10bee7e3561c18c9dcd104c904bdc26d9f5034853f7c75

Observation 0faaade5-8f3f-4af3-a586-0d186385b1ab · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.752587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.752587Z digest=sha256:e11d5cb8453f5b37651182bff5774260ad8add32208ab34cfc64fcfbd271a02f

Observation a61c7e02-bda4-4af6-9a4d-542b7246ea2c · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.756811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.756811Z digest=sha256:3c813ae063f6428986e7b66ec2971f7b9ad36f54ec64f88c009bccd05c47e0c4

Observation feb64b77-e522-4dca-b0aa-798fe3fb0323 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.762891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.762891Z digest=sha256:28b32edf16e626f486be6133debf18b244331bde6673882a55da018abde7df28

Observation 23acdd7b-6340-44a0-a435-65b815bff5fc · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.767567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.767567Z digest=sha256:3acecc7d450495aa274d0ecedf516b9f6578382cb7f958b7f6c5f84da5a07ba2

Observation 566fad45-6b69-461c-9233-29cfb1f62629 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.772256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.772256Z digest=sha256:38bbb3b5e18ff56abfaab6b73dfa99214aaadf90f7d927bf426eb632e3d66486

Observation 61148e1b-803c-4613-b3ff-dd1dd8afa080 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.777288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.777288Z digest=sha256:38e18c79c1456b401aa9fe2b210f61c54acf8993a161dc5fc9fde4f14d41944d

Observation 5074745a-3fc1-4556-95af-4a061d8d6e5e · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.781925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.781925Z digest=sha256:82b570ceeacf6bfcda603095c65a4e08b8b9b47c537743b17f9834566577fa36

Observation dd476d19-c547-42a2-b863-03bd1bda0ba6 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.786535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.786535Z digest=sha256:cfd11c46e9f6c233fa6949d5b90fb9b1520546bea4a09554420af4d4f25c3b5b

Observation 77b59946-7569-4cfe-8d9f-76164e9c046d · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

Effectively obtaining acoustic, visual and textual data from videos OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.791275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.791275Z digest=sha256:0ce210dec97da6d0eb98bd9b3842c1b404732d9dcfa74aaa808567262f8bb188

Observation fc848145-e1df-42a1-9291-03dd77cb3ef1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Effectively obtaining acoustic, visual and textual data from videos BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.796174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.796174Z digest=sha256:cc3e17410beb97ed23c39acb08842872479ca3114afd3666261d36d38bf33d47

Observation 655e029f-4504-4828-b44b-460328908220 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Effectively obtaining acoustic, visual and textual data from videos BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.801004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.801004Z digest=sha256:ffba0714c780e1e096c94e5da867ef158318fb58c3b255a4fc869583c4f94782

Observation 02170d03-b9a9-48b0-9946-3de063dc4f0b · outbound

This paper cites Word-Level Explanations for Analyzing Bias in Text-to-Image Models.

Effectively obtaining acoustic, visual and textual data from videos Word-Level Explanations for Analyzing Bias in Text-to-Image Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.805782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.805782Z digest=sha256:a39dc811b30369891296cfa5d27f1a27abbc52dd4e27c45402464c4e5fedb236

Observation bac5b7e1-71c3-425a-ad31-eb94a85b5459 · outbound

This paper cites Lawrence Zitnick.

Effectively obtaining acoustic, visual and textual data from videos Lawrence Zitnick

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.810527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.810527Z digest=sha256:a750055306e106460f525ba8f00551c59d250b5f17ad41b2479e76a3ae3d7d8e

Observation 724e1ee9-d848-4c4b-99ea-2e5f3deec07c · outbound

This paper cites A Comparison Between KeyFrame Extraction Methods for Clothing Recognition, 2023.

Effectively obtaining acoustic, visual and textual data from videos A Comparison Between KeyFrame Extraction Methods for Clothing Recognition, 2023

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.815034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.815034Z digest=sha256:b1cf9a0f9b6d651df52e6f88f96e05e37773769fe99eea1af132007ec4a4eb5c

Observation 3644ee48-de0c-47c8-bb57-9eed6a07f0d3 · outbound

This paper cites Plumbley.

Effectively obtaining acoustic, visual and textual data from videos Plumbley

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.819846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.819846Z digest=sha256:fe86aaa6fdc309ce30e8afa05970e83491a268501d67c3302f08ed659209b806

Observation 734295a9-b47e-49ef-a458-dffdfe7d8398 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Effectively obtaining acoustic, visual and textual data from videos Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.824327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.824327Z digest=sha256:3f7253040133a68ae2c633ab6e79ca7f7faf1105b5bda9258b7c3f6cccf8cd37

Observation ba28e53f-7947-48ae-8690-28f0e7e5beb0 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.829712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.829712Z digest=sha256:4e396e3da42609be8a810a7807e40b023f7d62ff3b331033b3be3c4a9834c3e2

Observation 94c9feb5-863e-4d14-ba60-ec85cecbc01e · outbound

This paper cites BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning.

Effectively obtaining acoustic, visual and textual data from videos BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.835254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.835254Z digest=sha256:4fc413cdc7df4349f5de59e9adc52f93f36962f482f878d5f5a0bce08e5905ca

Observation e4fd3a40-1031-4c09-8502-3071e09ade49 · outbound

This paper cites Stable Diffusion Akashic Records, 2023.

Effectively obtaining acoustic, visual and textual data from videos Stable Diffusion Akashic Records, 2023

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.840392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.840392Z digest=sha256:6372d7df36bf8da2bfba65ac4239592fb6ce48cb7a6ec0215b8cd041e0dc9236

Observation 9af69ee2-41fc-4bb7-9a40-bea10e6dbff3 · outbound

This paper cites GenRL: Multimodal-foundation world models for generalization in embodied agents.

Effectively obtaining acoustic, visual and textual data from videos GenRL: Multimodal-foundation world models for generalization in embodied agents

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.845122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.845122Z digest=sha256:456ec3c36444e8d67695fa94746fa87e89e6ba3c51e99815c157bc8eac05f051

Observation dce89b1f-caf0-430f-a02a-6568fa5b4c7e · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Effectively obtaining acoustic, visual and textual data from videos Mustango: Toward Controllable Text-to-Music Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.849835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.849835Z digest=sha256:78ce777919fe1de40e441aa81ad3449d91560a6994f667fff59458bc117d755c

Observation e2bddc2e-62e9-4019-8434-618c75b22c38 · outbound

This paper cites Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis.

Effectively obtaining acoustic, visual and textual data from videos Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.854428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.854428Z digest=sha256:d69f972f219b0735e39e7207bd51915aa83e68a4d2f3f6d928a80953e2a0dbfa

Observation 95916d8e-de34-48b0-ba78-b829705542a5 · outbound

This paper cites DALL·E 3 System Card, 2023.

Effectively obtaining acoustic, visual and textual data from videos DALL·E 3 System Card, 2023

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.860309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.860309Z digest=sha256:55fc21c6820f844d3d9a192d7a273fe0b28fda62da88296bf85ac3938a632525

Observation f1955f34-dc82-4f77-897c-572158524ffe · outbound

This paper cites Video generation models as world simulators, 2024.

Effectively obtaining acoustic, visual and textual data from videos Video generation models as world simulators, 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.864946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.864946Z digest=sha256:8b92fa799200a5a678393753d4360ea3ca0ea2f23a0d51b7720432b89de4dfc6

Observation 232e3c2f-2a47-448c-ae2f-a304e18c9c82 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.869533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.869533Z digest=sha256:950ec28ba0aa3c0dc8f59fd78a594e51a7cef1556e2e92d3b0bf9008950351a2

Observation 42c840a9-50bc-43b6-9010-7f8bb454a298 · outbound

This paper cites Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022.

Effectively obtaining acoustic, visual and textual data from videos Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.874061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.874061Z digest=sha256:a0dc6a7a56d996c0cffba39e2825f9a231ab5f816fab101fd6f17ec55a02d792

Observation 071a64e6-495c-4522-bfe9-fbbd871374f7 · outbound

This paper cites AudioSetZSL, 2019.

Effectively obtaining acoustic, visual and textual data from videos AudioSetZSL, 2019

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.879208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.879208Z digest=sha256:5905e7fc12138fb74414acec1600e44135cf8ee49440addbd7398f0eb11ac928

Observation 2223b9a0-d786-4723-b199-03d5d455ed09 · outbound

This paper cites Coordi- nated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classifi- cation and Retrieval of Videos.

Effectively obtaining acoustic, visual and textual data from videos Coordi- nated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classifi- cation and Retrieval of Videos

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.884006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.884006Z digest=sha256:67a80a517d65a8ec00b2f4ef5757b931ecbfbd324490dfaaa49299883469c2f7

Observation 4ccd66b9-3ce4-4bb6-8ba1-378cccafe673 · outbound

This paper cites Pijanowski, Luis J.

Effectively obtaining acoustic, visual and textual data from videos Pijanowski, Luis J

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.888564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.888564Z digest=sha256:a54ed39ce1cc8834c5602ad5638fce0e0409aa9677365526c2775f67fdde56fb

Observation cbee204a-803b-4c13-bc2c-e870549cd3b2 · outbound

This paper cites Plummer, Liwei Wang, Chris M.

Effectively obtaining acoustic, visual and textual data from videos Plummer, Liwei Wang, Chris M

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.893458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.893458Z digest=sha256:dbdba52cf07b6e5c83108813f3a974a0a459909eef658236747ec5ad0d3f7fb3

Observation df676aed-bb72-4ac0-b295-29901731b91d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Effectively obtaining acoustic, visual and textual data from videos SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.898253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.898253Z digest=sha256:16bc30d24223fcffc6acddc780d1afe6c7b0ca8fcd0f055c1d5f4f8db3a99d22

Observation d13fe9b0-f7e7-44bd-b383-7d52c3a1a9aa · outbound

This paper cites Does mixing of speech signals comply with central limit theorem? International Journal of Electronics and Communications, 62(10):782–785, 2008.

Effectively obtaining acoustic, visual and textual data from videos Does mixing of speech signals comply with central limit theorem? International Journal of Electronics and Communications, 62(10):782–785, 2008

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.722935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.903031Z digest=sha256:665f554e25fc5ea05a753f5eaa5f778c8b6699f785332ad423f5b0f932d8ef93

Observation 27f08d89-5358-4c74-84b1-f2cb26e6d60a · outbound

This paper cites MirrorGAN: Learning Text-To-Image Generation by Redescription.

Effectively obtaining acoustic, visual and textual data from videos MirrorGAN: Learning Text-To-Image Generation by Redescription

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.907592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.907592Z digest=sha256:b7bfde3d818e1f98d5f5654cdb6f6d2b81bce6946eefe270882bbd4804e5fb0b

Observation e3c07864-0d86-489c-95ab-81798b370b51 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Effectively obtaining acoustic, visual and textual data from videos Learning Transferable Visual Models From Natural Language Supervision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.912294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.912294Z digest=sha256:eee43dcd4f1fed02f778f82d55e7a97b920e01899848b6516e1d4dfdd9aa6c32

Observation aedb676f-01de-41a3-996e-a5061d3e48ee · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Effectively obtaining acoustic, visual and textual data from videos Robust Speech Recognition via Large-Scale Weak Supervision

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.917064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.917064Z digest=sha256:4b86542372d245e3a4887e5a5164634a696087fd64aa133ef931cf15bee380c8

Observation 87d9060f-42ac-4c74-b9f8-de95ef1ef747 · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Effectively obtaining acoustic, visual and textual data from videos Zero-Shot Text-to-Image Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.921533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.921533Z digest=sha256:13c36ba93e81a10e002853d9a99a75daf108d40e01fb1bb47836adbb27972b55

Observation f0de62db-f4ac-4b36-87d6-36ac83a0c53d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Effectively obtaining acoustic, visual and textual data from videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.926838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.926838Z digest=sha256:256117d049d47cc48f5fbc914f7b2e98136d565d2aee42284f385931266e485d

Observation becfe866-8825-413b-8892-eb51594631fe · outbound

This paper cites Stable Diffusion, 2021.

Effectively obtaining acoustic, visual and textual data from videos Stable Diffusion, 2021

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.683239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.932052Z digest=sha256:252e79dddd0c0f140a54f309053ef57edbfb913196c1c6871cbb4dd3602098d9

Observation 0465db97-e607-4ada-a9ba-331eec2aa695 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Effectively obtaining acoustic, visual and textual data from videos High-Resolution Image Synthesis with Latent Diffusion Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.937443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.937443Z digest=sha256:a23d29bd71a511006948c7b7c4668da5b056b544bf849d7e202726414b5438a0

Observation 8e505e95-660f-4781-bda9-a9e2b108ebb9 · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024.

Effectively obtaining acoustic, visual and textual data from videos Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.942534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.942534Z digest=sha256:6a6a274243f5603f733baa393216cdc4b94f23f0904506e4fea1de65f7e3089f

Observation 74544f37-caf6-4c0a-ba80-96e6065ba4e1 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Effectively obtaining acoustic, visual and textual data from videos Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.947834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.947834Z digest=sha256:fa1890cc5440d8b3164249ac586fa7bc592fec05963a7556f8c6f7519843c27f

Observation 3fd18f94-87ce-447f-86a3-1f94f9743bd8 · outbound

This paper cites ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition.

Effectively obtaining acoustic, visual and textual data from videos ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.644398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.953007Z digest=sha256:7fdc2da416f7ef67ba6c54e08a81070a791ce9af804189aabbc2f24c25a95209

Observation 6c2abd28-5c1e-43b3-abfb-6a13cf33016d · outbound

This paper cites Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey.

Effectively obtaining acoustic, visual and textual data from videos Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.957524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.957524Z digest=sha256:53254788e8879ac1b35468844cf31b93dc0d6e17bf43d24f88c21798f3c94e8d

Observation 1e99fdc3-62d6-4ba8-beb5-859e76015007 · outbound

This paper cites What is noise?Geophysics, 63(4):1122–1124, 1998.

Effectively obtaining acoustic, visual and textual data from videos What is noise?Geophysics, 63(4):1122–1124, 1998

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.628464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.962343Z digest=sha256:29edb2c6667c8cb1d7a113661a4120e912fca36bc531983a0ec37889bcb9baeb

Observation f86c7b08-e085-4929-a849-b350da166928 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Effectively obtaining acoustic, visual and textual data from videos LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:01:38.612113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T05:01:36.967095Z digest=sha256:4222fd4807fcb7069c8a686092040ccd8829f34722548766cfd2e329a1b8d651

Observation 1141ad1b-87bf-4cf4-8205-39e917d797f9 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Effectively obtaining acoustic, visual and textual data from videos LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.971628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.971628Z digest=sha256:7b5b284f169757ff27b9302829ceb2718856536ee0972754984a311c3d98c572

Observation 181fc8f0-d346-40ef-867e-3e508f577b82 · outbound

This paper cites an unresolved cited work.

Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.976501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.976501Z digest=sha256:d7bedde79315bee4d0ce7357a5a6f915139a77e383789e39ca95ae7f208f0cc4

Observation 7bae9aa2-c149-499e-a4ec-8a463198ecb6 · outbound

This paper cites I Hear Your True Colors: Image Guided Audio Generation.

Effectively obtaining acoustic, visual and textual data from videos I Hear Your True Colors: Image Guided Audio Generation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.981088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.981088Z digest=sha256:8ad71eb725a35a8be2dcc6f6e2680c78381f2e26df9703257847031877936fe8

Observation b956306a-aaff-4ab1-bd9c-34721c77bd74 · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Effectively obtaining acoustic, visual and textual data from videos A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.985927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.985927Z digest=sha256:6b5bc5f57e6bb0316347e79dc1aba05772f0665f7ed3379bbe424a2bb6c7c7b8

Pith citing papers

Observation 7f772987-5f54-46cf-9866-ec6efe4b5788 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Effectively obtaining acoustic, visual and textual data from videos

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:25:28.242608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T21:25:26.928879Z digest=sha256:d7dca520f88ad0bf457b36eea916210a258fea90aedb3a8985da598764146a1c