Pith. sign in

Paper Citation Record · LEDGER

From Sound to Sight: Towards AI-authored Music Videos

As of 14 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2509.00029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00029 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T18:23:31.599121Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35b71881-cbfc-4364-bd1e-bc64100746e9 · outbound

This paper cites Secure & Personalized Music-to- Video Generation via CHARCHA, 2025.

From Sound to Sight: Towards AI-authored Music Videos Secure & Personalized Music-to- Video Generation via CHARCHA, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:42.054602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:26.897893Z digest=sha256:aa47bf6658e6b48c138158538fd13c3b3c7df5595995ca21e8f9e90281a36ae2

Observation a61a3c3b-bbed-48d2-94ba-e2d358488d00 · outbound

This paper cites ImproveYourVideos: Architectural Im- provements for Text-to-Video Generation Pipeline.IEEE Ac- cess, 13:1986–2003, 2025.

From Sound to Sight: Towards AI-authored Music Videos ImproveYourVideos: Architectural Im- provements for Text-to-Video Generation Pipeline.IEEE Ac- cess, 13:1986–2003, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:41.898128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:26.959917Z digest=sha256:29cb56c201a178e0c8d10c50a738fd958770bd944f389ae3c848cada3a6f67ac

Observation e8302f5f-55e2-4c94-90e2-e384f004d410 · outbound

This paper cites The MIT Press, Cambridge, Massachusetts, 2021.

From Sound to Sight: Towards AI-authored Music Videos The MIT Press, Cambridge, Massachusetts, 2021

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:41.702044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.037973Z digest=sha256:77bb814c76b3527e5fffb43745b7ca0586ced3896e411255151088f3556f1c45

Observation 5ba3e4db-b1de-4f28-aca2-68fa68d69f68 · outbound

This paper cites Boden and Ernest A.

From Sound to Sight: Towards AI-authored Music Videos Boden and Ernest A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:41.545856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.112468Z digest=sha256:a43ed2d76ec90bed540068257427cf263e6d2be08fff2a276d46651daf7cd394

Observation bdad2ee3-a30d-475a-b094-ded1e023e0cd · outbound

This paper cites Review of Gottschall (2012): The story- telling animal: How stories make us human.Scientific Study of Literature, 2(2):317–321, 2012.

From Sound to Sight: Towards AI-authored Music Videos Review of Gottschall (2012): The story- telling animal: How stories make us human.Scientific Study of Literature, 2(2):317–321, 2012

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:41.416285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.193693Z digest=sha256:9fb7c08407bd866980da3e81e5158b5cc46fd65ced53d840a7c4971268537047

Observation c291cf6f-437e-4aab-912a-8393c68895a1 · outbound

This paper cites Diffusion Models as Artists: Are we Closing the Gap between Humans and Machines?.

From Sound to Sight: Towards AI-authored Music Videos Diffusion Models as Artists: Are we Closing the Gap between Humans and Machines?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-05T18:23:32.266214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.331853Z digest=sha256:49feabf23e3b1c56aa8d2a79ca7ece680c5c7f73fc412e15ecf097682d4e721d

Observation 4c861e84-251a-4ec6-a422-bb2a7f26388b · outbound

This paper cites Crossmodal associations between naturally occurring tactile and sound textures.Per- ception, 53(4):219–239, 2024.

From Sound to Sight: Towards AI-authored Music Videos Crossmodal associations between naturally occurring tactile and sound textures.Per- ception, 53(4):219–239, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:41.268228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.415389Z digest=sha256:2dcc4f940548ae9911bc33ccfa70fdb3ebb8e54854978c49fad787f99729ab95

Observation 1383f726-cf8c-4abb-b8f7-54389b6414c0 · outbound

This paper cites Cancino-Chac ´on, Maarten Grachten, Werner Goebl, and Gerhard Widmer.

From Sound to Sight: Towards AI-authored Music Videos Cancino-Chac ´on, Maarten Grachten, Werner Goebl, and Gerhard Widmer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:41.039657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.487946Z digest=sha256:99ce8f5a0775703ae6581647943545129e9f09f746fb8571e3101a95e1f3cbca

Observation edca7429-49e1-4a90-b226-69ee5cb045ed · outbound

This paper cites ”scary robots”: Examining public responses to ai.

From Sound to Sight: Towards AI-authored Music Videos ”scary robots”: Examining public responses to ai

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:40.868573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.566047Z digest=sha256:e6c0eb5d48318d11d597866704e38ff061fc061c1feeb0945ea8f3bb91dec061

Observation b0ca1abb-0212-4b0c-b9df-58d12a4c64f9 · outbound

This paper cites Understanding and Creating Art with AI: Review and Outlook.

From Sound to Sight: Towards AI-authored Music Videos Understanding and Creating Art with AI: Review and Outlook

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T18:23:32.063832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.698936Z digest=sha256:8c1dd75f4208cc0e765cb9d62e13998659cba4f5c40265056401accdbb18c2e2

Observation cebe88bd-f23b-4be7-8dd7-154a2ad1f1e6 · outbound

This paper cites Dasovich-Wilson, Marc Thompson, and Suvi Saarikallio.

From Sound to Sight: Towards AI-authored Music Videos Dasovich-Wilson, Marc Thompson, and Suvi Saarikallio

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:40.679610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.784957Z digest=sha256:f5679ef18cd2c31ca72960460413e7454a459a8364a1c17ff7e25279588a6101

Observation 54998d99-14d2-4bad-bb6b-7cf9c35eff00 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

From Sound to Sight: Towards AI-authored Music Videos DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:27.845221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:27.845221Z digest=sha256:2210abb9eac4f0a737498316d2181c48608a8f58e703d30d8a413c37f74721c4

Observation 2b292f5d-f0b2-440a-b3b8-14ef5615f0c2 · outbound

This paper cites Pengi: an audio language model for audio tasks.

From Sound to Sight: Towards AI-authored Music Videos Pengi: an audio language model for audio tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:40.508262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.892730Z digest=sha256:9622c1d6f1a49f377b2068e4b9f7bc7622369b87819230f861a6b1a78c5f8991

Observation 5438db99-ddb5-4b7e-b188-6f788828f549 · outbound

This paper cites Berkeley Publishing Group, New York, New York, 2005.

From Sound to Sight: Towards AI-authored Music Videos Berkeley Publishing Group, New York, New York, 2005

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:40.316719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:27.988577Z digest=sha256:d6df55ce67ceabdd030c42af64008cdd86f92d47686b95268cc85983cbff18df

Observation ec55e3db-b401-425e-a0c1-e5f396de7d9e · outbound

This paper cites EasyVid AI Video Maker, 2024.

From Sound to Sight: Towards AI-authored Music Videos EasyVid AI Video Maker, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:40.146291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.068936Z digest=sha256:e4ed09777950bc42cb408631bdc42f449e98b2e69fe5714b86c2b57b3a198db5

Observation f7ef522e-3c84-4be7-90a0-ec5035bf86e8 · outbound

This paper cites CAN: Creative Adversarial Networks, Generating "Art" by Learning About Styles and Deviating from Style Norms.

From Sound to Sight: Towards AI-authored Music Videos CAN: Creative Adversarial Networks, Generating "Art" by Learning About Styles and Deviating from Style Norms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:28.167494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:28.167494Z digest=sha256:ccb2253663b543787058b524f67716a3aca5c7cfc587c438acec01d37642cc8f

Observation 690b92dd-d1ba-415f-b0e0-47f770d65da6 · outbound

This paper cites CLAP Learning Audio Concepts from Natural Language Supervision.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2023.

From Sound to Sight: Towards AI-authored Music Videos CLAP Learning Audio Concepts from Natural Language Supervision.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:39.930859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.237066Z digest=sha256:5505faf1065e264029e2450afa2e50846cc2445d2c5dd0bb8f79312610e76a06

Observation d0d87f1c-8fa1-40e8-9acc-a015489e44bd · outbound

This paper cites Rand, and Iyad Rah- wan.

From Sound to Sight: Towards AI-authored Music Videos Rand, and Iyad Rah- wan

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:39.659126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.330599Z digest=sha256:0401526a861386f80926cc53e79b102bd440922cdd0274c4a68060f265f97372

Observation 28a48ced-bc6f-4edc-a4ae-bd6ef9e7845e · outbound

This paper cites Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith.

From Sound to Sight: Towards AI-authored Music Videos Frank, Matthew Groh, Laura Herman, Neil Leach, Robert Mahari, Alex “Sandy” Pentland, Olga Russakovsky, Hope Schroeder, and Amy Smith

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:39.392080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.455570Z digest=sha256:f3255259539ed5080fcbe8c3ca85cc04f0054d3e17c4b51b9be2d7e86c236a76

Observation b03fa92e-d669-4b4c-b3cc-618d47076bdd · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

From Sound to Sight: Towards AI-authored Music Videos GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:28.522442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:28.522442Z digest=sha256:6eabd8617e9cc3b5475539d9d3c2e792ee0782186cd4cf0c97c9f9ed0bbc6283

Observation 6a5f500d-6737-4254-a0b2-d12bd44783da · outbound

This paper cites From Ragtime to Swingtime: Fifty Glittering Years of Stage and Song.

From Sound to Sight: Towards AI-authored Music Videos From Ragtime to Swingtime: Fifty Glittering Years of Stage and Song

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:39.058349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.614731Z digest=sha256:1021427dd6b9ef5ca183fb4c04d44a964ecb18796603937e45a4ae615a50bc2b

Observation cb10c913-26cb-4406-ba05-edefb477c907 · outbound

This paper cites Generative adversarial networks.Commu- nications of the ACM, 63(11):139–144, 2020.

From Sound to Sight: Towards AI-authored Music Videos Generative adversarial networks.Commu- nications of the ACM, 63(11):139–144, 2020

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:28.676982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:28.676982Z digest=sha256:37fa3b05dd519c2860fb0bd1461491b6aeeca1d95956445033bfb3eb830e3b60

Observation d5233c9e-2fff-4183-9f49-a040759ab52c · outbound

This paper cites Graesser, Murray Singer, and Tom Trabasso.

From Sound to Sight: Towards AI-authored Music Videos Graesser, Murray Singer, and Tom Trabasso

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:38.781115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.760858Z digest=sha256:9932b30d9cb42da6dfb08fde328d7c95047bbaa6390525decfc2c669e5836182

Observation 8e3b659f-8a81-491c-9cbb-9245ba06337e · outbound

This paper cites Beware of fictional ai narratives.Nature Machine Intelligence, 2(11):654–654, 2020.

From Sound to Sight: Towards AI-authored Music Videos Beware of fictional ai narratives.Nature Machine Intelligence, 2(11):654–654, 2020

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:38.617464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.852552Z digest=sha256:10f62968eb8deba480d9801399b6530e46c5dc95afd1d23035f72f8d86060752

Observation c7a6a6e8-13fa-4429-99e4-29a1ce06a8b9 · outbound

This paper cites Can Computers Create Art?.

From Sound to Sight: Towards AI-authored Music Videos Can Computers Create Art?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T18:23:31.876234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:28.947855Z digest=sha256:36503614efdd210b7f70ff9015ba7336de6487088d51a7df73cd797cd72112a9

Observation 88c7f4d6-e73e-4761-89ed-973b8d6cd02f · outbound

This paper cites Computers do not make art, people do.

From Sound to Sight: Towards AI-authored Music Videos Computers do not make art, people do

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:38.444000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.024474Z digest=sha256:114d148a492dc9d071f0c90796c084277dc2e370bf2d6cd739c79152b805e7ca

Observation 8474ac28-c3e5-4463-9bb9-edf5a15669f1 · outbound

This paper cites Cascaded Diffusion Models for High Fidelity Image Generation.

From Sound to Sight: Towards AI-authored Music Videos Cascaded Diffusion Models for High Fidelity Image Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:29.137643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:29.137643Z digest=sha256:651e8ff2805f04da9f328862497860b372128ec7f9c071565fa12a29020231d2

Observation 2ece083d-74c1-4451-ab72-1d3abc6687e3 · outbound

This paper cites Video Diffusion Models.

From Sound to Sight: Towards AI-authored Music Videos Video Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:29.193146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:29.193146Z digest=sha256:b3660c52e133c66e832de88a23e006fa81fd72462338aba336c3f0811710f9e2

Observation 4240c902-f576-46d6-a20b-3c28a669161d · outbound

This paper cites Artificial Intelli- gence, Artists, and Art: Attitudes Toward Artwork Produced by Humans vs.

From Sound to Sight: Towards AI-authored Music Videos Artificial Intelli- gence, Artists, and Art: Attitudes Toward Artwork Produced by Humans vs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:38.287509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.276918Z digest=sha256:4bcdf77830663c13b4966ec0154bbb08e18b644460073ca36b3fe667259c19e4

Observation 503c4c33-4102-42e1-8021-e4b0c846cba7 · outbound

This paper cites Blaine Horton Jr, Michael W.

From Sound to Sight: Towards AI-authored Music Videos Blaine Horton Jr, Michael W

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:38.054367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.328957Z digest=sha256:46dff1171657a6ac4548a31dbc59b22f671d6ac1459d844cb8dbaa209001af7f

Observation 41388663-6713-471c-ac63-3afa9d8e79ef · outbound

This paper cites VBench: Com- prehensive Benchmark Suite for Video Generative Models,.

From Sound to Sight: Towards AI-authored Music Videos VBench: Com- prehensive Benchmark Suite for Video Generative Models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:29.398758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:29.398758Z digest=sha256:a83702ad8c1ad7df1948c9212753ad01aaf8c202c9de56d3e21cb9602f681eb5

Observation 6b1ec288-881f-4ad2-baf4-f984b31c6e91 · outbound

This paper cites Cross-modal associations between paintings and sounds: Ef- fects of embodiment.Perception, 51(12):871–888, 2022.

From Sound to Sight: Towards AI-authored Music Videos Cross-modal associations between paintings and sounds: Ef- fects of embodiment.Perception, 51(12):871–888, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:37.848811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.573162Z digest=sha256:7f406b0b45ef2b51bcb026b18d780f7bce653d7dea64a9c8aa324adc3047b1dc

Observation 15bb7bbc-1eac-47d0-83b2-3fd9b6249836 · outbound

This paper cites Kaiber AI: Generating Videos with Superstu- dio, 2025.

From Sound to Sight: Towards AI-authored Music Videos Kaiber AI: Generating Videos with Superstu- dio, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:37.644475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.639115Z digest=sha256:a406a6052be65f6413201ddd9f227e090240f0feed1b8e95bb75cb6f2c89bfd7

Observation 8704e49a-128b-44ba-a303-d5f9636c681a · outbound

This paper cites Artificial Intelligence and Copyright: Le- gal Quandary in the Digital Age: Some Musings.SSRN Elec- tronic Journal, 2021.

From Sound to Sight: Towards AI-authored Music Videos Artificial Intelligence and Copyright: Le- gal Quandary in the Digital Age: Some Musings.SSRN Elec- tronic Journal, 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:37.465831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.678001Z digest=sha256:7642ca4e579757dcd10a011854e55d2ad7e189c95f46c6a714e786378741ea76

Observation 698aae6f-f86f-4f8e-b48d-636c01161c93 · outbound

This paper cites ”AI enhances our performance, I have no doubt this one will do the same”: The Placebo effect is ro- bust to negative descriptions of AI.

From Sound to Sight: Towards AI-authored Music Videos ”AI enhances our performance, I have no doubt this one will do the same”: The Placebo effect is ro- bust to negative descriptions of AI

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:37.288366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.734159Z digest=sha256:8796cef18d9c61bf1c00ae8c1f7ed4e1052c5be03b0e8dfd23af330da99e77fa

Observation 5fa8b5dc-de0b-45a2-b11c-af5887acc0f9 · outbound

This paper cites The (R)evolution of Music Video in American Music Industry.New Horizons in English Stud- ies, 8:163–176, 2023.

From Sound to Sight: Towards AI-authored Music Videos The (R)evolution of Music Video in American Music Industry.New Horizons in English Stud- ies, 8:163–176, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:37.075736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.780560Z digest=sha256:8cbf3c57c95579a2a6420d57920cb8a2de47411f163b9c061407cc48974436a2

Observation 022a4c7f-7d85-4e6e-9fd6-3b600fcd24f6 · outbound

This paper cites The Placebo Effect of Artificial Intelli- gence in Human–Computer Interaction.ACM Transactions on Computer-Human Interaction, 29(6):1–32, 2022.

From Sound to Sight: Towards AI-authored Music Videos The Placebo Effect of Artificial Intelli- gence in Human–Computer Interaction.ACM Transactions on Computer-Human Interaction, 29(6):1–32, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:36.853711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.812172Z digest=sha256:737fe510d91ed28344482558cc72246e4ae023a56ff31662e924c7f2666328dd

Observation 4be64de3-089b-4070-8776-a39526b934c4 · outbound

This paper cites Robust One Shot Audio to Video Generation.

From Sound to Sight: Towards AI-authored Music Videos Robust One Shot Audio to Video Generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:36.615800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.912793Z digest=sha256:9479401aeee8a02f4fb8210c6c0a06c0cdc21e86546b3ae388f73e8a29bcf559

Observation 74bd9a09-60a2-4fcf-95e4-2f5fcce6270f · outbound

This paper cites an unresolved cited work.

From Sound to Sight: Towards AI-authored Music Videos Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-05T18:23:36.408063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:29.975981Z digest=sha256:ca33a568ca4e6cac1552aa4f6a91f33657250ebf3f903af8099a45748acfd85a

Observation 3d797bd9-75fa-468f-a6c2-3c4459d4bd54 · outbound

This paper cites Lima, Carlos G.

From Sound to Sight: Towards AI-authored Music Videos Lima, Carlos G

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:36.189412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.043470Z digest=sha256:63b785bce7a033e90266a98fc61cf55cc1c77bd9e9f79f321cf7fb39047a9fc5

Observation ac8a5eb0-6a88-4dae-92d6-665e84ccc9e7 · outbound

This paper cites In ai we trust? effects of agency locus and transparency on uncertainty reduction in human–ai interac- tion.Journal of Computer-Mediated Communication, 26(6): 384–402, 2021.

From Sound to Sight: Towards AI-authored Music Videos In ai we trust? effects of agency locus and transparency on uncertainty reduction in human–ai interac- tion.Journal of Computer-Mediated Communication, 26(6): 384–402, 2021

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:35.961495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.099710Z digest=sha256:e600ca25991531e01728edb607f425ce41d1a2d1a604d4a9113524ca6d6a0370

Observation 3230a59a-0e88-4646-b8e0-6bce11a5b60c · outbound

This paper cites Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation.

From Sound to Sight: Towards AI-authored Music Videos Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:35.771559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.149947Z digest=sha256:5d55004d03c7f37c3ee36fc6b088790be5421da5449542a6f78f646ea0e2754f

Observation f95758a0-0e62-46ba-b3df-1982fbfe6de0 · outbound

This paper cites Are Emergent Abilities in Large Language Models just In-Context Learning?.

From Sound to Sight: Towards AI-authored Music Videos Are Emergent Abilities in Large Language Models just In-Context Learning?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:30.219156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:30.219156Z digest=sha256:879b1a3416d22aa6d4b4f8df44d1f96feb3e5da7b9cf9da1c005ff3fa060c6b7

Observation 00238615-afab-462f-8c81-689dc8f2f935 · outbound

This paper cites Art, Creativity, and the Potential of Artificial Intelligence.Arts, 8(1):26, 2019.

From Sound to Sight: Towards AI-authored Music Videos Art, Creativity, and the Potential of Artificial Intelligence.Arts, 8(1):26, 2019

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:35.597409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.292545Z digest=sha256:bc179dc4396e8b75fc8696665b4b70b97ee907cc2ea3764568a5b999b6733c8a

Observation 538e82b9-738e-4623-80bd-8ac68965494d · outbound

This paper cites Windows Media Player, 2025.

From Sound to Sight: Towards AI-authored Music Videos Windows Media Player, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:35.424087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.365220Z digest=sha256:90d9a572390b1e4660bbed5fb004228626e460b4f9f7e196a1c9103f77aa0ff0

Observation d9d86d38-155b-4fab-9b62-ab2753341979 · outbound

This paper cites Com- putational Music Structure Analysis (Dagstuhl Seminar 16092).

From Sound to Sight: Towards AI-authored Music Videos Com- putational Music Structure Analysis (Dagstuhl Seminar 16092)

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:35.253847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.413253Z digest=sha256:729bec5c0137d81507c7c1dde5a567a7a89012a8c69cb9330383ad35ded3001f

Observation 8439a621-4d74-40e0-b9f8-c498a573d51c · outbound

This paper cites AI Music Video Generator, 2025.

From Sound to Sight: Towards AI-authored Music Videos AI Music Video Generator, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:35.081445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.475329Z digest=sha256:adc4061fd3778a2c89f7b9dd04da650c681d3d3482b59a1829b052072ed95593

Observation 4b2417f8-7576-4bb6-a7bd-c4c2be20e346 · outbound

This paper cites Hello GPT-4o by OpenAI, 2025.

From Sound to Sight: Towards AI-authored Music Videos Hello GPT-4o by OpenAI, 2025

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:34.888134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.514243Z digest=sha256:1eea00a8d15f903fa8e28893f814a342a601c88cfc39c683636236e50ca7e532

Observation d06b6828-bac5-48b9-92cf-986da4e8dcf0 · outbound

This paper cites Synaesthesia.European Neurology, 57(2):120– 124, 2007.

From Sound to Sight: Towards AI-authored Music Videos Synaesthesia.European Neurology, 57(2):120– 124, 2007

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:34.657679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.606125Z digest=sha256:4de94b055277925b674152439c890c982307bb07d2d129ac5c68492d4bca1a78

Observation 4476f215-5682-4707-a602-e679e3427c66 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision,.

From Sound to Sight: Towards AI-authored Music Videos Robust Speech Recognition via Large-Scale Weak Supervision,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:34.467417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.667108Z digest=sha256:5a294ab93e990ae2adf077af2706ed1203287afab20cbe6c3793b562054872e9

Observation 1e2aeb90-a502-4ce8-9300-b14926c3376d · outbound

This paper cites Self-supervised Dance Video Synthesis Conditioned on Mu- sic.

From Sound to Sight: Towards AI-authored Music Videos Self-supervised Dance Video Synthesis Conditioned on Mu- sic

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:34.356357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.785288Z digest=sha256:671527f1825e77236e5b5e597a2a5a638e435e2be0d305de8d19c88ddd231c65

Observation 932469f3-9384-422f-b1db-2cbf574f3cf6 · outbound

This paper cites High-Resolution Image Synthesis With Latent Diffusion Models.

From Sound to Sight: Towards AI-authored Music Videos High-Resolution Image Synthesis With Latent Diffusion Models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:34.209998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.911969Z digest=sha256:99732f1f0d7f2c494d568473ff2b72ebf089d29e096d9b8fdf5d623088b0e915

Observation c9259e24-855a-4884-84ad-ea6c0f4f8482 · outbound

This paper cites Oxford University PressNew York, NY , 1995.

From Sound to Sight: Towards AI-authored Music Videos Oxford University PressNew York, NY , 1995

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:34.033741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:30.966040Z digest=sha256:399f1e5200b4570c4e8a27dd8f78ac2a97d2fa9d00108721cec4afe590ebdca9

Observation 329a479d-5675-4df6-b4d0-b612c99fed59 · outbound

This paper cites Machine Learning Processes As Sources of Ambiguity: Insights from AI Art.

From Sound to Sight: Towards AI-authored Music Videos Machine Learning Processes As Sources of Ambiguity: Insights from AI Art

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:33.859082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.056150Z digest=sha256:fcccb5b2e2680340968a76147491b4352ecdaac6eb38c26f5b22225a1f2dca55

Observation 26ef0629-2ced-490d-b627-9b126447c808 · outbound

This paper cites Mochi 1 by Genmo Team, 2024.

From Sound to Sight: Towards AI-authored Music Videos Mochi 1 by Genmo Team, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:33.748561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.097961Z digest=sha256:760838579ef54e38f88be658c14d3ae1da9ddc1c5ba4797ae46b121c32adca18

Observation 6efd3b9f-d5b5-4315-8054-5d2f427bca2f · outbound

This paper cites Revid.ai, 2025.

From Sound to Sight: Towards AI-authored Music Videos Revid.ai, 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:33.639793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.151817Z digest=sha256:48eb09424bbe2d89665140bf33af0b4f757d0e600042cb80d22859c3063bf367

Observation 14680611-1f71-4210-b9cf-19d4446e1bcf · outbound

This paper cites Specterr: Music Video Maker Online, 2025.

From Sound to Sight: Towards AI-authored Music Videos Specterr: Music Video Maker Online, 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:33.425690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.179963Z digest=sha256:0597afdf2396c39b110151add6abc95da12d3084fcc105f364b4999621d990fc

Observation e2ae8bdf-dd3f-46a0-813c-da86bef81a7f · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

From Sound to Sight: Towards AI-authored Music Videos Wan: Open and Advanced Large-Scale Video Generative Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:31.222822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:31.222822Z digest=sha256:9591af4e1a6bb9ac087e51223fc1d3ac4f7433e1c8fc6e52b4a4e796c1003c35

Observation 41e18c0c-b122-424c-b254-b67062be121a · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

From Sound to Sight: Towards AI-authored Music Videos Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:31.278343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:31.278343Z digest=sha256:07bc89b731a4492faa385473ce194ac16818ed0a8f63cf2e6a55ec11eb3ed310

Observation e7073e7c-785f-4103-891a-2cb15cfe1d07 · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

From Sound to Sight: Towards AI-authored Music Videos A Survey on Knowledge Distillation of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:31.358582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:31.358582Z digest=sha256:478d7a1e8b2bf52bd251ce783f4f4c8ca732ec51e814145be39dcad8ac6f5dd9

Observation 49183d08-19ef-4db6-8e03-37ab707352bf · outbound

This paper cites Wordcraft: Story Writing With Large Language Models.

From Sound to Sight: Towards AI-authored Music Videos Wordcraft: Story Writing With Large Language Models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:33.201784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.419333Z digest=sha256:747af4c1d3af759d72faf830fcc1c89eb4750bfeec997213b2e7204b321594df

Observation 9634ef36-91ee-42ab-a47c-4f8eb2540052 · outbound

This paper cites Let’s Play Music: Audio-Driven Performance Video Generation.

From Sound to Sight: Towards AI-authored Music Videos Let’s Play Music: Audio-Driven Performance Video Generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:33.006421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.483700Z digest=sha256:912d1ced16ec7a5c52cfcf60f1849309b2fa2e79b8ada45a4ba38cc173a0563e

Observation 298bad70-e43a-43d4-9bcb-25f33bafe350 · outbound

This paper cites SCENE #:.

From Sound to Sight: Towards AI-authored Music Videos SCENE #:

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:32.741767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.551008Z digest=sha256:b66add2ba39b66a5030473de5bb4ac9f0e3435db68ef814d3d222c0272f8a0a7

Observation 49386794-f5d9-4d91-b458-d8ee3e6eee88 · outbound

This paper cites All items were rated on a 7-point Likert scale: where 1 = Strongly Disagree, and 7 = Strongly Agree.

From Sound to Sight: Towards AI-authored Music Videos All items were rated on a 7-point Likert scale: where 1 = Strongly Disagree, and 7 = Strongly Agree

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T18:23:32.510045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T18:23:31.599121Z digest=sha256:c498ea4cc47feaea3b45bbb6fe98f8cad62dc3da02137f760866e4f5ddc7b84f

Observation 93003541-2ae2-4dd4-8393-3136a5382996 · outbound

This paper cites an unresolved cited work.

From Sound to Sight: Towards AI-authored Music Videos Unresolved cited work

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:30.852381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:30.852381Z digest=sha256:87fc041d9e83f4868b7169f5947f99a6f4bacbe85d045be498cffa20ece9fd87

Observation c7a423f1-b732-419d-a460-ae3e23fc4ffa · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

From Sound to Sight: Towards AI-authored Music Videos Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:30.736114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:30.736114Z digest=sha256:62006a23191844e33ec194b83663d5ed7f1751662f355fe589015e79e76d8ef0

Observation 24b92ba4-8eb7-4a3e-897e-9e3f3a89ca43 · outbound

This paper cites VBench: Comprehensive Benchmark Suite for Video Generative Models.

From Sound to Sight: Towards AI-authored Music Videos VBench: Comprehensive Benchmark Suite for Video Generative Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T18:23:29.510645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:23:29.510645Z digest=sha256:32ec790be958c2e89b3e5b4fe17ff2331945f55592fdb638d7ce77a638da25f4

Pith citing papers

No inbound Pith citation observations are available.