Pith. sign in

Paper Citation Record · LEDGER

Testing chatbots on the creation of encoders for audio conditioned image generation

As of 16 August 2026, this Paper Citation Record lists 100 of 123 outbound references and 0 inbound Pith citation observations for arXiv:2509.09717.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09717 v1

Coverage vector

measured 100 of 123 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:25:27.288300Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 123 outbound references displayed

  • verified exact5
  • verified fuzzy22
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2773827-75a3-47c3-820d-11bd236c39a0 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.568746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.568746Z digest=sha256:e67ba1a52fa42aa052ff957b79c8398bedbb6576162e513a8edbdce805b0af78

Observation a3196e02-8a17-4eae-b2ca-7b4c4fccbee3 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.575691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.575691Z digest=sha256:4ef87c10947e90aa1f39e08759779c86e583ff4c70bf10b933468614af1a1fbb

Observation a05e9955-7c72-4f18-a048-e542adba6200 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.582775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.582775Z digest=sha256:0da35465ae262a06b69a946ca8f8f4670482f04c1d2db5d4c65fe7a236c623e6

Observation 3e732957-04b3-454e-a69a-2081c3d8a33a · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.590262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.590262Z digest=sha256:796e1204c0d2b2af5a32efacadd03c0f1322b86daf08743c60a3f7baaab7d7cd

Observation 73eda8d4-c2cc-4b6a-8a19-cbb7a2817182 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.596268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.596268Z digest=sha256:3426ae8bb66745c9c667319cf8b9edabdd4056c919cfe5f7640928385367e487

Observation fdecdf9e-7d72-433a-a58f-cc59b3d5ccaa · outbound

This paper cites MusicLM: Generating Music From Text.

Testing chatbots on the creation of encoders for audio conditioned image generation MusicLM: Generating Music From Text

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.603213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.603213Z digest=sha256:2396a57131009b1bfefa5734f14113753c8663984703a6050d892663be741aee

Observation 30b1b9fa-3cbc-4db3-8af1-1fa9a452d78d · outbound

This paper cites Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering.

Testing chatbots on the creation of encoders for audio conditioned image generation Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.610281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.610281Z digest=sha256:ed77a0bf4e7759f5534d4415794093a665e0c158bc89c9f24f6849af20bbbdd6

Observation 9dc492ba-08cb-4972-9fc4-c7814d9eb3fa · outbound

This paper cites Mistral Models, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Mistral Models, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.616045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.616045Z digest=sha256:3f360771f0ccf384f16b9ff68ae52ebeafaa2467a6166081f0939db2fe0d57e0

Observation a0aa1203-489f-4ccc-a811-9e9ff8a99709 · outbound

This paper cites Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019.

Testing chatbots on the creation of encoders for audio conditioned image generation Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.622613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.622613Z digest=sha256:6b770af19b341ecec3408d108abb0c891c2bbfd321121ffd8919c64b70a6b3a6

Observation f6c6ed46-b1e8-4bda-80ed-e5bb48a66962 · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.630294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.630294Z digest=sha256:2a6792f63f23195889b26bee55e984a6257735cef0fe6c074da844da4ae75d22

Observation 264247d8-e039-464b-8b52-798c1768a358 · outbound

This paper cites Claude 3.7 Sonnet and Claude Code, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation Claude 3.7 Sonnet and Claude Code, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.635144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.635144Z digest=sha256:241fde57363946e22d02a79a2c911e76c78b815736abb9e2879d75cafc0f109f

Observation af19f626-c13b-44ea-a39a-4bc7577bf874 · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.641339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.641339Z digest=sha256:5608f2c2ba23673887058d04596f147c7e3caa4a624694e82435c105559fa281

Observation 1b3e2439-fd12-4a7b-b049-34ec221eb8f9 · outbound

This paper cites Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.649251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.649251Z digest=sha256:6268c92d37da7c5f81c1b6336567fff50c61ab36f64f61fc65d1674b3259521b

Observation c9e1f668-242a-4ffd-b03c-3a86f74b18cd · outbound

This paper cites Ballester.

Testing chatbots on the creation of encoders for audio conditioned image generation Ballester

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.654893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.654893Z digest=sha256:8d8f83498052e44d5724ef7ed565c41d857a1d75f86bf244f8aa5ca02180200f

Observation 1fd800a7-ff39-4098-8915-37b54e294be5 · outbound

This paper cites Improving Image Generation with Better Captions.

Testing chatbots on the creation of encoders for audio conditioned image generation Improving Image Generation with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.661088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.661088Z digest=sha256:31c2b2ed68627bd52656c528d46553aaf0a644fe6734e811d0adb336dfc6b2a1

Observation 20547dc6-b008-4d2c-947b-a66803d7055c · outbound

This paper cites RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model.

Testing chatbots on the creation of encoders for audio conditioned image generation RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.666579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.666579Z digest=sha256:b55a92bfba7c8b04e6a030cac51930b8f2812c1a17ba8a494da52dc159465733

Observation fcc6e1ec-fb5d-4440-9ba7-70abb84a231e · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.674010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.674010Z digest=sha256:94558acdc15dbb180ec571f43d5cb6b80b9500d4921871486a82b0e9fb580513

Observation ab37063a-de8f-43b6-bfe1-d16d0dc28503 · outbound

This paper cites A contemporary review on chatbots, AI-powered virtual conversa- tional agents, ChatGPT: Applications, open challenges and future research directions.

Testing chatbots on the creation of encoders for audio conditioned image generation A contemporary review on chatbots, AI-powered virtual conversa- tional agents, ChatGPT: Applications, open challenges and future research directions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.680984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.680984Z digest=sha256:8786f825847b509725e3946ecdf697a5ab3ed7b08c1f0a88b661a3ec3237dc31

Observation 82dcdb46-bde2-4d42-b8ce-53cdd125d978 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.690305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.690305Z digest=sha256:b5fdbb60f5db1ae5f626a0d58e282886ed36144cf414d561e825b1a2cef7db8e

Observation bc3c8152-5272-4f11-a4e6-d8dcb2bad7cb · outbound

This paper cites Veo, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Veo, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.697479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.697479Z digest=sha256:2330cb2755c20b5e963740e1601691146ef75ad8d58391854c1be8b8f99d8067

Observation 309f9b93-2acc-4a5f-a337-cf7a05d6abc1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Testing chatbots on the creation of encoders for audio conditioned image generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.706108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.706108Z digest=sha256:d621233d87d40fb47aee536c7caab81fa098b07b7d6e560fe7731c007fd110aa

Observation 8de0c676-9fd4-422e-ac60-8fa77d590a9b · outbound

This paper cites A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021.

Testing chatbots on the creation of encoders for audio conditioned image generation A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.712700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.712700Z digest=sha256:0480a4411f1c65ae5df1f85bfc2374402fd10b49d7d3d722d6fbaf6ae64c6dc9

Observation 2ef99852-0c8f-45c3-a4b6-81286b25bc31 · outbound

This paper cites Jukebox: A Generative Model for Music.

Testing chatbots on the creation of encoders for audio conditioned image generation Jukebox: A Generative Model for Music

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.718047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.718047Z digest=sha256:40e0c9c8b6650e4be98792845ba540f92e58f2acdf85db5a6e4012181821fed4

Observation fc4d511e-d4f1-4943-bbe4-4729dfd730d9 · outbound

This paper cites The Llama 3 Herd of Models.

Testing chatbots on the creation of encoders for audio conditioned image generation The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.723712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.723712Z digest=sha256:3b93685b5aba6503db02fee67e730d93018f56b0109a0e9c84678246329a3ff5

Observation 3da38730-450b-4774-b7e8-e809ea7f9be0 · outbound

This paper cites Grok, Gemini, ChatGPT and DeepSeek: Comparison and Applications in Conversational Artificial Intelligence.

Testing chatbots on the creation of encoders for audio conditioned image generation Grok, Gemini, ChatGPT and DeepSeek: Comparison and Applications in Conversational Artificial Intelligence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.730575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.730575Z digest=sha256:3efb40e61d31599bcfa261c380d821cabb98cb708bb006ef7a42fa6ea0ea5fe6

Observation 9f2c13e2-b696-45c9-82cd-af72e946ec29 · outbound

This paper cites Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022.

Testing chatbots on the creation of encoders for audio conditioned image generation Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.738911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.738911Z digest=sha256:8b7b9e1d220cb50f66efff7d94c53a8da198409a6ff2e59d7bb4c1885e39a8d7

Observation fa9b58c6-6e37-4aa8-ba87-ba78c7eb8c3d · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Testing chatbots on the creation of encoders for audio conditioned image generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.744606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.744606Z digest=sha256:e4744a2ac48548a93e6a8e8481724e6437221b15b75fbe287098a25a19aaffcd

Observation ee78b31a-84de-4bf3-bd58-33552e137ab4 · outbound

This paper cites Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching.IEEE Transactions on Multimedia, 26:3884–3896, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching.IEEE Transactions on Multimedia, 26:3884–3896, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.750603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.750603Z digest=sha256:fb2b6db2b71ee9d16addc83bd63aef7d82247674e25f8451b1d8387f33ec1d9f

Observation 047ae5d0-4bd9-4baf-8a9d-3790b8fe4c05 · outbound

This paper cites Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.758106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.758106Z digest=sha256:72b82db4a5304411b8a3ab34bf286299346ca2055aebc86a08ee6960919a5c4e

Observation ad942ac4-e177-4d3f-921c-3aa1448bd327 · outbound

This paper cites Creativity and Machine Learning: A Survey.

Testing chatbots on the creation of encoders for audio conditioned image generation Creativity and Machine Learning: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.764195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.764195Z digest=sha256:94afbdf5ecffffb94566cec7bd522bf44af4e2d8306661e40a14a30cf8c36f38

Observation bc1a4d8b-6bd8-46f1-9916-ab35fab94150 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Testing chatbots on the creation of encoders for audio conditioned image generation The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.772408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.772408Z digest=sha256:33a3f31ab57a1c43e3fff56a34a254ca9d10584740f33f62de548941b9a9d588

Observation 09e8c540-3a19-48e0-8d72-22511bc0c912 · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

Testing chatbots on the creation of encoders for audio conditioned image generation ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.782500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.782500Z digest=sha256:6aaf3d74fc2751d54804508491bab0bffc939285af3eab485739083d686fa74f

Observation 119adbf1-bf88-4802-853f-34a2ac03bf28 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Testing chatbots on the creation of encoders for audio conditioned image generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.792335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.792335Z digest=sha256:89889ed9f60ae5a3251dc917cf850a6008c45c2f9809cc97b2defee1b417875e

Observation 406eddbf-0f26-4e61-a2d9-0a886814416a · outbound

This paper cites AudioCLIP: Extending CLIP to Image, Text and Audio.

Testing chatbots on the creation of encoders for audio conditioned image generation AudioCLIP: Extending CLIP to Image, Text and Audio

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.798563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.798563Z digest=sha256:2a98918fe402a25dbad4646ae90e51d2044aba1741aa890641cc0524a991153c

Observation ffbee515-f413-429f-9cdd-85b23af3b599 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Testing chatbots on the creation of encoders for audio conditioned image generation Deep Residual Learning for Image Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.806461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.806461Z digest=sha256:c532a63222540ca007e8f978e36d4d8e12ece031f86002209f4db7ce9eb01924

Observation d0ee86c4-a0bb-4610-862c-89a99d38fcfb · outbound

This paper cites Ringle, and Rudolf R.

Testing chatbots on the creation of encoders for audio conditioned image generation Ringle, and Rudolf R

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.812512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.812512Z digest=sha256:37cf848413134124757fff51255904880c5aab74653f7140fc84b823662a83a6

Observation 4d61857a-ddd6-4619-a87e-a0626dc1af99 · outbound

This paper cites Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model.

Testing chatbots on the creation of encoders for audio conditioned image generation Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.818748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.818748Z digest=sha256:bc43af7ce386aff6ae622662d14abffba86bf25a1c848c789e20cc8b2eeac35a

Observation cc085099-2c40-44a1-8e42-403256fd10a6 · outbound

This paper cites Make-an-audio: text-to-audio generation with prompt-enhanced diffusion models.

Testing chatbots on the creation of encoders for audio conditioned image generation Make-an-audio: text-to-audio generation with prompt-enhanced diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.825041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.825041Z digest=sha256:34e45d9f5381b065e9a4337c67b01d7f3be1af0c7fed39e27594454f3c7dce66

Observation df0d6c3f-633e-4981-8fb9-65f2ecdd678f · outbound

This paper cites NLIP: Noise-Robust Language-Image Pre-training.

Testing chatbots on the creation of encoders for audio conditioned image generation NLIP: Noise-Robust Language-Image Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.832038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.832038Z digest=sha256:ab32689abf465fa82769a0ab95c1c078ac7d5f99922d1b7028c4b884fd22decc

Observation 98e98fc0-7756-49d1-a249-36c0ef3a0af8 · outbound

This paper cites Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications.

Testing chatbots on the creation of encoders for audio conditioned image generation Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.837379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.837379Z digest=sha256:171ec6ccea8439b3f427fc12a3e7000cd4b15dc6ae4ac9f42640b17f08e45462

Observation 46dae24a-4596-4be1-a69b-016b472057f1 · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.844739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.844739Z digest=sha256:04ad6aa2856f45a48ecf706f01df1067f0f62cca1e52060a2b9e3fc3be734c7b

Observation 0a4d095f-b77c-4ee5-a5aa-9c072ebfc696 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation A Survey on Large Language Models for Code Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.851346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.851346Z digest=sha256:decccc2f26ddbb9ebbe23f9cd9787274e15818f418a6f97b7400906e975298c2

Observation e2250698-490c-43ca-9507-c52ddad3006b · outbound

This paper cites TimbreCLIP: Connecting Timbre to Text and Images.

Testing chatbots on the creation of encoders for audio conditioned image generation TimbreCLIP: Connecting Timbre to Text and Images

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:28.343551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:26.858805Z digest=sha256:048df05a5d7902c4983dbaa48acf1731ba184bfea8d5c9b70e1deefc71fa7da5

Observation 6c058880-0ef4-4c78-af9a-c283e4103de4 · outbound

This paper cites Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning.

Testing chatbots on the creation of encoders for audio conditioned image generation Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.864893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.864893Z digest=sha256:c692a0a4d204d88920d28f5238a956ddc02382195639c7b23966b147f3c94f39

Observation 69b1f822-cf4d-45df-ac9c-c41495f623ae · outbound

This paper cites Gemini 2.5: Our most intelligent AI model, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation Gemini 2.5: Our most intelligent AI model, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.871127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.871127Z digest=sha256:abe4cae1179b959f22ea5fbdf6c063f186c26e88be6910a6688c05a74e029311

Observation 4d7c3d45-3127-4197-aa5a-f7c0b94c6f3a · outbound

This paper cites an unresolved cited work.

Testing chatbots on the creation of encoders for audio conditioned image generation Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.877668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.877668Z digest=sha256:87fc9a36387d990168d990dc9db172d5f38bfc6e6d878de7c4c31f65825065c5

Observation d253f024-d21b-4968-b2e5-759b970f9219 · outbound

This paper cites Kingma and Max Welling.

Testing chatbots on the creation of encoders for audio conditioned image generation Kingma and Max Welling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.887383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.887383Z digest=sha256:1cdd676ce460f48ebe70fbd197f33bdc694ccf9447eb2d7a130cc1a4b43c9c1f

Observation e4a5ec78-d9dc-4d66-ba76-3f0446d240da · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Testing chatbots on the creation of encoders for audio conditioned image generation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.892235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.892235Z digest=sha256:99ac3e1849de5ae5c215509f2f43481e46401399f507ceefb3c6a15567aa7f67

Observation 5d376a71-001d-4b19-8ad8-e663da3d7499 · outbound

This paper cites Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Proceedings of the ACM on Software Engineering, 1(FSE):2261–2284, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Proceedings of the ACM on Software Engineering, 1(FSE):2261–2284, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.898143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.898143Z digest=sha256:255bc516fcb82c9f91311a69434a7025d7e911b688e394403724f86e2798adaf

Observation b2ed3cc5-127e-42e8-a0e2-220132449bde · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation AudioGen: Textually Guided Audio Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.909866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.909866Z digest=sha256:ac7b33d17d040810811c5d24cf1d0f7ed40b6e848b2c87016917fe4d12359ace

Observation 435325f9-0ec9-4167-a8c6-b8d4486fe4ad · outbound

This paper cites BindDiffusion: One Diffusion Model to Bind Them All, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation BindDiffusion: One Diffusion Model to Bind Them All, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.916774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.916774Z digest=sha256:e8d253ccf36dca88bc1bc1d8372e9104267fdc381296d9be2a1de37ce12cc452

Observation 976fec45-ace3-47a1-b262-0ec5f103148f · outbound

This paper cites FLUX, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation FLUX, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.923550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.923550Z digest=sha256:218e29ea027d010fdd25777ed59018e79c3ae0ace4ab6c04a703af5e0b01ca8c

Observation 7f772987-5f54-46cf-9866-ec6efe4b5788 · outbound

This paper cites Effectively obtaining acoustic, visual and textual data from videos.

Testing chatbots on the creation of encoders for audio conditioned image generation Effectively obtaining acoustic, visual and textual data from videos

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:25:28.242608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:26.928879Z digest=sha256:ac98b4cf8c592201f0c3345af381ed5443ab335b94953b32e5c8b396a6991ff5

Observation e5a97f0b-c847-4cce-9647-fee7cfe9f742 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.940451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.940451Z digest=sha256:8a63ffa40f5e5c0233a97b86fb98678aa44625194e2491e333233e8b52a82c1d

Observation 41c23dbc-f4c8-4775-a078-abb7d87613f1 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.951332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.951332Z digest=sha256:450a95bfec5dc02611a82a26a00c913c1aaf8b42499807163de1737477dda122

Observation d24a96c5-b7d6-4c71-ab90-f8bde6247612 · outbound

This paper cites Word-Level Explanations for Analyzing Bias in Text-to-Image Models.

Testing chatbots on the creation of encoders for audio conditioned image generation Word-Level Explanations for Analyzing Bias in Text-to-Image Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:28.146878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:26.959197Z digest=sha256:1eaf1686a57a7a2e17bbfb0699d6cc0310266df156db5aa2b5f1f314244f36a7

Observation 1294e260-46d8-432a-bb08-7ca16290258f · outbound

This paper cites Plumbley.

Testing chatbots on the creation of encoders for audio conditioned image generation Plumbley

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.966726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.966726Z digest=sha256:840d5be5fd0fafd28ac90da127de897a0b98d983ad958ccd97e53685727e507e

Observation 21019a34-676b-45d6-8db4-d5e6457671e8 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Testing chatbots on the creation of encoders for audio conditioned image generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.973235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.973235Z digest=sha256:8e3cc4365557212c054b98b8ae8268ea3d816d8fb3c76f1b9f2ca0b6c08c0f57

Observation 34c387b4-362c-40dd-be69-e07d440041c2 · outbound

This paper cites Michaud, Max Tegmark, and Mike Williams.

Testing chatbots on the creation of encoders for audio conditioned image generation Michaud, Max Tegmark, and Mike Williams

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.980025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.980025Z digest=sha256:edfa0876eb8a2d600de3e4360625ebeffc00960b20477464af0c0b597433a4c6

Observation 9724b24e-3ecf-4c90-bd92-f7a47ee7a40e · outbound

This paper cites BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning.

Testing chatbots on the creation of encoders for audio conditioned image generation BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.985740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.985740Z digest=sha256:323ade50368befa507322d5b014ff6dea0522d49a3e238e11aa8591aa97eef38

Observation 2da851cd-2682-42f5-88a5-89d4d333f690 · outbound

This paper cites Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering.

Testing chatbots on the creation of encoders for audio conditioned image generation Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.993264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.993264Z digest=sha256:b4ae94cfc3bdc4d2247ec133afa48eb7e5c4a601d6f296a9e2a6030120380cee

Observation 8d17c90f-f472-4576-a9dc-85ca822f9bd4 · outbound

This paper cites Stable Diffusion Akashic Records, 2023.

Testing chatbots on the creation of encoders for audio conditioned image generation Stable Diffusion Akashic Records, 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.003073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.003073Z digest=sha256:4ceb768cec72bc5c6c7b76f5b2e896cab828db6508aff3e797f6b7830d6bb9c7

Observation 3a15c995-acde-409f-ae15-874ae5110b08 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.IEEE Transactions on Artificial Intelli- gence, pages 1–18, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.IEEE Transactions on Artificial Intelli- gence, pages 1–18, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.008405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.008405Z digest=sha256:ed2ded28f280463731b12107869e1c34b06498e3272ffca9c4555564d0b4d3b7

Observation 9df21f03-dd77-4c91-b5d1-001243792a5a · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation Mustango: Toward Controllable Text-to-Music Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.015737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.015737Z digest=sha256:f8729218062baf5de38729d8d092450391f508a6d31eca34443237f23e99cbb3

Observation 4ff0f8b1-b540-4fad-a911-2893145cd957 · outbound

This paper cites Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis.

Testing chatbots on the creation of encoders for audio conditioned image generation Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.028569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.028569Z digest=sha256:c5277c47290325f8aa35d866c01e1a294f1ecba5b80d1c1fa9b095e8d01b6e92

Observation 1e474672-964f-487d-9555-73187522931d · outbound

This paper cites DALL·E 3 System Card, 2023.

Testing chatbots on the creation of encoders for audio conditioned image generation DALL·E 3 System Card, 2023

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.040043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.040043Z digest=sha256:d3a425323cf0742d034d6e3a52775ac2bf68adc944521bc4c3f2dc671ad16b41

Observation 68eebaa9-7df4-4b8f-932d-d2b1201983c5 · outbound

This paper cites Video generation models as world simulators, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Video generation models as world simulators, 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.046611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.046611Z digest=sha256:a805b5d0561aacd4152f569818a2ebacc5558a22df7a8216658a1f5add25de70

Observation d34ea8b4-de07-4fa4-9651-94139d0d5480 · outbound

This paper cites OpenAI o3-mini, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation OpenAI o3-mini, 2025

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:30.124640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.055357Z digest=sha256:f457b529a89535926234b6662dd575c622ada191e6efaae9fac6f374b8076dff

Observation fc5b0e53-a0ec-41a7-bd91-a40affe9a61e · outbound

This paper cites Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022.

Testing chatbots on the creation of encoders for audio conditioned image generation Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.944903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.064624Z digest=sha256:9b7e9241fe020494a809c62f268304dab00a794cbe071a9c48acd41551fd16a2

Observation d2ad6525-3af7-4fbf-bffa-451c752a362e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Testing chatbots on the creation of encoders for audio conditioned image generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.072098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.072098Z digest=sha256:da05caa285355fe68bb0b05ee5be6321b1106f3fc72b7ebbd112de95e7ab3969

Observation 9873f815-964c-4990-a342-ec5cc674ded1 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Testing chatbots on the creation of encoders for audio conditioned image generation Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.802763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.079736Z digest=sha256:0951c69aa64f3b001a44f29be40333d5b6bfc366757e961746a0b2fd77b0ee87

Observation ae62acc3-d4f7-43ab-8426-d82bfc8b4266 · outbound

This paper cites MirrorGAN: Learning Text-To-Image Generation by Redescription.

Testing chatbots on the creation of encoders for audio conditioned image generation MirrorGAN: Learning Text-To-Image Generation by Redescription

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.668998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.087641Z digest=sha256:5817db897fc4e892327c41088a197ff15600e1262551081d2bedbbd2f750d12b

Observation 94515e86-5902-4e92-a229-dc1711ab73c1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Testing chatbots on the creation of encoders for audio conditioned image generation Learning Transferable Visual Models From Natural Language Supervision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.100683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.100683Z digest=sha256:286bfc259069ac612f0593d3aca8d753743d390d952b1555df212168c68601ad

Observation 560ca1e8-28fb-46bf-b316-0d2626334408 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Testing chatbots on the creation of encoders for audio conditioned image generation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.604367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.108332Z digest=sha256:f6153ba3d24f80e6e97441f078470aba434876ec1d2f010a693cf80754286cc3

Observation e4bfc21b-1c53-4578-98e0-b9f4ab7a1ca6 · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation Zero-Shot Text-to-Image Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.114736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.114736Z digest=sha256:ababaf7ae75cb3a14ebdb926cd159b2c00253b0d7f632707ea593d7cb9e1f7b3

Observation 189ec3ff-a078-41c5-bb52-ac6561aa6189 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Testing chatbots on the creation of encoders for audio conditioned image generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.120753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.120753Z digest=sha256:f4051f6cafe8d319ccc36f045ecc455d7c65ea62e77093d66f25bb372ff235a0

Observation cd2fb3e2-711a-4f2f-83cc-c9d162c93154 · outbound

This paper cites Stable Diffusion, 2021.

Testing chatbots on the creation of encoders for audio conditioned image generation Stable Diffusion, 2021

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.560658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.128148Z digest=sha256:aeeef8958f8a87239d03a934b10422722ab796c6cc43dd2a517d4beae372e701

Observation b390fbb8-d714-4c5f-b665-0ee1f85481c4 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Testing chatbots on the creation of encoders for audio conditioned image generation High-Resolution Image Synthesis with Latent Diffusion Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.135223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.135223Z digest=sha256:bda0b531a1524b380ea63fd8ab6723782bb6f40d6e6cfcd7c235167eb8900c89

Observation 876a173b-c180-4813-9a4a-8f98572b99b7 · outbound

This paper cites Stable Diffusion v1-5 Model Card, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Stable Diffusion v1-5 Model Card, 2024

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.532726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.141626Z digest=sha256:9ef8b7be34b9b3c00fbdbedf0f96cfeb6ecab7024908790226a443c8a55b623a

Observation 427732d4-70e3-434b-8784-53b9af8ba850 · outbound

This paper cites U-Net: Convolutional Net- works for Biomedical Image Segmentation.

Testing chatbots on the creation of encoders for audio conditioned image generation U-Net: Convolutional Net- works for Biomedical Image Segmentation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.498665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.146424Z digest=sha256:cf62996ef1532c9661f9c3e961379686b0e2c93b3ce083ce2094930f18cc3a42

Observation f5a8a677-2f4f-4297-aa56-0c07c574c3ff · outbound

This paper cites Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.468796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.151806Z digest=sha256:e1c58c892a99bab2d21b6593c946161e8fe2864ae0df54dfb4ced941354e03fe

Observation 580452de-ae29-454a-a195-0e68d6aabc2e · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Testing chatbots on the creation of encoders for audio conditioned image generation Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.447853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.157197Z digest=sha256:002bf3e6559d32f140e9c4a5de8e5f48e3b011e2ba7e79b6e7ce330bac663e97

Observation 448a132b-e5c9-440d-8066-d6aa0d5b81b5 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

Testing chatbots on the creation of encoders for audio conditioned image generation A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.165116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.165116Z digest=sha256:b5a6ac7798f37a410c87131cd149e42b097881ea490ee64727a2c5c0a4fbf077

Observation 039102f3-8421-4ccd-b322-a7b9d1b4189a · outbound

This paper cites Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey.

Testing chatbots on the creation of encoders for audio conditioned image generation Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:27.892925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.175663Z digest=sha256:ab91f3417f456f1db1065f4a070c408dc4e0cdd948fe9a3943c7e955e6d0ccd9

Observation 4df9dea7-e41e-41d1-82cf-cc5dbb4adae1 · outbound

This paper cites What is noise?Geophysics, 63(4):1122–1124, 1998.

Testing chatbots on the creation of encoders for audio conditioned image generation What is noise?Geophysics, 63(4):1122–1124, 1998

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.427153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.182122Z digest=sha256:f2032f1f74c61fe2539867e41809abaf335f7a203543ab3a803c3bcc53796773

Observation 9e8efcb1-4fd6-484c-a962-df22d26cfcad · outbound

This paper cites Large pre-trained language models contain human- like biases of what is right and wrong to do.Nature Machine Intelligence, 4:258–268, 2022.

Testing chatbots on the creation of encoders for audio conditioned image generation Large pre-trained language models contain human- like biases of what is right and wrong to do.Nature Machine Intelligence, 4:258–268, 2022

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.406716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.188129Z digest=sha256:e4a0321dcadb9c2867e308c9aca4369670b6b952275647a94d4e6606c925fea2

Observation ca7199ee-0ff4-4c83-836d-9c28b6654a4b · outbound

This paper cites A comprehensive review of large language models: issues and solutions in learning environments.Discover Sustainability, 6, 2025.

Testing chatbots on the creation of encoders for audio conditioned image generation A comprehensive review of large language models: issues and solutions in learning environments.Discover Sustainability, 6, 2025

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.386873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.193122Z digest=sha256:d5d7281b0641364fd9d4ab0726b6e626bd32d1065dbafb17b03431562c6baec5

Observation 09179d60-4806-4adb-a7d8-ec64697aef52 · outbound

This paper cites I Hear Your True Colors: Image Guided Audio Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation I Hear Your True Colors: Image Guided Audio Generation

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:27.862179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.198569Z digest=sha256:184b3fecbbfebfa48f00b5739ffa72f64bc496c47dc6c8abf341d7c6ce865388

Observation a0fc0fbc-155f-4321-8f35-c0266598a0db · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Testing chatbots on the creation of encoders for audio conditioned image generation A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:27.823348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.206894Z digest=sha256:76ec574979b2ce0dd1004d87c4ed6509ac9de7aad18fd29fdff11c95cb3c9aa5

Observation d7b109c7-918a-4162-b654-d7f697ebdd39 · outbound

This paper cites Audio-to-Visual Cross-Modal Generation of Birds.IEEE Access, 11:27719–27729, 2023.

Testing chatbots on the creation of encoders for audio conditioned image generation Audio-to-Visual Cross-Modal Generation of Birds.IEEE Access, 11:27719–27729, 2023

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.353725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.214946Z digest=sha256:9eee4ce9b4d94a4b0620e7f335675603585ec6679336853c305f7947bf8022d8

Observation 9c67e0b6-73e8-4529-9a42-bf3c2d3cf6b2 · outbound

This paper cites Outpainting Images and Videos using GANs.International Journal of Computer Trends and Tech- nology, 68(5):24–29, 2020.

Testing chatbots on the creation of encoders for audio conditioned image generation Outpainting Images and Videos using GANs.International Journal of Computer Trends and Tech- nology, 68(5):24–29, 2020

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.323392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.221269Z digest=sha256:3b1cb13500b7e81a032e9fa0937fabca01b685a7b6062446643154155476cce9

Observation 8ee53418-5151-4228-92ba-0de62cba72a1 · outbound

This paper cites Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emo- tion Recognition.

Testing chatbots on the creation of encoders for audio conditioned image generation Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emo- tion Recognition

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.294833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.227831Z digest=sha256:9fc45fed99cd06bd643d0c42f49d57f77bacbb05da60255ed1f633e80ca44152

Observation e529c811-477d-4dc6-8775-99272ed96449 · outbound

This paper cites A survey of multimodal deep generative models.

Testing chatbots on the creation of encoders for audio conditioned image generation A survey of multimodal deep generative models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.262507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.234737Z digest=sha256:65ce502007c6b99e51113efdc6ecc09bdddd1f2d460d6909b1a6fb7479b1c24d

Observation e31a462e-99d7-463c-b0e8-11cae17856c8 · outbound

This paper cites CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation.

Testing chatbots on the creation of encoders for audio conditioned image generation CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.241054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.241054Z digest=sha256:6a5411688869a90c8def6d43bc0d09c0fd741872ca85e5d9e0c8b134c4986d40

Observation 1abe22b2-cfb2-4ecf-b81c-80a6af345805 · outbound

This paper cites Any- to-any generation via composable diffusion.

Testing chatbots on the creation of encoders for audio conditioned image generation Any- to-any generation via composable diffusion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.242371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.255462Z digest=sha256:417bfd113c2ec85bbaa89424e340a459481b751139649abd973a275b09f1b908

Observation 13a05cb0-55f7-4dbb-991c-ba76d164d6bf · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Movie Gen: A Cast of Media Foundation Models, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.222820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.262025Z digest=sha256:ab4e9bbba8b607aabfea72d28c425e0e869fa8adc9af4f9be29d59ebd14be7b3

Observation 58e740db-a03e-4fc7-a7ac-05156baf126a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Testing chatbots on the creation of encoders for audio conditioned image generation LLaMA: Open and Efficient Foundation Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.268025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.268025Z digest=sha256:a17c8a740e3d5a6280343347556a0d362ef9d7e7603f25d94dcb1a44b336ae86

Observation f6aec77f-4114-4f64-a806-6f059afab01a · outbound

This paper cites Structural Equation Modeling in Information Systems Research Using Partial Least Squares.Journal of Information Technology Theory and Application, 11(2):5–40, 2010.

Testing chatbots on the creation of encoders for audio conditioned image generation Structural Equation Modeling in Information Systems Research Using Partial Least Squares.Journal of Information Technology Theory and Application, 11(2):5–40, 2010

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.200270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.274583Z digest=sha256:9df7a74991bccabbc6365254692097c6ab345c92028db9c571cefdb328a210b6

Observation ef90ac3f-497b-487e-8f25-76419f114f45 · outbound

This paper cites Fugatto 1 - Foundational Genera- tive Audio Transformer Opus 1, 2024.

Testing chatbots on the creation of encoders for audio conditioned image generation Fugatto 1 - Foundational Genera- tive Audio Transformer Opus 1, 2024

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.174044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.281670Z digest=sha256:10f45684d0d8dd0443d79598a13588752c7ded2423d3c1217e43a39955ef2721

Observation 1310c348-71c0-43bc-81d7-ac37937b96b5 · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Testing chatbots on the creation of encoders for audio conditioned image generation Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:25:29.152219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-04T21:25:27.288300Z digest=sha256:6cca07dc1db2372940db4ce533dd83d2a09903b1839c1e87a39c144c0ec9cb11

Pith citing papers

No inbound Pith citation observations are available.