Pith. sign in

Paper Citation Record · LEDGER

LLMs can see and hear without any training

As of 10 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 4 inbound Pith citation observations for arXiv:2501.18096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18096 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:08.276886Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:41:39.719320Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T17:53:11.781235Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07461521-f5d5-4dbc-ba99-b9eca04c3b9d · outbound

This paper cites write newline.

LLMs can see and hear without any training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:07.963230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:07.963230Z digest=sha256:045ebcaebb4ad3a9a567600fabd82a53c85111aab11cbef9a6db5974aae7d802

Observation f40734fc-2f39-4238-b0e9-b4ac7c89a2c0 · outbound

This paper cites Pixtral 12B.

LLMs can see and hear without any training Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:07.970023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:07.970023Z digest=sha256:9852f195b471154b19a50734a837847c0d52ce1f93ed92164de5e42b1f51dfb8

Observation 4f5ce520-c345-42cb-8b39-daedf969618c · outbound

This paper cites SPICE : S emantic propositional image caption evaluation.

LLMs can see and hear without any training SPICE : S emantic propositional image caption evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.269755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:07.975582Z digest=sha256:4a4993532fdc694d27f2e97b21c3323dd033ca127aca58a2305bcd726d504b10

Observation 259fe07d-7b7f-43c5-89ef-e5b091e28ad5 · outbound

This paper cites and Lavie, A.

LLMs can see and hear without any training and Lavie, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.253716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:07.980903Z digest=sha256:02e700bb17a7ea85a647c3b4fd80f980a32cc71d02ded3780be46c4dccaeafc3

Observation f7f04a0f-559b-4c16-a6f7-d1f502c76582 · outbound

This paper cites Improving image generation with better captions.

LLMs can see and hear without any training Improving image generation with better captions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.237162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:07.986739Z digest=sha256:5a2a3104b964b16210d7fd1ac8407fe3e994378e9448ef04ba05cc1bb0f32863

Observation cb5f418e-b65d-4717-9e60-9717e8746d21 · outbound

This paper cites W., Fidler, S., and Kreis, K.

LLMs can see and hear without any training W., Fidler, S., and Kreis, K

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:07.991979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:07.991979Z digest=sha256:bf1a36ef436307e8f6cdf5a489bae1973b64704715f24f3776c95d1ed26beb56

Observation 8d80eda5-09d1-4c14-a4db-130921020026 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

LLMs can see and hear without any training Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:07.998074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:07.998074Z digest=sha256:db105bc3986edac80ce4cee46fd8599cb6e7e4701b5217c30d65031c8fe4a0d7

Observation 67571005-a8df-4362-aab6-cacfa53cad31 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

LLMs can see and hear without any training Imagenet: A large-scale hierarchical image database

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.004275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.004275Z digest=sha256:f4aa7db78ab60c71db7abeba52a098a1ec0bc549c85296cd3631a300b5b172b1

Observation 9ec30906-4b7d-4993-9436-66a84a2a57e2 · outbound

This paper cites Clotho: An audio captioning dataset.

LLMs can see and hear without any training Clotho: An audio captioning dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.197196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.008906Z digest=sha256:f331a5adc6a03ece9807a46387db0c5095d9b8cf371df1d95bdc24c3b03b4d05

Observation 8c0cfefa-d255-443d-87bb-1d43bbba6d22 · outbound

This paper cites The Llama 3 Herd of Models.

LLMs can see and hear without any training The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.017455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.017455Z digest=sha256:006cc05b2051a7f517cf57aa10d5865325daa95cc58899e469775e8fa34d0ce6

Observation d19565a3-573a-4ece-93c4-787d72b96cfe · outbound

This paper cites M., Jain, A., Schmidt, L., Toshev, A., and Shankar, V.

LLMs can see and hear without any training M., Jain, A., Schmidt, L., Toshev, A., and Shankar, V

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.181512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.023267Z digest=sha256:8e09f28fb15b936aae644612d33f0ada45e4f01c63410d6e69acae42de5777d5

Observation 8dda7dcb-7a38-4118-9091-8a4fc12923ac · outbound

This paper cites Interpreting the Second-Order Effects of Neurons in CLIP.

LLMs can see and hear without any training Interpreting the Second-Order Effects of Neurons in CLIP

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.028501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.028501Z digest=sha256:d167b79ff7c5be92e2de879719561af8d77607642fcbe74772a0493658737c04

Observation 24436c06-49ae-487f-bbbd-1af9ed9e7331 · outbound

This paper cites A Neural Algorithm of Artistic Style.

LLMs can see and hear without any training A Neural Algorithm of Artistic Style

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.033691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.033691Z digest=sha256:800f9d0dc7c06e4f31805ee5a96f549fd5022d48c9a4d2e7c029be1bc33863ad

Observation ee6ca5fe-c59e-4946-ab2d-0d88b34aa8a9 · outbound

This paper cites On the content bias in fr \'e chet video distance.

LLMs can see and hear without any training On the content bias in fr \'e chet video distance

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.165970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.038837Z digest=sha256:c6e94a704643ba2d120b999d3f650e379ff5fbb05738317acbefd612c57b712c

Observation 54066178-ac0e-4576-b742-67253af18245 · outbound

This paper cites F., Ellis, D.

LLMs can see and hear without any training F., Ellis, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.150324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.043404Z digest=sha256:b839203598756ee6203b22bc6d0c407fa5454d5468c983da2237fc20b12868c9

Observation bfefb861-2462-44e1-91b4-ca65e854ed40 · outbound

This paper cites V., Joulin, A., and Misra, I.

LLMs can see and hear without any training V., Joulin, A., and Misra, I

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.132059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.047758Z digest=sha256:8c1035b2a4488d7d2037e4b562a4e6bad41caede25d3f8f27ceda2d17da12fe6

Observation da8027aa-beb1-4bf2-a24d-36f98151b87b · outbound

This paper cites S., Shah, A., Yin, X., Parikh, D., and Misra, I.

LLMs can see and hear without any training S., Shah, A., Yin, X., Parikh, D., and Misra, I

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.052450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.052450Z digest=sha256:f9625182f7451aa26bce3503cbc3d0131cbaef4a635111d14355ffcce2475c7f

Observation fdd064a1-cf92-4143-9fd5-2bf110ec5155 · outbound

This paper cites Mmg-ego4d: Multimodal generalization in egocentric action recognition.

LLMs can see and hear without any training Mmg-ego4d: Multimodal generalization in egocentric action recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.106134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.057400Z digest=sha256:4a1046ee46b118ffe99168305444b3b9f5b4d34d2a586cb0f07ff539f7e029f8

Observation 2aa4606c-ac40-4d58-b040-be7ec5246a39 · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

LLMs can see and hear without any training Audioclip: Extending clip to image, text and audio

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.091082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.062041Z digest=sha256:3f39a98b898b8a93ba08349939ee8eb770e6e86e32854a04624d4813eb0d0d36

Observation 0741e06d-d761-4090-95da-3600bbee57e9 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

LLMs can see and hear without any training Imagen Video: High Definition Video Generation with Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.066514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.066514Z digest=sha256:d63a391e640a58011b73ad296598d9251c83557fa2ea71341fa589cde8f9ec06

Observation 20abe8e5-1b44-4608-a80a-2f1ca44b0f92 · outbound

This paper cites Openclip, 2021.

LLMs can see and hear without any training Openclip, 2021

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.074929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.071638Z digest=sha256:2ebadef626ee3361b1b089b11cb1f79d5f76ab746762f84987aad7ec235bad9b

Observation 50a54a14-ec58-44f2-b3be-4933370cbdc2 · outbound

This paper cites Rethinking fid: Towards a better evaluation metric for image generation.

LLMs can see and hear without any training Rethinking fid: Towards a better evaluation metric for image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.059583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.076412Z digest=sha256:c54bba33f38fcb9a52120176bdf8b76104e13442068c95bf1d5c7fcc2e952810

Observation e327ff0a-71c0-455a-8b32-53de2061b349 · outbound

This paper cites Mistral 7B.

LLMs can see and hear without any training Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.081016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.081016Z digest=sha256:f67d4a41e17c37627ef8be1f5cf437c8937ed0ed4a33df1f85ed681f53e45a59

Observation 9e04e41e-d73b-49e9-a3d0-baeb2bf2f162 · outbound

This paper cites and Fei-Fei, L.

LLMs can see and hear without any training and Fei-Fei, L

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.044785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.085982Z digest=sha256:18aca33cd255b9aca328947ed1087ef1b5dd5af52a4595d2cd5808141614d931

Observation 45df32dc-b331-40fc-8f74-81b7e508733e · outbound

This paper cites What do we learn from inverting CLIP models?.

LLMs can see and hear without any training What do we learn from inverting CLIP models?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.090384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.090384Z digest=sha256:3556229683cf19ad2dcce4bdc51170b6588a43b0c97f17e4ec5aba6ebf799162

Observation 45345009-5da6-443e-8f01-318e7bb14201 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

LLMs can see and hear without any training Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.028909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.095404Z digest=sha256:f987c087c5b919e77911447c45099540c9ece9ee11a499ae120438174b386dfc

Observation 50f1f474-5d21-4fa5-9df9-987b465c0df9 · outbound

This paper cites S., Reid, M., Matsuo, Y., and Iwasawa, Y.

LLMs can see and hear without any training S., Reid, M., Matsuo, Y., and Iwasawa, Y

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:09.012021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.100070Z digest=sha256:8ad43d8a9e3e4527c25fc4faa7b5953e6f7fcd9eb8ddb0d628bb3b47a26e6b24

Observation 688e3352-7798-4a11-b2f5-d99be38bd7b3 · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities.

LLMs can see and hear without any training Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.997000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.105072Z digest=sha256:2603161f05655eb998b8aa7b06c3444d5dc3041b46070d3a2f0d636faa27bc02

Observation 36a6ddfb-a057-4d2b-b35b-cbcf5ab1102b · outbound

This paper cites an unresolved cited work.

LLMs can see and hear without any training Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:46:08.977632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.109523Z digest=sha256:24445724b4434b030f3c59df2cd9c16567b0a8a912f4f7c3a15455aa38d60d37

Observation 52c1ea4a-947f-4dd1-ad4c-5160baaccb34 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLMs can see and hear without any training Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.961288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.114123Z digest=sha256:47a67c5b15fb46d63d7499b48d2f3d0e7531e6df8622fc4803ed75daa83095b7

Observation ac1e12be-fce0-4996-a55a-dcd92156cbe6 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

LLMs can see and hear without any training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.118465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.118465Z digest=sha256:aada44bb2470de3d378061b9b19d0726e92b6dfedc1a3676fd839f4614679b2f

Observation aaddedd7-5f35-4e5b-9764-503fad42fcb9 · outbound

This paper cites DeCap : D ecoding CLIP latents for zero-shot captioning via text-only training.

LLMs can see and hear without any training DeCap : D ecoding CLIP latents for zero-shot captioning via text-only training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.945732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.123749Z digest=sha256:3f1ef2774f62526b24d45c7fb0ef94e0900eef3dd4f4e809eaf2e3116694087d

Observation 14a4ec48-b4ce-41cd-b469-16522d447a2b · outbound

This paper cites an unresolved cited work.

LLMs can see and hear without any training Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.128461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.128461Z digest=sha256:1287991058ac5c3ff024d5429b167e458497e1d77b18ab0e33028c3f6b4b717c

Observation 2cbe38f1-f7c2-43ad-ad6f-6b85627af581 · outbound

This paper cites Flow Matching for Generative Modeling.

LLMs can see and hear without any training Flow Matching for Generative Modeling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.133554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.133554Z digest=sha256:50c85d441dc938e815cdcd9aa29160ef1367fb592e805f5eb40bc2dedc82f5bc

Observation 34264a0c-d475-4189-b526-fced26b8aac1 · outbound

This paper cites Improving Text-to-Image Consistency via Automatic Prompt Optimization.

LLMs can see and hear without any training Improving Text-to-Image Consistency via Automatic Prompt Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.138982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.138982Z digest=sha256:9f8c15c54bd4a66e7b929cd4544cce5660754f77e2de5cefea7ed8277767c805

Observation 593a6f57-3e95-44bb-a976-34d8cea1126d · outbound

This paper cites Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities.

LLMs can see and hear without any training Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.143919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.143919Z digest=sha256:bb4bb39dd5e0f05fa44b3f4823ef531fb92f65bc7c021317706e96747b695ae0

Observation 1033a21d-83e2-4edf-aeba-44da51e4bdd5 · outbound

This paper cites How T o100 M : L earning a T ext- V ideo E mbedding by W atching H undred M illion N arrated V ideo C lips.

LLMs can see and hear without any training How T o100 M : L earning a T ext- V ideo E mbedding by W atching H undred M illion N arrated V ideo C lips

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.917534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.148791Z digest=sha256:0a336adb047a6d1d178a1f320816e93d63d2c266b1137197f99d6fe5710f2518

Observation aa943601-8824-46bc-8f50-368d8e29c6bf · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

LLMs can see and hear without any training Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.153195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.153195Z digest=sha256:db5588469057d187348010485163e735e5d4b6f829324ee083cfdce9c3a87e96

Observation 396a3532-bf3b-4daf-a936-b2a2bd6c225c · outbound

This paper cites H., Seybold, B., Hauth, A., Manen, S., Sun, C., and Schmid, C.

LLMs can see and hear without any training H., Seybold, B., Hauth, A., Manen, S., Sun, C., and Schmid, C

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.902188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.158096Z digest=sha256:1d855c377dec47fca21e6a9c5cc35cfd2381c706a8339cc9381e002d31ec10b6

Observation 5dddf899-95dd-414b-b6d0-038afb215443 · outbound

This paper cites an unresolved cited work.

LLMs can see and hear without any training Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:46:08.885727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.162744Z digest=sha256:331d3bd73a6b77ea3ed31b881af1880a3603ab0ff2ea38cb7e8eb9ff45becd8d

Observation 10d47875-b6e7-4f60-83ca-41e7e4e945bf · outbound

This paper cites Introducing openai o1-preview.

LLMs can see and hear without any training Introducing openai o1-preview

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.867355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.167420Z digest=sha256:43c8acf60d583c7626bd91ddb3f5459b5f039636bf8447fc0777e93c0d45de50

Observation b7918d94-7097-4632-a4d4-a19cd2c21e94 · outbound

This paper cites BLEU : A method for automatic evaluation of machine translation.

LLMs can see and hear without any training BLEU : A method for automatic evaluation of machine translation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.849424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.172819Z digest=sha256:4d744ee2bf8a7c862edec5fd52116a45929fd75b9cb8b1025ba453bec85bdf25

Observation c4876983-387a-4576-a046-bea434a1ba8d · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

LLMs can see and hear without any training Movie Gen: A Cast of Media Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.177469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.177469Z digest=sha256:4e4ec8442274fd075005449002c6cf80e943109bda31e952b447797f16621c4c

Observation d52ba24e-5c4a-4635-ac20-4c094d5ec01d · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I.

LLMs can see and hear without any training W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.182174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.182174Z digest=sha256:e3bc167b735538da835294cab79f01138a7af77c0717a7769d24171b6004068a

Observation 9baacafb-a287-4ef2-b54d-1cbb561e5bed · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

LLMs can see and hear without any training Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.186659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.186659Z digest=sha256:b8f8bda841688bfae5896db85eb1e938d191ba7e8fa3057ff209cfcdc92d7bea

Observation aa7d852a-d09a-49d4-bb44-801c457c92ca · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

LLMs can see and hear without any training High-resolution image synthesis with latent diffusion models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.191571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.191571Z digest=sha256:1a6bcee1962ee877be956f283c66916c312a55f465893d45de749d825218fbfb

Observation 0874f5f2-fb01-4a21-af90-4158959c883e · outbound

This paper cites L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., Ho, J., Fleet, D.

LLMs can see and hear without any training L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., Ho, J., Fleet, D

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.812005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.196060Z digest=sha256:f2e30c4d8c6d3a95319e6606e987fc3776ac7e062ec709618bf2002baba6a186

Observation 7728eb58-c5f2-4394-9cd9-8e7d72f1312f · outbound

This paper cites Zero-shot audio captioning with audio-language model guidance and audio context keywords.

LLMs can see and hear without any training Zero-shot audio captioning with audio-language model guidance and audio context keywords

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.796723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.201000Z digest=sha256:c5c5f0d5c5d6f1794b0103a5c7397c6371df7786ae78a99b69afe34647934de3

Observation 566a6112-59e8-49d9-b567-3629ee02ecd8 · outbound

This paper cites Zero-Shot Audio Captioning via Audibility Guidance.

LLMs can see and hear without any training Zero-Shot Audio Captioning via Audibility Guidance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.205600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.205600Z digest=sha256:3e346d7aa4a0217fd8487154f573e4b754a460704f3cb25f50ea7f20867b3ba1

Observation da075a6e-f406-4cee-b340-a85ccdcbdc6c · outbound

This paper cites H., Tan, H., Bansal, M., Rohrbach, A., Chang, K.-W., Yao, Z., and Keutzer, K.

LLMs can see and hear without any training H., Tan, H., Bansal, M., Rohrbach, A., Chang, K.-W., Yao, Z., and Keutzer, K

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.780741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.210966Z digest=sha256:ddcd65fe683347711955b7417d86741eca2e3754cdba4f3d48a30fe2ae29902d

Observation 266e9844-c78d-4d8b-9996-41f9ca698977 · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

LLMs can see and hear without any training Emu edit: Precise image editing via recognition and generation tasks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.764846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.215648Z digest=sha256:a3aa0d6123318a667570622fc7c7043e9267a04d6e8354f8b4a85bf717cc818b

Observation 93a29c0a-a145-4515-bed3-dfa693881bb7 · outbound

This paper cites and Zisserman, A.

LLMs can see and hear without any training and Zisserman, A

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.220243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.220243Z digest=sha256:adac5c8334802ab84f32f6fc5caf8fa0b7859f2a1fa1f46dccd396df6979be9a

Observation 5f0924f5-8430-486a-9718-8c1f6a86e3d4 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

LLMs can see and hear without any training Gemma: Open Models Based on Gemini Research and Technology

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.224799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.224799Z digest=sha256:8f588748db4ef5a41a0fa8e7598b00c6d08c327098d236202f6053988af354c8

Observation 45254f0e-ab4f-49be-85de-a6f2c7572252 · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic.

LLMs can see and hear without any training Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.738738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.229609Z digest=sha256:148a19d20f1be9b8d13c7176963bcf9fd76e65883301c697b84279052393fd81

Observation 5138eade-5d48-4783-bde5-0238c7b8605c · outbound

This paper cites CIDEr : C onsensus-based image description evaluation.

LLMs can see and hear without any training CIDEr : C onsensus-based image description evaluation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.724624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.235076Z digest=sha256:7b6a27a95914db03a63a785f88f560b21691e29bccd6d12c12c940793892f0ff

Observation ba63aaaa-e306-448c-9b55-4807c0e9e945 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

LLMs can see and hear without any training Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.708850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.239669Z digest=sha256:a311d7d6589c94689fb17f96e83cd601649236e67baf7b5c2be7b8f57066fe39

Observation 3c11ff3c-043c-47d7-b8b2-5a1e44f33196 · outbound

This paper cites V., Zhou, D., et al.

LLMs can see and hear without any training V., Zhou, D., et al

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.692953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.244253Z digest=sha256:8cddbd1eb578f95eb41d5de4e4bc3d0eb0cadb0ce266df1223dfbb1b7495c239

Observation 44818b35-f4e7-42e0-bd36-ffc37988b3c4 · outbound

This paper cites E., Huang, P.-Y., Howes, R., Sharma, V., Li, S.-W., Ghosh, G., Zettlemoyer, L., and Feichtenhofer, C.

LLMs can see and hear without any training E., Huang, P.-Y., Howes, R., Sharma, V., Li, S.-W., Ghosh, G., Zettlemoyer, L., and Feichtenhofer, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.675933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.249134Z digest=sha256:157b1f588671e8bfc7f2d05de2aea3705816010aa35721d943eeadbafa929101

Observation dc653a20-dfca-4130-9302-d9e9532d66bb · outbound

This paper cites MSR-VTT : A large video description dataset for bridging video and language.

LLMs can see and hear without any training MSR-VTT : A large video description dataset for bridging video and language

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.660540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.253807Z digest=sha256:03a7357a970b1216f1b8d02440a13561412d5172df1610483d6bd3c38ac164ed

Observation 69bb8738-4a48-403b-ab0e-cbe856820531 · outbound

This paper cites V., Zhou, D., and Chen, X.

LLMs can see and hear without any training V., Zhou, D., and Chen, X

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.645005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.258318Z digest=sha256:912b45fc1c749713b49a374e4f0af41d651b228461190782bebd4fe60b13f984

Observation 3c9f03ad-4cec-483f-ad2e-a15fc1827c61 · outbound

This paper cites ConZIC : Controllable zero-shot image captioning by sampling-based polishing.

LLMs can see and hear without any training ConZIC : Controllable zero-shot image captioning by sampling-based polishing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.630321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.262763Z digest=sha256:0aa2636630902d6b5389b507ec930a81bc3d824cc1d5c36acbab4dd340b9de53

Observation 4ed7780e-fb70-4000-b3f7-e11c6360ea04 · outbound

This paper cites Meacap: Memory-augmented zero-shot image captioning.

LLMs can see and hear without any training Meacap: Memory-augmented zero-shot image captioning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.614899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.267580Z digest=sha256:f3c8402d9521f6a6c10e99bbf2d91fccd47790b9c8525c4f8fd64bf35c5ddc59

Observation fc8b6ceb-0b37-4262-907d-dc8bf1a118a5 · outbound

This paper cites Sigmoid loss for language image pre-training.

LLMs can see and hear without any training Sigmoid loss for language image pre-training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.272230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.272230Z digest=sha256:36c64b8a39a97e94e24a42af150a7a04079366cae698cff68042a9afdad6b1a7

Observation 88cde246-7d4e-4957-9ac3-b77d3112e4f5 · outbound

This paper cites a henb \.

LLMs can see and hear without any training a henb \

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:46:08.588755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:46:08.276886Z digest=sha256:3aedecb5bdcfa0ba7877a21d95c19f50ef3bcf7844aae24af752670ce846393d

Pith citing papers

Observation 07f8e3c3-0c4b-49da-9c4b-4fc59bc741a6 · inbound

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents cites this paper.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents LLMs can see and hear without any training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.719320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.719320Z digest=sha256:a766851a20732e389a2fb0aded91223a9266ae08c774089d181183188a0b0fad

Observation 234690bb-8c5c-4572-865e-afc404bf07a3 · inbound

Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models cites this paper.

Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models LLMs can see and hear without any training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:43.276306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:43.276306Z digest=sha256:c81edd9b88602af53cd1a94e509bdbbde7d0d25114914e08ef0e85d9bc235175

Observation f7cc7435-15c9-461a-aad3-117776b30975 · inbound

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models cites this paper.

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models LLMs can see and hear without any training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:53:11.782964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T17:51:32.439732Z digest=sha256:343beb5088d8cafee72afcbacbc59fb25c32e2f8945e2a50991dc45f2ee0d91f

Observation 03fd9dca-c859-42c1-bd48-dfcfb2d0f5bb · inbound

Personalizing Text-to-Image Generation to Individual Taste cites this paper.

Personalizing Text-to-Image Generation to Individual Taste LLMs can see and hear without any training

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:02.475416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:07:31.236729Z digest=sha256:7b1b8fc57b06d129482baca81dc3fe9a14713873db5653a5f0419b931fe173d7