Pith. sign in

Paper Citation Record · LEDGER

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2505.14562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14562 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:15.094780Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.551896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:41:14.663492Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99e7f0db-e75a-41eb-b099-380885e3ee43 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning transferable visual models from natural language supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.549401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:13.452064Z digest=sha256:c18c4e5540bdfd29c9a40d647e1969acea844b42bef1873a82f14d5b169ae563

Observation 2b93de3d-88e2-43dd-a87b-ba987bbb8709 · outbound

This paper cites CLAP learning audio concepts from natural lan- guage supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities CLAP learning audio concepts from natural lan- guage supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.408674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:13.496702Z digest=sha256:6735b4b65adce9c42feabf2b7f56421ad0d247d82317909a01425fddf0ef0990

Observation fabf8d23-e5c5-46de-8acd-fa912c52bd6b · outbound

This paper cites Multimodal learn- ing with deep boltzmann machines,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Multimodal learn- ing with deep boltzmann machines,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.235540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:13.595143Z digest=sha256:aabc009c70b3d7bd0c8f3e4118f7afe45dfa660e58592c958a33a454f9d88e20

Observation e96eb8f0-4007-4a5a-aa42-9441bc759fe2 · outbound

This paper cites Learning visual features from large weakly supervised data,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning visual features from large weakly supervised data,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.052093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:13.679426Z digest=sha256:496b77fd265c50eeb22e8a4ebc6e0d3a2e0c54ea6b296420fe10c06c092a360a

Observation db6b673f-9dbc-45ab-9532-31d92af84d9d · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.912027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:13.759296Z digest=sha256:c345d6228b0c5c7acec96736788888e6af0e12ee27382e3ce2937af912191de7

Observation 32e3baa8-88bc-4130-b9d7-aaa8079621cf · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Florence: A New Foundation Model for Computer Vision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:13.877031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:13.877031Z digest=sha256:b39ba9bc1f8031257ca1283a6582257d9cf0ca740df4c235c7bb1ee279cc4b07

Observation bb57bd20-a237-4aa2-84a2-75e9ed2c7cdd · outbound

This paper cites Wav2CLIP: Learning robust audio representa- tions from CLIP,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Wav2CLIP: Learning robust audio representa- tions from CLIP,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.768700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:13.958279Z digest=sha256:2f19f4865b4e7dbb7f82f2257a8c72785f447adccc47b73e6dc60f6fc7b1ecf1

Observation 396192f0-8fa6-4f63-a623-313544346d70 · outbound

This paper cites Microsoft COCO: Common objects in context,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Microsoft COCO: Common objects in context,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.591992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.079730Z digest=sha256:fff0fc5a62fd7e0fbdbced76032afce9db6d88e003e0f482fa56339c10720a5a

Observation 64a99d06-0983-431a-90d9-1fd6fb60e89d · outbound

This paper cites Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.430857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.223807Z digest=sha256:8466860dfc1631dd8f727e057ebd670a86567ac06db9e3f2565e7096ff43a241

Observation 07f760c5-d9c6-4ada-86c1-3a2436579f30 · outbound

This paper cites YFCC100M: The new data in multimedia research,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities YFCC100M: The new data in multimedia research,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.243583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.325160Z digest=sha256:c757adf64da1a81cbd0992f1d6257525c7eedc00d71b9892ee40d2f5673fd1d1

Observation 9f471ed8-3cb0-4790-a9ae-65f3e2d60f93 · outbound

This paper cites FSD50K: An open dataset of human-labeled sound events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities FSD50K: An open dataset of human-labeled sound events,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.106720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.486000Z digest=sha256:b6a95d5b8770b2b9a3382763783f36c7fa9158d526a6fd38e5b6505ff167d2c2

Observation 68537183-119f-44b4-955e-a282dceffadf · outbound

This paper cites Clotho: An audio captioning dataset,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Clotho: An audio captioning dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.547487Z digest=sha256:4db72e9eed387dca9858e2888a3e56ec6ceb8d268cfa81eb4f673cce2b93a5aa

Observation 67bd587a-e2c3-4fa3-9363-a1501bb6a625 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities AudioCaps: Generating captions for audios in the wild,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.814031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.665385Z digest=sha256:b5d9919346444ffb81713a437da16ecb0e17dea7cec98bd973dbb00b582f4fe7

Observation 0aeba42d-176e-414e-9584-807e916a9439 · outbound

This paper cites What is the ground truth? reliability of multi-annotator data for audio tag- ging,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities What is the ground truth? reliability of multi-annotator data for audio tag- ging,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.632294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.761764Z digest=sha256:7c6bfe0a8bcacea888c6ab9ff6bcf187d78aff8e3c1573c47c31f7072d4a83ec

Observation d814e34b-9f97-4733-98df-56ee70e7696b · outbound

This paper cites Audio- CLIP: Extending CLIP to image, text and audio,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio- CLIP: Extending CLIP to image, text and audio,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.484053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.816806Z digest=sha256:7e1e467966263daf95ef81acf7fdddfe05d35609a120d9dc41e97d5612686ea6

Observation 2ece6d4f-4a63-40ea-a543-ceba1cabcdaf · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio Set: An ontology and human-labeled dataset for audio events,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.380850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:14.933049Z digest=sha256:a3683ca0abf895c39e428faa44bac1b57eb7c62799e183e115a55309eae65d0d

Observation 3dbc9405-fb1f-4ae3-a221-602119900054 · outbound

This paper cites Sudarsanam, I.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Sudarsanam, I

Reference 17

Resolution
verified exact
doi, observed 2026-08-07T15:36:15.236385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:15.021689Z digest=sha256:25f974e6d9cecb7215d3e778ae3bb2fa5b7c4939a62b011142c58b85a7ba6f0b

Observation 7a0b3938-31fc-49ce-ae4f-2a5ff4c5c74e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Representation Learning with Contrastive Predictive Coding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:15.094780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:15.094780Z digest=sha256:c03453121a941ac5f70e3410ee379704313e8aa842561bb3374c87225cb8d7f0

Pith citing papers

Observation 367aff17-6b04-4fb6-9b8f-f32f317f4eed · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.551896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.551896Z digest=sha256:097c7e60dad49e6f7eff44229a8c49601095c38d129f5c90b4aaf3e9be6c32e3

Observation 3301051a-947f-4ff9-90c4-2bd38cb9021d · inbound

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception cites this paper.

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:41:14.716289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:41:11.839628Z digest=sha256:1cbdc0b92b0de81a6b046b3322b91edaf12c4ba0874e229073b3b2dd3512ba53