Pith. sign in

Paper Citation Record · LEDGER

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2505.14562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14562 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:15.094780Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.551896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:41:14.663492Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99e7f0db-e75a-41eb-b099-380885e3ee43 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning transferable visual models from natural language supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.549401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:13.452064Z digest=sha256:a9dd8dffd06b661fd2c58c78bb821c34a1be3cda90e137f68f89a915779ed9e0

Observation 2b93de3d-88e2-43dd-a87b-ba987bbb8709 · outbound

This paper cites CLAP learning audio concepts from natural lan- guage supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities CLAP learning audio concepts from natural lan- guage supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.408674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:13.496702Z digest=sha256:c03f14cea0ccf4a1243121620f544b247bd169377c06fde8660c5e99a3a5cc8c

Observation fabf8d23-e5c5-46de-8acd-fa912c52bd6b · outbound

This paper cites Multimodal learn- ing with deep boltzmann machines,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Multimodal learn- ing with deep boltzmann machines,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.235540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:13.595143Z digest=sha256:9f78e2c69d7646ed241755e9ad0ae0cd3db3322cd526c908439cf4d0144df971

Observation e96eb8f0-4007-4a5a-aa42-9441bc759fe2 · outbound

This paper cites Learning visual features from large weakly supervised data,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning visual features from large weakly supervised data,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.052093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:13.679426Z digest=sha256:dcff4ef7e45078ab4896d22ac6cfdb1ebb0a9acb0a86ba5e592ee3dd989ae51a

Observation db6b673f-9dbc-45ab-9532-31d92af84d9d · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.912027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:13.759296Z digest=sha256:0b3650868188f404fb1efecc0ad747e9ae8835160c9c7f0115b15a1b15c1ead9

Observation 32e3baa8-88bc-4130-b9d7-aaa8079621cf · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Florence: A New Foundation Model for Computer Vision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:13.877031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:13.877031Z digest=sha256:b39ba9bc1f8031257ca1283a6582257d9cf0ca740df4c235c7bb1ee279cc4b07

Observation bb57bd20-a237-4aa2-84a2-75e9ed2c7cdd · outbound

This paper cites Wav2CLIP: Learning robust audio representa- tions from CLIP,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Wav2CLIP: Learning robust audio representa- tions from CLIP,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.768700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:13.958279Z digest=sha256:16d08100379fdc6bb3d92a8a2f7fb41e028e4bdce08ffc7efb2d8b289eeb18c2

Observation 396192f0-8fa6-4f63-a623-313544346d70 · outbound

This paper cites Microsoft COCO: Common objects in context,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Microsoft COCO: Common objects in context,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.591992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.079730Z digest=sha256:614e112c035d01bb8b53ebf5aa66ad0aaa6f7a1bd78dee899d323e668e4b2aee

Observation 64a99d06-0983-431a-90d9-1fd6fb60e89d · outbound

This paper cites Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.430857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.223807Z digest=sha256:fa541fa6f16eebe59e93e6c3d17fc59d8f2d99e43d2686b2fa5563b3ef31457a

Observation 07f760c5-d9c6-4ada-86c1-3a2436579f30 · outbound

This paper cites YFCC100M: The new data in multimedia research,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities YFCC100M: The new data in multimedia research,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.243583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.325160Z digest=sha256:fda85df41e97757f6de19d4c348017a3e315d3c53d825ee5ed3f619965b38220

Observation 9f471ed8-3cb0-4790-a9ae-65f3e2d60f93 · outbound

This paper cites FSD50K: An open dataset of human-labeled sound events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities FSD50K: An open dataset of human-labeled sound events,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.106720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.486000Z digest=sha256:343bb8856caf79572e1005fdf243d5c06310fb4c0c53a226fddd452e0c6b6700

Observation 68537183-119f-44b4-955e-a282dceffadf · outbound

This paper cites Clotho: An audio captioning dataset,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Clotho: An audio captioning dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.547487Z digest=sha256:636b0a7b103d4634f1d410b117d28f95f9f603cd6545f6fd966eb5b0ce925174

Observation 67bd587a-e2c3-4fa3-9363-a1501bb6a625 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities AudioCaps: Generating captions for audios in the wild,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.814031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.665385Z digest=sha256:0a73a2d4d881dce5b800d70fb63ec254113fe6d85a268cb35a3596f0c97eb44c

Observation 0aeba42d-176e-414e-9584-807e916a9439 · outbound

This paper cites What is the ground truth? reliability of multi-annotator data for audio tag- ging,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities What is the ground truth? reliability of multi-annotator data for audio tag- ging,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.632294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.761764Z digest=sha256:34dadaf5e28b3814657d0cce8cb63c8ce02253fa1a6e2dad25407db521132294

Observation d814e34b-9f97-4733-98df-56ee70e7696b · outbound

This paper cites Audio- CLIP: Extending CLIP to image, text and audio,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio- CLIP: Extending CLIP to image, text and audio,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.484053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.816806Z digest=sha256:a25995df2f095f0d8b4168dbfe28a37925d795b509633f2033d3de4551becf5e

Observation 2ece6d4f-4a63-40ea-a543-ceba1cabcdaf · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio Set: An ontology and human-labeled dataset for audio events,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.380850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:14.933049Z digest=sha256:a44083b5d89f7e2d17fa5fd813c9f5d7b9ffaf9f6d164c6c245d18ba7df4ae39

Observation 3dbc9405-fb1f-4ae3-a221-602119900054 · outbound

This paper cites Sudarsanam, I.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Sudarsanam, I

Reference 17

Resolution
verified exact
doi, observed 2026-08-07T15:36:15.236385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:15.021689Z digest=sha256:008017f36faada6398b0f12dfed0c5928bcbb03b43132e4d858752a1fac937c7

Observation 7a0b3938-31fc-49ce-ae4f-2a5ff4c5c74e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Representation Learning with Contrastive Predictive Coding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:15.094780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:15.094780Z digest=sha256:773d9869d040448bf9bfe4dc6533508155c5ba9cb196b36cbc93612a4c5ead98

Pith citing papers

Observation 367aff17-6b04-4fb6-9b8f-f32f317f4eed · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.551896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.551896Z digest=sha256:097c7e60dad49e6f7eff44229a8c49601095c38d129f5c90b4aaf3e9be6c32e3

Observation 3301051a-947f-4ff9-90c4-2bd38cb9021d · inbound

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception cites this paper.

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:41:14.716289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:41:11.839628Z digest=sha256:bd5d28aff40ef7e16ca501a4353bcf8035d4cc2d0e716022401778c6b48b21ed