Pith. sign in

Paper Citation Record · LEDGER

ImageBind: One Embedding Space To Bind Them All

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2305.05665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.05665 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:37:33.205620Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 000f3bdd-248f-4333-abc4-be74ef8d2baf · inbound

PandaGPT: One Model To Instruction-Follow Them All cites this paper.

PandaGPT: One Model To Instruction-Follow Them All ImageBind: One Embedding Space To Bind Them All

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:01:37.127118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:01:37.096206Z digest=sha256:066c380feb128af657ae0d08c6dc99c42a469c2f0a5ede4708535a01fc405249

Observation c08849fd-f62f-447d-9d13-85cc380c45a7 · inbound

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study cites this paper.

Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study ImageBind: One Embedding Space To Bind Them All

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:37:33.205620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:37:33.205620Z digest=sha256:4e16269aa6d5b6b4a0d445bc5b6a28cc0864c2b1462b83cb5c0318877df6d19b

Observation 8420d55f-4461-4204-8bee-fa62af796319 · inbound

Learning from Limited and Imperfect Data cites this paper.

Learning from Limited and Imperfect Data ImageBind: One Embedding Space To Bind Them All

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T13:09:57.253731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:09:57.253731Z digest=sha256:c176d844401072f891c05e0c852faaec13093378f3c7ab4224c2e2ddb42c537a

Observation cc9105c9-c3f9-462e-85db-995cc7d6b868 · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.656278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.656278Z digest=sha256:1a5dc4af962d890daad72ee628ae95bc25f353e086547153cee134d7c28eba81

Observation 09e8c540-3a19-48e0-8d72-22511bc0c912 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation ImageBind: One Embedding Space To Bind Them All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.782500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.782500Z digest=sha256:052fc84256279d2ef626cbc87fc65b06ded59f0ff7a6423a113f503907960466

Observation c72b6f62-dcd0-4c93-a0f1-c5b58f530078 · inbound

Artificial Phantasia: Emergent Mental Imagery in Large Language Models cites this paper.

Artificial Phantasia: Emergent Mental Imagery in Large Language Models ImageBind: One Embedding Space To Bind Them All

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:44:22.697166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T21:41:39.111769Z digest=sha256:04f0d90dcdd132092600625db6d0cdeb648cc8add8cc7220ebbd796c08a7e372

Observation 5a106b2e-9977-4464-bba4-fa6e18ed3986 · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models ImageBind: One Embedding Space To Bind Them All

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:46.287773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:46.287773Z digest=sha256:d6f8da0354512a9fab450ac8dbf1955da6832083147d9e11e80383c218240a73

Observation 965d9d39-ac50-4bd5-a220-74b627f6fd4e · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models ImageBind: One Embedding Space To Bind Them All

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.505525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:7688ef4049ad15c983aa6f5e406646bb263429d6e6b3ab9c6c312a89504af071

Observation a3751b5a-98a1-45f1-a77b-efceaf80bb50 · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models ImageBind: One Embedding Space To Bind Them All

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:02.077149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:529980fe8023f5954ba8938a4fbfa307f3319140e24c72eb3425999c170af3b6

Observation 4f017f60-e75c-4d55-b1b2-38c786c9510c · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models ImageBind: One Embedding Space To Bind Them All

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.761245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:c40d6bde5b6e21eae99995048c9930d6980c72d15023778413638e0f91a7226d

Observation 174e19d0-a978-49c4-bf2a-f34fa0a26dd1 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models ImageBind: One Embedding Space To Bind Them All

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.931907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:cdf715577ee32fba7806ee9770f4676aebf3aeb4254e219f38aac87805e69719

Observation a2e579a7-ca37-40b8-8b01-43f79547c9f5 · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning ImageBind: One Embedding Space To Bind Them All

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T14:10:57.788301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:5c8341e84ab5194661572355a97ea77d5730a8c0ce3fca05119caeb06a3a9c27

Observation 8eeb4785-dba7-400a-97bc-01df1b8d7591 · inbound

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning cites this paper.

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning ImageBind: One Embedding Space To Bind Them All

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:14:56.772592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:14:56.772592Z digest=sha256:8ab2eca4273d01f97e466ab41a55d423f7f543f7f16bf1bb3811a4d26e70937c

Observation ca4e6d03-8add-4756-a62d-18245f5a083d · inbound

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio cites this paper.

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio ImageBind: One Embedding Space To Bind Them All

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:46:35.520484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:46:35.520484Z digest=sha256:290edc665010fb7a03daaf470d0d1c47b4a4dd07733701c5bbc2c1b5b25f533d

Observation b7e0f27c-b379-42e9-8c91-8930a251271b · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation ImageBind: One Embedding Space To Bind Them All

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-30T10:34:59.307223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:34:59.307223Z digest=sha256:a353d34154bdba4a8973a16047da8330b5a196951794c114461d4ec5729d7883

Observation 83822ea9-a08a-4d7b-b085-7fa3729229f7 · inbound

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation cites this paper.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation ImageBind: One Embedding Space To Bind Them All

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:50:57.995004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:50:57.995004Z digest=sha256:f1467afc7f8e808c6cb8344de2aa854e98aa257698dae47f63a0e1964af52055