Pith. sign in

Paper Citation Record · LEDGER

Long-CLIP: Unlocking the Long-Text Capability of CLIP

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2403.15378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.15378 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:40:21.599948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.890361Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03a6e17e-2b29-4780-b90a-e008cf09c589 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.852840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:de6e75a205cecd4e2c4cb01a404d92342e7f6801b99a3b2272bcf97f92882658

Observation e14f8811-1c17-4665-b344-4424d20d57a6 · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:20.952536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:368a869c6e58f41b815fcb52652b9829d141c0e25d00a8be15d3f2028fc7e463

Observation 6cb0ee77-9ec0-4717-8e75-9b406f2599c7 · inbound

Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency cites this paper.

Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T23:40:21.599948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:40:21.599948Z digest=sha256:4f40ba1c6cf501f3ccfaff9be33a7dae3d24ecfd26829e4e82d51f3215a7a236

Observation 5b51070d-d63d-4fc0-976a-77c714bd43a5 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.217443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.217443Z digest=sha256:ea9680ee2d75ba7393334c81a7dc1ad1a3c5b0f5a252a8297f64a02d14b2bf85

Observation 1d6f4e5d-57fc-41b3-9b1f-fd9d40e586a1 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:55.385363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:55.385363Z digest=sha256:45eb7b4ebdb50fe5f4799b6387655d6bf6b7a501113e61bb787a377c558306b1

Observation 766675df-dbfd-4115-975d-536f1a31855b · inbound

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations cites this paper.

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.089363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.089363Z digest=sha256:d32333662b168e993a07fdcf1723c039844f66c90096f4a41db3a29538309e3e

Observation 5a9871a1-57b7-40e4-bd16-eda9ed779b15 · inbound

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model cites this paper.

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:47.273964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:47.273964Z digest=sha256:039b03d8c568286fcd749129049e9aa8ddb26a479f2e3425b14634781872cae9

Observation 38f709d8-3a45-402a-a049-1a51bc49f3f1 · inbound

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles cites this paper.

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.804567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.804567Z digest=sha256:0fddda46d45cbbf2c4ee4195c67ae49d254320f99387b214bc2b4a8b9c494ad5

Observation d91d90a1-acd1-4d8e-ad0a-ea4de23d1568 · inbound

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text cites this paper.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.501373Z digest=sha256:a3ccccdb7d816d34332a447f697a52e547d96905381acaca08e01928912a8320

Observation 7fecc64b-1d59-4c7c-9aa8-12589220b431 · inbound

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation cites this paper.

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:57.092222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:57.092222Z digest=sha256:e8b6ac667e4b8863c29bf37848b48beb705424f61442ddd8734f598a671a2a41

Observation 4160fdf5-77db-4682-8f0b-cf6b232dfb42 · inbound

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement cites this paper.

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:22:49.772744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:22:49.772744Z digest=sha256:c6bfaa9e14602dc375bab557e359bcf4ef702fa36f82a06d6de42be738ad3232

Observation c4c08ade-a215-4f17-8092-44de9aca2019 · inbound

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach cites this paper.

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.671528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T12:39:56.396231Z digest=sha256:c2bfceec50e195abd34cda63fee98292f0de75cfacad31ccd23f8309802930fe

Observation 21865d75-2172-4a81-abe7-a9c3a11ce13c · inbound

LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning cites this paper.

LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:22:52.400078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:21:19.606631Z digest=sha256:bd0a6644c92da815bbb23f4e2fb4c710e39149c75690876abf6f886e3f7d461b

Observation b9a0f02a-d2d4-4809-99f8-99d96cc9c187 · inbound

AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting cites this paper.

AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:50:56.206901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:50:56.206901Z digest=sha256:bc2005080aefaf6ad0e14ed359a3b90ac1db02283c79e49ac9b78b68f7459b43

Observation 5b837b95-8915-4b84-aca9-6f6f0cc107e9 · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.288367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:96b548aeb66ebe04296b918437debb7810058fda0c6c1f670a3ee49d1bf7da05

Observation 84f41b62-ca6e-407d-8f7b-90a806f33b21 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.725478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T01:52:14.874049Z digest=sha256:c673ca129924d6119609b01bfc3bf99594fea6dec57a23165e9106a05a1c1d10

Observation 62d63b4f-850f-4b2b-8183-983e4e591e26 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.461323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T21:54:33.902256Z digest=sha256:ca9f110b17b495e568eb6bf0a0ea6d75266fddec496588c41deb1643f39baad9

Observation 873a7819-028c-47e2-a0c4-a30081ab958e · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.624844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T09:12:35.777810Z digest=sha256:d6745df8bb258cd81eaeb22a28f2e47a3c5edc3f7bf31ebb3876c2494e3b3d41

Observation 65d9bc3f-5700-4c3f-847a-0ce1a24e5f20 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:45:05.533477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:44:21.427570Z digest=sha256:4c2740e7a01a830bcb70d935306eb2b187290ddb931767f5cd6458eb3defa0f3

Observation ac12bc4f-551e-49c0-9619-29fecc202c2d · inbound

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation cites this paper.

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:43:38.647922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T18:41:12.533556Z digest=sha256:e0f8105e9067eabd59ef6af8ef9c79e94d872b20c74a87cd243c9fb6f97736c0

Observation 29d9322a-d926-4c8b-a69c-81a353b8bb01 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.630748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:5b09587b4ca78e44dc0018ff8b56846b7d6c6cb0d00fe8001474a197229558cf

Observation ab0c928c-5f9b-456d-95c7-e6a77702e2f0 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.891855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:29:41.291832Z digest=sha256:b1404bcb659ceac495e9290a6eca247986b603577b2b4efaa864a021bfbc39b2

Observation f8971470-a521-43ab-9a07-60b75aa612e5 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:03:32.224194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:a62c5afeb05308bfbe650a9dbf968429af8b84bb05c0dd4e2e155eefec783379

Observation 016f124c-dda2-4a77-b01e-e8e688f0ee17 · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.757577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T01:02:20.024793Z digest=sha256:3f77c3513927e183d4c2193a58ed579438cda8b8cf8ea6262c5238ff3e90e2ad

Observation c9eee4a0-f03e-48e6-b829-770ae20f3b26 · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T11:46:50.520083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:46:50.520083Z digest=sha256:1dc52cc46b5ff3c10c21e4d42eac44b0d9e76264ecf47d5f6ae21af48bd5f6fa