Pith. sign in

Paper Citation Record · LEDGER

Contrastive Localized Language-Image Pre-Training

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.02746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02746 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:19.724749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:19:13.788988Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10317e8e-dfcd-4da6-b919-f54e55f45b5d · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Contrastive Localized Language-Image Pre-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.724749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.724749Z digest=sha256:cb0e19547d92d601a6e9e167623e80c18a96351e563c373de711df2a2e60b2cb

Observation db769080-4e8a-4ec9-bb79-9fcfc021e317 · inbound

Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures cites this paper.

Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures Contrastive Localized Language-Image Pre-Training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:31.659608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:31.659608Z digest=sha256:8cbf0752298ba1821a6f8312d7f5ea71d4969452a3b1f8f7e4a4b718033bbd7f

Observation 97a0e847-3537-4f22-8915-6afb4ad9b595 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Contrastive Localized Language-Image Pre-Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.225148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.225148Z digest=sha256:3ec5ec027e6670bbdef4e77ca91e4486a4b09bd869afe88c91a06eeea618193f

Observation 41ed5110-7327-4e3d-bafa-1a1622ef2ee5 · inbound

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting cites this paper.

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting Contrastive Localized Language-Image Pre-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:04.661105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:04.661105Z digest=sha256:df5696597acdb933bfd012ce6ab80b58c21597c1926454f49e1b2bbeeba6e02f

Observation 21d5ce15-807a-4a5b-b0db-72252ee48611 · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Contrastive Localized Language-Image Pre-Training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:42.682062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:42.682062Z digest=sha256:d04cbed25523cc6146a42060406c0cb97f2f363bef8584ffc99f0afca299571e

Observation 206d0ac7-b75d-4f82-bf73-4bb50d3769c9 · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval Contrastive Localized Language-Image Pre-Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:30:48.231935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T09:29:32.250418Z digest=sha256:8d1c95f435bc258fb474b1fcad450f6a83e7a5e0973cc3f5549709f6139d098a

Observation b6db835f-6bf8-4700-a978-b909f73139ca · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval Contrastive Localized Language-Image Pre-Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:07:40.022288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:07:40.022288Z digest=sha256:a12b74414466b743f687f63f9bf9682f178246843ddc42d5c928f2da76ba016f

Observation 14c2ff8d-487a-420e-9634-def8d146bdca · inbound

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning cites this paper.

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning Contrastive Localized Language-Image Pre-Training

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.562958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:16:57.329872Z digest=sha256:396ea1fecf098db3707c1926cc7f6a63eeff34e5c99c71c550ae1653c017b280

Observation 3c81927d-03e5-4188-a63f-84cfb6a3c2b0 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search Contrastive Localized Language-Image Pre-Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.876040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T15:18:17.427707Z digest=sha256:053b7863a539b741d9decdd97f265b8b8d75ed2b5f7a79d3cf2a1019a5583242

Observation 2c58a7a9-7f40-4e70-a6dd-a1069d22512b · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search Contrastive Localized Language-Image Pre-Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:28.864002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T23:17:03.456746Z digest=sha256:1233bddc2f1c5ca5315eb9166e98aaf3f0cf8eecee5bd926f5444361588458c7

Observation 7cb20dff-3047-436b-b11c-4a52cd1dafc0 · inbound

Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks cites this paper.

Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks Contrastive Localized Language-Image Pre-Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.790413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:16:21.252304Z digest=sha256:87aed9d917281ffb1cad99aa8571e4f1a782e1471e0cf5510f5be8da7936d205

Observation 29fc88c7-d200-4558-8e6a-94ca6844798f · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition Contrastive Localized Language-Image Pre-Training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T11:31:04.191369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:31:04.191369Z digest=sha256:7fdb2d83b23a44091bad5b0ca88545e94035ea73111867d912e3612f77c061f0