Pith. sign in

Paper Citation Record · LEDGER

Improving fine-grained understanding in image-text pre-training

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2401.09865.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.09865 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:31.506605Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45f8f3ed-0b7b-434a-80da-61b24efeec82 · inbound

QuARI: Query Adaptive Retrieval Improvement cites this paper.

QuARI: Query Adaptive Retrieval Improvement Improving fine-grained understanding in image-text pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:31.506605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:31.506605Z digest=sha256:e59efcb354864f5fd9bac68f01c4753726adeb8e5dcbd5c4707637a7c7347704

Observation debbb057-9995-4c62-95c4-2984a88e9df7 · inbound

AGA: An adaptive group alignment framework for structured medical cross-modal representation learning cites this paper.

AGA: An adaptive group alignment framework for structured medical cross-modal representation learning Improving fine-grained understanding in image-text pre-training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:49.182317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:53:49.182317Z digest=sha256:ae3c4294d91a4f2b641d6310c9451db8b0d1cbd1c72645934dcb987af5d7a9f1

Observation 61d1197a-075e-4d58-852b-429f3ce781c9 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval Improving fine-grained understanding in image-text pre-training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.304921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:c7a5fac1cc357774f71429173b90093422a06b53d5bf94030dbce6c467f06bc5

Observation bc6b390b-0fa7-4974-ba07-ac700ebf8f6d · inbound

Xray-Visual Models: Scaling Vision models on Industry Scale Data cites this paper.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Improving fine-grained understanding in image-text pre-training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:57.884381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:57.884381Z digest=sha256:53e30d64aa8e0ab64dc29ae07f67095881104faede68a06db9333d70a80a37d1

Observation 5eb1d7db-baf2-4eef-8c1a-9a3f303c495e · inbound

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding cites this paper.

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding Improving fine-grained understanding in image-text pre-training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:31:03.805393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:26:55.369840Z digest=sha256:52986fae536ecbb62120b86747887f9eb45d5d2875364341cf3ef74ae5d3bfcf

Observation 5f5feb4e-38ea-4f69-92a4-f072077772b6 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Improving fine-grained understanding in image-text pre-training

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T19:05:10.550927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:2a43d690b896825642158b2bd62fd1310d98b4f2d26fa4d39ca240dd3b452b6f

Observation eac0458d-bad5-49f5-9cdd-f0cc0579e1a3 · inbound

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning cites this paper.

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning Improving fine-grained understanding in image-text pre-training

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:23:15.566401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T08:16:57.329872Z digest=sha256:56d1f50daba62211083e2a17fed23e9bf8083d24a753f6c4162ee494990e6668