Pith. sign in

Paper Citation Record · LEDGER

FG-CLIP: Fine-Grained Visual and Textual Alignment

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2505.05071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05071 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:54:59.478902Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.758647Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aa57dabe-d4a1-4abf-9dbc-0aa92135ffcd · inbound

OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning cites this paper.

OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:54:59.478902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:54:59.478902Z digest=sha256:d5b98b6db5a8e340a06742adf5756cbbfb4d60676b13f306799a043ad3aa76c9

Observation 5da17b1c-2b8b-44f0-98a6-3cd8ef72f4e1 · inbound

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation cites this paper.

SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:58.099238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:38:58.099238Z digest=sha256:414eddcebcec41eb84a78c4e3fd3de370406c091573280a38ee2daab4ce14eb3

Observation 4611edf2-8eb8-40cc-b249-c6b37710c0e2 · inbound

Emu3.5: Native Multimodal Models are World Learners cites this paper.

Emu3.5: Native Multimodal Models are World Learners FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.615646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:564e827ceef1a93c71daddf02ffc054830db4651c4f8e7aecbde0f5cbb16d595

Observation d670af6b-6e6d-4657-919d-66937fa6d802 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.287769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:89a1658546d56113398bf4f23c5bd1f6907073121bed1657c2c31c89b6c33948

Observation 5286bb90-ce9e-4975-8e5e-ddd122d783b9 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.553656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:7b35777dbb7aed450e7a31bf669338a0e34e5ad90d8dba240710317211aab46d

Observation c175fb72-8b60-4568-847c-979d9611ee67 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.036252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.036252Z digest=sha256:ac3ab672f54c3a40ece1fbb43e87865ef6c28bd5951b1e6e454c1aaad62be342

Observation 7373a016-8256-41ca-82f0-b07f3affd1e9 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.903003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:1fd385dae081d0aafc340c479e77288e414095bad5c2aefe915fd077c69a35e3

Observation 6bac1155-41c3-4421-80a9-d973d6a95d1c · inbound

Mitigating Multimodal Hallucination via Phase-wise Self-reward cites this paper.

Mitigating Multimodal Hallucination via Phase-wise Self-reward FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.259432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:10:45.144421Z digest=sha256:2f3d11ea5f938349d0575941381f4e54b714a19317716a92c49c0018eb1606c3

Observation 9f7720e3-99ec-4222-b6b1-9508dbd3ab9d · inbound

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce cites this paper.

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.736470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T00:55:56.146885Z digest=sha256:547c2d5a8246cdf45bc8b9c977c86b503deaeeb17b5f07b2848d1862b4f77c22

Observation fc4fc96e-58b6-4b8e-ad78-3a6af6b69bcd · inbound

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations cites this paper.

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:09.981957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T19:33:24.911067Z digest=sha256:5061205718277e919756781bbeda677786cf10dea16f29357ef3fca71edc2319

Observation d7bc418a-5b55-4a14-8883-44847e771448 · inbound

LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment cites this paper.

LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.757322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:38:09.299209Z digest=sha256:c633320e345b89b232e60995f328f6c409b3a8798a583a4b18adc0267755ee50

Observation d593473c-b77d-4a0e-ac67-f2dd3a9717c1 · inbound

L2P: Unlocking Latent Potential for Pixel Generation cites this paper.

L2P: Unlocking Latent Potential for Pixel Generation FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.727524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:12:28.181595Z digest=sha256:0c73b8e48fdf0fde4886e7ceb44ad7d36132c7196a3867aef0a8bb07ef232100

Observation b4d4030b-f343-42a7-947b-d622d4549ec1 · inbound

CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection cites this paper.

CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.582680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:13:34.998844Z digest=sha256:35827271a47d464c5fb767637bf65ea22e6820eaad23725cffee3210dcafc4c5

Observation b30c20c8-4764-47fd-9777-738b298868d9 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:40.789072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:14:35.109298Z digest=sha256:5ae18dbdd3d1da0c06c6fb4946424b4e206851578b92726054be91d1b1374c1f

Observation a44653e4-c4fa-4e70-b2de-5d8a0cb884d9 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:39.629755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:29:51.635039Z digest=sha256:0aa9484ce00f54ddaba482cff4c73ced0fc8762dad0ad608097d62621dcb38e3

Observation 04bb5bc8-c02d-4539-9e83-84470608a121 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.760118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:704aaa8700b3675bbf1eab4532efe0f848c15bb9db535b89dd62ce12132f4e56

Observation 85a38dce-0093-4b5c-9272-47eb371fb75e · inbound

InstanceControl: Controllable Complex Image Generation without Instance Labeling cites this paper.

InstanceControl: Controllable Complex Image Generation without Instance Labeling FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:45.109676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:37:41.030752Z digest=sha256:209fe72dbb6041dec9e83d26033dbe2266251d9ebb9e477b0c3993cbcae5da41

Observation 8addccce-036f-48d6-98bb-def375b1dbfc · inbound

DialogueVPR: Towards Conversational Visual Place Recognition cites this paper.

DialogueVPR: Towards Conversational Visual Place Recognition FG-CLIP: Fine-Grained Visual and Textual Alignment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T14:39:42.933706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:39:42.933706Z digest=sha256:e13a88bc0ca1a39ef9affd707e319abe948f24651a72f06625f245bd030a95da