Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2505.05071.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:54:59.478902Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:39:50.758647Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation aa57dabe-d4a1-4abf-9dbc-0aa92135ffcd · inbound
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da17b1c-2b8b-44f0-98a6-3cd8ef72f4e1 · inbound
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4611edf2-8eb8-40cc-b249-c6b37710c0e2 · inbound
Emu3.5: Native Multimodal Models are World Learners FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d670af6b-6e6d-4657-919d-66937fa6d802 · inbound
Attention Grounded Enhancement for Visual Document Retrieval FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5286bb90-ce9e-4975-8e5e-ddd122d783b9 · inbound
POMA-3D: The Point Map Way to 3D Scene Understanding FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c175fb72-8b60-4568-847c-979d9611ee67 · inbound
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7373a016-8256-41ca-82f0-b07f3affd1e9 · inbound
RGB-Pointmap Pretraining for Unified 3D Scene Understanding FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6bac1155-41c3-4421-80a9-d973d6a95d1c · inbound
Mitigating Multimodal Hallucination via Phase-wise Self-reward FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f7720e3-99ec-4222-b6b1-9508dbd3ab9d · inbound
AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc4fc96e-58b6-4b8e-ad78-3a6af6b69bcd · inbound
IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d7bc418a-5b55-4a14-8883-44847e771448 · inbound
LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d593473c-b77d-4a0e-ac67-f2dd3a9717c1 · inbound
L2P: Unlocking Latent Potential for Pixel Generation FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b4d4030b-f343-42a7-947b-d622d4549ec1 · inbound
CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b30c20c8-4764-47fd-9777-738b298868d9 · inbound
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a44653e4-c4fa-4e70-b2de-5d8a0cb884d9 · inbound
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04bb5bc8-c02d-4539-9e83-84470608a121 · inbound
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85a38dce-0093-4b5c-9272-47eb371fb75e · inbound
InstanceControl: Controllable Complex Image Generation without Instance Labeling FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8addccce-036f-48d6-98bb-def375b1dbfc · inbound
DialogueVPR: Towards Conversational Visual Place Recognition FG-CLIP: Fine-Grained Visual and Textual Alignment
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.