Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:38:48.718208Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 3 inbound Pith citation observations for arXiv:2412.12940.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:38:48.718208Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T21:28:58.415452Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f692c8d6-53e4-471b-b6b3-4483b6121e35 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abb1fdc-6ca9-497d-817c-3b9a83c9aa5a · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff5f3a8-6e18-4471-966e-a5f7ff2fc138 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3ee29b-54d9-4056-b79e-1a3b50a14088 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c50e91d1-5df6-41cf-9dde-fe077f4d97f3 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2524bb3-5842-4a0c-8db3-b5e20ff37f14 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1afdd27-89e9-4517-965a-1939588c699a · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training E.; et al
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50fe8086-f1be-40f8-9841-b3716de52634 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c308d1-7ce9-425d-b961-4344e7ebf56a · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training PaLM-E: An Embodied Multimodal Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d4a999-ef3f-4ef2-92e7-bbbead272f72 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 76ed9def-f302-4623-91f1-2ae4399f6f76 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training GPT-4o System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e3ff2c-b675-425c-9aeb-2650e2bad0d0 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8183eb8-cd66-4691-96e3-ed86ff1e76a5 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f271c5-3221-49f4-be25-7b2ad3227dd8 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5fa7fda-a0b8-4da7-8e7c-458c2bb54727 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ccddd14-43e4-45fb-aa4c-81fa8f4657ae · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1acb917f-de01-4d4b-91b3-42fc795eb1af · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Carbon Emissions and Large Neural Network Training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 407c2d59-c536-4c40-a3ee-03e257cc2d8f · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f9900e-44ce-4094-870f-d90016f9b138 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a75ae1-5564-49f1-af93-d6791e11b48b · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ceefd9-4a8a-44e0-857e-af25f44d6e12 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e78a243b-1b4a-4bfd-ad87-5843e7feadc0 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aa65b62-02f3-48c3-8fb4-67c2de24ef24 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe62e698-93f6-4ecc-91e9-179aa452b368 · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training , " * write output.state after.block = add.period write newline
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb4a3ed-ce2c-4ac9-963b-a573c338a0ca · outbound
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training write newline
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f61ac8-4882-48ed-8eae-3bf6c8df07b5 · inbound
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4591128f-01c6-41d4-91b5-7adcbe14be72 · inbound
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8786e51c-aef8-47c0-9dcb-f02a802a4832 · inbound
ESC: Emotional Self-Correction for Reliable Vision-Language Models Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.