Pith. sign in

Paper Citation Record · LEDGER

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

As of 12 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 3 inbound Pith citation observations for arXiv:2412.12940.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12940 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:38:48.718208Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.415452Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f692c8d6-53e4-471b-b6b3-4483b6121e35 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.419921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.419921Z digest=sha256:8163eb8e995dfbc6f124a9eaf6803a8c5f44743fa039f065bd1995b417887a79

Observation 8abb1fdc-6ca9-497d-817c-3b9a83c9aa5a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.424361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.424361Z digest=sha256:192322c906750530cf236ae8cd478691cca4cb48a6b54f1f5b7df9b76f3b80e8

Observation 4ff5f3a8-6e18-4471-966e-a5f7ff2fc138 · outbound

This paper cites Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.427892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.427892Z digest=sha256:e0d38e6460a7295a4e875b7e4d217a01194b94cb855cdd1455d05d2761f12cbe

Observation 9f3ee29b-54d9-4056-b79e-1a3b50a14088 · outbound

This paper cites Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.431347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.431347Z digest=sha256:b8e06f46047eb49db593cf99d89c44aa90ab03c98199666094b5fa41ed8be8f9

Observation c50e91d1-5df6-41cf-9dde-fe077f4d97f3 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:49.123955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:38:48.435663Z digest=sha256:a1d2b2f96708e2a14bb797aca8aa4125379e0605c5c6940ef25d4a1c0a161bcf

Observation d2524bb3-5842-4a0c-8db3-b5e20ff37f14 · outbound

This paper cites Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.439701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.439701Z digest=sha256:1ae107b3919566ae9f17c38d54ad6c5c940d006dbba0c06b820d1fa4154d52e6

Observation c1afdd27-89e9-4517-965a-1939588c699a · outbound

This paper cites E.; et al.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training E.; et al

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.443399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.443399Z digest=sha256:b7c657125ae7035c41a4b9f6362fb556ad33b0eb04fd4a6db30372addfd77c82

Observation 50fe8086-f1be-40f8-9841-b3716de52634 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.460520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.460520Z digest=sha256:0a89cda3d8692ee3af646a3de4ca1185b72b28d189a5a4de274e8f688c8d6304

Observation 32c308d1-7ce9-425d-b961-4344e7ebf56a · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.505212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.505212Z digest=sha256:77af982ce6409720396fa0fda7041a9a7600455fdc88d21a5be73c416b01f7a6

Observation 77d4a999-ef3f-4ef2-92e7-bbbead272f72 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:49.107337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:38:48.546688Z digest=sha256:deb1e9b8c69771e2367582575d8ef5e7ee1b400ec74c7ad881ff2215c869779b

Observation 76ed9def-f302-4623-91f1-2ae4399f6f76 · outbound

This paper cites GPT-4o System Card.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.572034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.572034Z digest=sha256:f61dbafacf6e812c03a10931f94911bf083d371b8528ac2a09e93915d3f71585

Observation 49e3ff2c-b675-425c-9aeb-2650e2bad0d0 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.640228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.640228Z digest=sha256:b77b897d298b0fe653b50c2b75e08fd203d3311ac0090d9bc5ddc75dd6c1df2a

Observation f8183eb8-cd66-4691-96e3-ed86ff1e76a5 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.674294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.674294Z digest=sha256:132d3b58e052a13832934c95108efa7d85776be2ab06c1c5b31cfccf3ddb5dd0

Observation 28f271c5-3221-49f4-be25-7b2ad3227dd8 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.677915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.677915Z digest=sha256:c3025731fd378c11003ffe0e29bd384323f844d6bd3d18bed46c57f2f6b384a4

Observation a5fa7fda-a0b8-4da7-8e7c-458c2bb54727 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:49.083532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:38:48.681281Z digest=sha256:539a25119b2e05d89ae24bb58fc32001fcf672c57c6908410917ff7b9d52e2d9

Observation 8ccddd14-43e4-45fb-aa4c-81fa8f4657ae · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:49.072948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:38:48.684405Z digest=sha256:fc513291c4fa4715789b416525a1d6f7dea28913f5a94e2b6fd3bf40a740ae15

Observation 1acb917f-de01-4d4b-91b3-42fc795eb1af · outbound

This paper cites Carbon Emissions and Large Neural Network Training.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Carbon Emissions and Large Neural Network Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.688127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.688127Z digest=sha256:e2c1da695a895671c7fe2da67d906ac984244f06f9e5391c6799c6783bd39903

Observation 407c2d59-c536-4c40-a3ee-03e257cc2d8f · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.692182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.692182Z digest=sha256:62b6c7718f3b7a22050a89761d6404364d71763baebfb950cb74c7c2e75271e0

Observation 37f9900e-44ce-4094-870f-d90016f9b138 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.695364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.695364Z digest=sha256:cfe9e8c41d8e40aedda01fecd1e808222d94e9b169e1a6e1fbeb667c26eba01b

Observation c3a75ae1-5564-49f1-af93-d6791e11b48b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.699233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.699233Z digest=sha256:dfac7852cf739ae67b9d4ce19d1de190193255d1d4d565f1e683473f8ef9c5be

Observation a7ceefd9-4a8a-44e0-857e-af25f44d6e12 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:49.054240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:38:48.703428Z digest=sha256:dacbe59df294ab3a1b6ebfaf960f319d3edabaa35973e76e662af7f7e22a51c6

Observation e78a243b-1b4a-4bfd-ad87-5843e7feadc0 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.706644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.706644Z digest=sha256:dbf64a73de576dda086026752781853ebcf2e479a51d41196542fde2b17e52b3

Observation 1aa65b62-02f3-48c3-8fb4-67c2de24ef24 · outbound

This paper cites an unresolved cited work.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:48.941388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:38:48.710069Z digest=sha256:d2a7d2e5217c1f2aa018706921d0093f291d768342fe26ac3f076e843b36f4ed

Observation fe62e698-93f6-4ecc-91e9-179aa452b368 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training , " * write output.state after.block = add.period write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.714006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.714006Z digest=sha256:64ee78c6ec4cd66429c47a315e5768a316c69302ddfd55a4b03e865693b128ee

Observation 4eb4a3ed-ce2c-4ac9-963b-a573c338a0ca · outbound

This paper cites write newline.

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training write newline

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:48.718208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:48.718208Z digest=sha256:0444e75199abf604ab855d9b795f89cd0d3b5412b06a1f0f84fa8f13fe79fb42

Pith citing papers

Observation f4f61ac8-4882-48ed-8eae-3bf6c8df07b5 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:51:05.108698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:e8f9bdad7a13528b56d4601e8b9ce05a2431503faba945362e906ac0afaa9569

Observation 4591128f-01c6-41d4-91b5-7adcbe14be72 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:36:25.017520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:8a6a9300e9f084baee1148e4a0e0a039512fd4d561e29081856a132e680e6ef0

Observation 8786e51c-aef8-47c0-9dcb-f02a802a4832 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.416837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:1d107fc03333eae22881f14e7d3353e62f2c5234e129f36fb2b07e5dd6b79426