Pith. sign in

Paper Citation Record · LEDGER

Optimizing Vision-Language Interactions Through Decoder-Only Models

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.10758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10758 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:07.654957Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de7930ef-f9ba-4cdf-8795-e67183fa8da3 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.405342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.405342Z digest=sha256:0783450406706fc62773cefbe1bbc8d47123673cbfc1a84482b4233d12edae78

Observation 0f3c027a-23f5-49e5-bf7b-c9340a66c595 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.418467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.418467Z digest=sha256:e66899c95c8aa3dda5eac573d94c9c6d81758e305a111a902a03cc27d55e80d9

Observation 29fe954b-93b0-43eb-ba0b-5604e6484cf9 · outbound

This paper cites Visual in-context le arning for large vision-language models,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Visual in-context le arning for large vision-language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.424605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.424605Z digest=sha256:d70646201dce3a0b325a999fcbdb9a8b3a87da6969df47250f917fe2321ebfce

Observation 7c6a1830-2b1b-448d-9a27-2f52f892de80 · outbound

This paper cites Claret: P re-training a correlation-aware context-to-event transformer for eve nt-centric gener- ation and classification,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Claret: P re-training a correlation-aware context-to-event transformer for eve nt-centric gener- ation and classification,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:08.041191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:41:07.434018Z digest=sha256:58fe086e00cd5fc475e27e7db331412abccf6a4cbc398ff717bdc8b7eb9c7e8a

Observation 884311d6-0dfb-43ad-acea-83f94f7efda6 · outbound

This paper cites Eventber t: A pre- trained model for event correlation reasoning,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Eventber t: A pre- trained model for event correlation reasoning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.439801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.439801Z digest=sha256:4a15ccb4ed23840961cb410ddae8d52c754eae8cc3ef21394f7deea55c8cb6e2

Observation 0f360099-d185-4b02-80c6-1149ae41d554 · outbound

This paper cites Visionllm: Large language mode l is also an open-ended decoder for vision-centric tasks,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Visionllm: Large language mode l is also an open-ended decoder for vision-centric tasks,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:08.006708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:41:07.445111Z digest=sha256:c769f6ad64ca3231383eef91ce1a0c1b5f86cb4140afa15a47069d02af5e6810

Observation 8c35f26c-cc7b-4475-89e5-dd16497b03d6 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Optimizing Vision-Language Interactions Through Decoder-Only Models VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.450255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.450255Z digest=sha256:7ffb9155fe114e42e56c2bdf4d2901c7d1d2dbde220bfc6fbb6d1dbc1c79aefc

Observation d2fb4fb8-c10f-4b83-9891-1ae0033fe3ed · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.456101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.456101Z digest=sha256:3c8cea0ae73f69ae003f9a2814bb838db117fd38c709de81734cab2a9139fde1

Observation dffc29d3-fe2f-412e-94bb-ab98399cc140 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

Optimizing Vision-Language Interactions Through Decoder-Only Models Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.461270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.461270Z digest=sha256:89547d768ae66e668b8765a0ef015fb392ba7585fbde3bfbfad9177b3f299a4d

Observation efdb86d5-6a57-4f20-8773-6fa05cfdb689 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Optimizing Vision-Language Interactions Through Decoder-Only Models Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.467266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.467266Z digest=sha256:e09056234e32cfbc233f5a169be1b6c097a5cc0d04602843c639fa634e1ffb37

Observation 69dd160d-b850-46cc-9e8e-b2abe60ea223 · outbound

This paper cites Triple sequence generati ve adversarial nets for unsupervised image captioning,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Triple sequence generati ve adversarial nets for unsupervised image captioning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.473518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.473518Z digest=sha256:c080f146cb3fe2bf214dee7bdfe7c58bcc322be959bb4d8e9076ad4bdc0eca77

Observation 79da6fd4-c041-4c53-b86b-d7ba60b2fd29 · outbound

This paper cites Sketch storytelling,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Sketch storytelling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.605049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.605049Z digest=sha256:ce86128b000063426773bae4c91a7dccd14c457add4fb3172da396185d8e62c8

Observation 08ee393c-8162-45b1-a029-1bcd41085a7a · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Optimizing Vision-Language Interactions Through Decoder-Only Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.611175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.611175Z digest=sha256:3e2097a820b8aebb8e5109681416e64102d8f8eb82dc2ad2dd41bb8c10e4a715

Observation 6d269bcc-721f-4cba-ab9b-9803bb332f29 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Optimizing Vision-Language Interactions Through Decoder-Only Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.634043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.634043Z digest=sha256:56197c6b1dd26828194b20dfc4660c6539b6c473eed1127aa41067584bf5673f

Observation 47267141-3a43-4ce5-9677-5de5bbb57551 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Optimizing Vision-Language Interactions Through Decoder-Only Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.628483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.628483Z digest=sha256:f288f6cdd7ec552a89a593a168b493d2e0a97448bb9f95bade2b7a1eef22c6f3

Observation 2e08b26e-817c-4f41-9012-f37a34ff83ae · outbound

This paper cites Multimodal event transformer for i mage-guided story ending generation,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Multimodal event transformer for i mage-guided story ending generation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.645686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.645686Z digest=sha256:e350ab723b8d235fd066d6eba112aa39bdbf5a8b3d9d0fa01f2be524d3913d3c

Observation 290f8d76-452e-4d2d-bfc2-4da99cfc6534 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

Optimizing Vision-Language Interactions Through Decoder-Only Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.639261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.639261Z digest=sha256:f52c592eaac4ad5521b4aa73a3e80733081b58e1667975a19372266d6b86c3d2

Observation aa2f60fb-8b6b-4693-8d29-832a263aaa79 · outbound

This paper cites Vilt: Vision-and-language t ransformer without convolution or region supervision,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Vilt: Vision-and-language t ransformer without convolution or region supervision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:07.935301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:41:07.654957Z digest=sha256:48a5c28a016b664d914ad7dd90971def825277f87ee85f0306a5cd7f9d67a423

Observation 5b40c818-fa1b-4aaa-a51c-8228ed812c57 · outbound

This paper cites Style-aware contrastive learning for multi-styl e image caption- ing,.

Optimizing Vision-Language Interactions Through Decoder-Only Models Style-aware contrastive learning for multi-styl e image caption- ing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:41:07.953949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:41:07.650306Z digest=sha256:a265cd77b46f10634305185b8928deda61f5c8a2507e899e591b48f234d7b5a3

Observation d9efcc2f-8978-4d23-97a8-97fad21a7f40 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.412311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.412311Z digest=sha256:4d5d21e283b44e06dede65cc745bc8b106711ba97c8aa195850514f8f8c9635c

Pith citing papers

No inbound Pith citation observations are available.