Pith. sign in

Paper Citation Record · LEDGER

Reconstructive Visual Instruction Tuning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2410.09575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.09575 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:32:46.962180Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:32.245115Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8805349b-7f10-43ce-95fe-b6f8bde3ef67 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Reconstructive Visual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.436682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:f6b5482575745d507179b2dfe2a2e19fd6eba514ef8108b4e0805c7958d09021

Observation 961c6b19-704e-4cf5-8b3d-43130886e1a4 · inbound

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver cites this paper.

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver Reconstructive Visual Instruction Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:32:46.962180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:32:46.962180Z digest=sha256:9cf4a6becdf0ee24f279053ef8393842e41b3fabfc6ad7a03c7f7e805cf0a7cb

Observation 88e3bb72-4031-4ae4-bb77-97095bef5fbe · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.181930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.181930Z digest=sha256:c1de9ca82744b8e1b3b1d132772247bb615116c8503c349bcae6ba968a1d43ca

Observation 9c34c403-4ceb-4df0-a2ec-b05a656ffa54 · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.467952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.467952Z digest=sha256:ab583311735797303d8f495a548a7c5ed654e408c2379dae8ee20a2ba39311c7

Observation a6ff7103-f3bd-4daa-a8d2-eb8018ed8ca8 · inbound

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving cites this paper.

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving Reconstructive Visual Instruction Tuning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:48:01.100732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T06:48:00.943591Z digest=sha256:5fbeff4e10f794467f18ec89a384af8f440d4de81677b45baa75507aa9c21b88

Observation 4f0c3670-1400-4d15-8359-c45d93d6fba8 · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning Reconstructive Visual Instruction Tuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:19.047232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:19.047232Z digest=sha256:3975e8a2cf35115d8bf3e01729334278c48b66e42c4c36e5467f8bdee55b54a5

Observation 727ea4d3-504e-4b96-b9c4-822755bbed99 · inbound

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation cites this paper.

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation Reconstructive Visual Instruction Tuning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:13.944749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:13.944749Z digest=sha256:0e347ec60b250ee10408073d8bf5df642767ded9006b10b16c17a347377f13e6

Observation f122d2ab-b75b-451d-9b42-eafd08f36b78 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models Reconstructive Visual Instruction Tuning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.735275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:66373b7b52b3888c07299f5fc6d772339a105b35d0a64645548dfa5dd9d4c1dd

Observation bacff6a8-81a5-4502-80e1-c01ddd75dceb · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Reconstructive Visual Instruction Tuning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.191326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:a87769474e7504d8ec5a7666b83b752ad31b84fe088fd046c7f7cf1295bdc17e

Observation 09bfac4b-1a20-4343-92f9-1c01c18a326f · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models Reconstructive Visual Instruction Tuning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.438803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:38:01.947806Z digest=sha256:3c512bfcca68cf922505b4c10098d4869281ea248b26c09e12e40cccfe62c6fc

Observation db4553d1-d21c-444d-9b2a-fa253baf28eb · inbound

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models cites this paper.

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models Reconstructive Visual Instruction Tuning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.162571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:42:46.422756Z digest=sha256:c28b5fc38b0db68ee09f0bfd05967273d72c6a405f32413a3f0cecebab1a7601

Observation c1c5f902-0a47-4c48-a955-f265a42a3164 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.237180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:b115bc98eeb8c4ad6004a77eae0210c196fab8307e21c47dce6559e8bb55f80a

Observation fb0388fb-fd3c-4570-b15b-345d8c0f86fc · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.342775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:e80db676b4b9b874040bef322c17f5f97418dc2fb3d5b926e5435a67929c15cb

Observation 6b3391d9-473d-4a90-a54f-0e402270f3bb · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Reconstructive Visual Instruction Tuning

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:32.246590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:e3b84be97e42659adad1fdebf007718100fc214d9163b874be53c3357aae8553