Pith. sign in

Paper Citation Record · LEDGER

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules

As of 19 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2412.18224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18224 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:59:40.093942Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a12ca233-a84d-43e0-8129-cf4e0d1198dc · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.940271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.940271Z digest=sha256:a93058e54a4ab136a0a46480068ee6f9a2cd60c710488de225e4fd1a9ddab3ea

Observation 21a104ab-71e9-4c69-ae84-22df2ea02b71 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.947577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.947577Z digest=sha256:75e486c1983183f5e63595dd3c751957af02aa68c3368ee4f3dad1e88e9e36b3

Observation 0dd6296c-12b3-4b62-937a-93348021ee57 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.953253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.953253Z digest=sha256:02b0aaa440c0bdce8903804882af3e75789167af2b3e1677756491c0c4fd9957

Observation c7c6ef93-46be-4db1-8fc0-f0134ab71dee · outbound

This paper cites an unresolved cited work.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.959509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.959509Z digest=sha256:0902470e646b4e1d145a09e8729b3e3e3eeccf3daa29dd6092e34c5b2c19a41d

Observation f8056d1f-9c45-4b18-8b8b-dd4881d7635d · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.965663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.965663Z digest=sha256:6ec8f506cdb5ee464a594cbd6896b8409352ee6de51af47a7ac9e05134f7c5e5

Observation d7b6a9f6-91a0-44f6-bffe-1aebd65ef34c · outbound

This paper cites Segment Anything.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Segment Anything

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.971453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.971453Z digest=sha256:a12be322458bfbff1319ce0e2d7c5d9069cd7ee1441ac843ee2f8c3db72ed67b

Observation eb50f5d3-c5eb-4a1f-94f5-7c78f3f039c3 · outbound

This paper cites A.; et al.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules A.; et al

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.978402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.978402Z digest=sha256:dba5556562f68d0b117c1dd08854536cdfeec79f26e531a98c1a60b875a5eda8

Observation ac5a35c9-b026-4bf9-b7f4-c0acf514aa0f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.984505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.984505Z digest=sha256:58ff1d610584418fb47345ca121a24e7bc9a63409bef5ea2af4b27d25ed26d96

Observation 537fc083-8379-4b40-998a-348decc9ae71 · outbound

This paper cites an unresolved cited work.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.991265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.991265Z digest=sha256:fd0d56a2e952c79595622bfc5812879e1429b7e550dcb52ac4f440015435d1b1

Observation e4212fd5-fa7b-476d-889f-a33958c21491 · outbound

This paper cites Visual Spatial Reasoning.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Visual Spatial Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:39.999600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:39.999600Z digest=sha256:e31423453cccf4aa259881a097ab3bc0186639a8792a47bc17f6b8b19904d8ae

Observation e26cc8a8-95f2-4d03-a0c9-b7f95cca870f · outbound

This paper cites an unresolved cited work.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.006951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.006951Z digest=sha256:d44714ef3e8610cbdd1c8c90bdd5dbb287fa20d21976ff87f0dc20161ddbefb2

Observation 5c79fe31-5031-4125-9be9-95366177cbcf · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Improved Baselines with Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.016681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.016681Z digest=sha256:a70c231146aec71f0a476a095ef36779843c46bd8f3a7d3c329dd0067efa3d46

Observation b1a014e2-0da5-473c-9d64-f15ea89af49e · outbound

This paper cites Visual Instruction Tuning.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Visual Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.022890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.022890Z digest=sha256:2f04f2cc160fbc8097909120087ed15ac4ce615252b358666a852b5ee13c090d

Observation 5509f950-34aa-4863-bea6-4a349e03f997 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules A Survey on Hallucination in Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.029055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.029055Z digest=sha256:c5c1e01c567f8b89bafb6d886e0d715bde9d4cdb766b6d7041f44db32c7e791e

Observation 95b93937-8e49-4e0b-8f55-c452f26a8325 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules MMBench: Is Your Multi-modal Model an All-around Player?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.035011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.035011Z digest=sha256:6443d5b0a5a16b9113d88423459043dc10a1ee3de0e31f8846365276dbc4cbe9

Observation b2d656d4-f509-4e4b-a15d-3ca5d9e54b11 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.040560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.040560Z digest=sha256:a8acca0ce6e220e208a8a0c865e0d10eb522a53e9cfc7d8d986abd7c04fedeec

Observation fd130baf-7745-49e6-a1ce-501512b0bbb7 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules DINOv2: Learning Robust Visual Features without Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.046035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.046035Z digest=sha256:9795bd638e50db4fa7b4cecf8574291e4c87a7c5f47817b0904f17088ce53765

Observation fae0410b-d515-45e9-8e77-1d61e6915e7d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.052543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.052543Z digest=sha256:51417dd4c7efaf980e91edeafb1d36b675b5c85486d28825a612a6e945920392

Observation 50f97423-349f-4132-a89d-00a828ac26c2 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Learning Transferable Visual Models From Natural Language Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.058409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.058409Z digest=sha256:f08f8dcca34fc3de3bff7fc6be32638cf8ba69df01cbe8e99b05f5197509d51f

Observation 6034182a-5eaf-40a7-96ce-3fb835466e92 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.064297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.064297Z digest=sha256:b57a86e4c2996a02f308de613eafe796f62b1204cc7232d957f47b20562cdf68

Observation 525aad55-e554-48e1-abdd-d2b9dd3f2352 · outbound

This paper cites an unresolved cited work.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:59:40.507251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T04:59:40.070841Z digest=sha256:bc279848617ac47c8751de59af95a73da33c8ae932c31568681ce9c38e33f0d1

Observation 2aadb8f1-789f-420b-9028-e3c0f3345b5e · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules Sigmoid Loss for Language Image Pre-Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.075766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.075766Z digest=sha256:7f90400281478df01a02d6d503d91da5906382a4b5035ead78cbd782e4e516d7

Observation cbaed8c9-f12a-41a4-8ab7-b81e087a8482 · outbound

This paper cites MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.081757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.081757Z digest=sha256:532adaa7b160266d67cbf1ec8116647ab3d27ce077ea6336ca6f39b4e851ecb0

Observation d8be22c0-5c69-4235-8daa-630f86303e3c · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules , " * write output.state after.block = add.period write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.087657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.087657Z digest=sha256:f6d7e4bdbbfd7df907a72f9ab32178893fedc4a19aae563ccfaae0efa7f55612

Observation 028e6688-0528-4ea9-8c96-c193feeb210b · outbound

This paper cites write newline.

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules write newline

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:59:40.093942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:59:40.093942Z digest=sha256:256ef4c9c9aaf865d858b386556204891da91b6c30d738261c0db75feb20e160

Pith citing papers

No inbound Pith citation observations are available.