Pith. sign in

Paper Citation Record · LEDGER

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.15144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15144 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:38:24.547522Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9600f08-70cc-46cc-8d16-9e5bd6b977ae · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.015154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.015154Z digest=sha256:70242600a6b0536fafbe638381231ffef934f36f8d0c2bff5fc1d8668205f03c

Observation 0a930986-248c-450f-9d8e-60d6af521c45 · outbound

This paper cites Alayrac, J.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Alayrac, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.890814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.020398Z digest=sha256:95c5082632979bba09043e089a1f2bdced16759cbf1591ce36de92fba238e04c

Observation fe7bcbf2-fa40-4db2-a84a-a1d8bbfed217 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.025536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.025536Z digest=sha256:365176861edd7302a4fa3d5d5d6095c9d3448769435d8894f4700e14cf880ef6

Observation 940a74c5-8ba4-48df-a740-0a32275092f2 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.045964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.045964Z digest=sha256:708e254aa629336838a480bb736395ccbdfc10c2a0949968950290e59ea7f555

Observation 10576d6b-f8c2-49ee-b744-fed08389c296 · outbound

This paper cites Carion, F.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Carion, F

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.875882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.156760Z digest=sha256:2c0b7e6822f4793e1a5ecd5526b782121cbfd48f733b2edde870d30327b38760

Observation 3c1a8817-28ba-4336-8eb7-c38a2dcd6382 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.861327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.193224Z digest=sha256:ab753b8285fdb29f728cbbd74d9c75ac45e3fe3d8916c4332beab1d605deedd2

Observation 4dceec40-7eab-41d6-9659-523db1b678d9 · outbound

This paper cites The Llama 3 Herd of Models.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.199332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.199332Z digest=sha256:9d42498e92308bad6eb2350211fb056f4a2b411e9ff744fd2b8c357bef6814fd

Observation 238f5e4a-4b42-4556-a9d0-95e355648d00 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.654140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.204219Z digest=sha256:be3c6b9fe05f0ce07728b3c3fe0211b7602c7e0cb09faf1c14100bc450c2653b

Observation 6b657000-0ba8-412d-aa29-55ab72f5b5b0 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.639103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.209906Z digest=sha256:236071cd0967a749d0a6793a6659c1ef1b811c8d505993742846bb7da199fdeb

Observation bec82237-c7d1-4272-81e2-49e398be6a08 · outbound

This paper cites Kv and A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Kv and A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.623751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.214634Z digest=sha256:3c3eae1867e1a88191a048eb571a3b072cbc3daa79dd82ef546ce46e92cc9f11

Observation 0977a93c-da7a-4db5-8d65-abb39d5a0576 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.609875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.219493Z digest=sha256:1a92470882cc42902bf867b62b51904b98853729c2468241464946e2520025ab

Observation 5d547e25-110d-46d1-9a00-7089337b122f · outbound

This paper cites Minervini, A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Minervini, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.420462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.223550Z digest=sha256:acbc66e3a9c403a7f08206406527f1bd28c47e95baf6f49506e4324e090ccb62

Observation b1114157-1104-4bb7-a156-a3ff464717d8 · outbound

This paper cites Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:38:24.771826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.228150Z digest=sha256:08b6931a38d35eaf94a597000ec2672c46e527b57ba476e1fc897e3a8cf94767

Observation d6e0c9c3-502f-4b1d-8630-292561fcc54d · outbound

This paper cites Radford, J.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Radford, J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.389601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.233127Z digest=sha256:3d8b24e17f95cf28df2178fe156dbf5b752c7fc22ee75afac3ba0806172547ed

Observation 4492f9b0-c03f-4c34-9738-dd928133df44 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.238672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.238672Z digest=sha256:af575fac88af924b3b7156771ae239ecf9231fe61a50396bbc8e81cfb878b0bc

Observation 235333c6-971a-48ab-9389-ae5ccdd6eb78 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.374655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.293516Z digest=sha256:25c468ff887dc3a1732d6858ad30380543f5d2005ed08766204a9c763fe2b5d4

Observation c63b0f55-e4ae-4d82-a09a-99c4ce743a26 · outbound

This paper cites Salewski, A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Salewski, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.358212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.411039Z digest=sha256:c805f85443c9ea21fbb598f339f58c08a2f097d5704eb4da007b8543b7861cfa

Observation 96a05e5f-1e9b-4574-be11-c733b22f262a · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.176242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.481439Z digest=sha256:32d1ab0cc101f1d688845c1761386ba1d72841b0695db254f10482479c8d50c2

Observation 6eef9f09-924d-460d-8b05-be5df45f10d9 · outbound

This paper cites FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.505902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.505902Z digest=sha256:df88ecb2e628db1d30d7cd354e968e94af4af8967f19499498869809ba95cec2

Observation e1cdf967-880c-4933-8791-c87a16da6ac7 · outbound

This paper cites Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.510393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.510393Z digest=sha256:b0348b0062e18bc6032f0c901063ad34342cc160467a5885bc32a9b75eb50e62

Observation d08b5476-9cec-49ff-883a-52c4d94ca8a0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.514755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.514755Z digest=sha256:7eb5eed7249d7244b4963d0fe138a32c3bd0b4d9d1cb7ac863e6c3c21c8ac5ac

Observation 0cb8647d-f12e-44d0-bd35-61754c089210 · outbound

This paper cites Qwen2 Technical Report.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen2 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.519030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.519030Z digest=sha256:8b12905d89c879565b41a4f54d1795483e1764052b3953d8281431fd5bc7e850

Observation a700a717-608f-4fbb-b18d-4bd2b1c7d44b · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.523145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.523145Z digest=sha256:87c02dbd13f973fbeef4fe56ece2d2ac990e9ab58dd511e6ab1e6a7dcbe8c24f

Observation 9eea5803-ea97-434c-ace7-79e3569581ad · outbound

This paper cites Specifically, the shape limit is increased to 5–6 shapes per image, and the occlusion limit is also raised to 5–6 shapes.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Specifically, the shape limit is increased to 5–6 shapes per image, and the occlusion limit is also raised to 5–6 shapes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.121899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.528507Z digest=sha256:ac25de676548eb42c93d0c6eb8124d224ec50de5a3a6eaaefdcf94c24f3ab335

Observation 06aa8683-e30e-4912-a802-ed735ad8f655 · outbound

This paper cites This setup tests the model’s ability to detect and attribute more shapes in configurations that were not present in the training data.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models This setup tests the model’s ability to detect and attribute more shapes in configurations that were not present in the training data

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.106796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.533223Z digest=sha256:c52f242640c2da5349c3576faa37bc460f24f9fd5c664e75d0e639d8cf9729ae

Observation 628f8bf5-7b8d-4947-b61f-b63869b57ebf · outbound

This paper cites This scenario assesses the model’s ability to accurately detect and attribute shapes under previously unseen levels of occlusion.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models This scenario assesses the model’s ability to accurately detect and attribute shapes under previously unseen levels of occlusion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.038081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.538152Z digest=sha256:6e1c3f3c18c8fcf57a959f6bd568ebb30a31af18bde29efad4d583285407dda5

Observation c0a8e2e7-c228-495c-b296-9ff3fdc85bc8 · outbound

This paper cites The ability to generalize to these out-of-domain rotations Model OD Comp.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models The ability to generalize to these out-of-domain rotations Model OD Comp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:24.879894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.543002Z digest=sha256:b86ce9a26d4627c7230bf9dc5f106fa178238d0e298bc1e231a41cd8baa106b1

Observation 1387a2fc-f641-4791-98f9-ac2311053ecc · outbound

This paper cites Sentence Format Tuple Format MiniCPM-V2.5 Figure S4. Sentence Vs Tuple Output Comparison for MiniCPM-V2.5 Models for Validation Dataset.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Sentence Format Tuple Format MiniCPM-V2.5 Figure S4. Sentence Vs Tuple Output Comparison for MiniCPM-V2.5 Models for Validation Dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:24.864409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:38:24.547522Z digest=sha256:81060694ebc83c286e03f46882aaa4d28b29893299c891c4a5c0d6f45c4ef368

Pith citing papers

No inbound Pith citation observations are available.