Pith. sign in

Paper Citation Record · LEDGER

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.15144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15144 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:38:24.547522Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9600f08-70cc-46cc-8d16-9e5bd6b977ae · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.015154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.015154Z digest=sha256:70242600a6b0536fafbe638381231ffef934f36f8d0c2bff5fc1d8668205f03c

Observation 0a930986-248c-450f-9d8e-60d6af521c45 · outbound

This paper cites Alayrac, J.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Alayrac, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.890814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.020398Z digest=sha256:3a10e04619fe1268ac8ae27825441415b7b8b8955c4bac37a65cbb74aa4ea256

Observation fe7bcbf2-fa40-4db2-a84a-a1d8bbfed217 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.025536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.025536Z digest=sha256:365176861edd7302a4fa3d5d5d6095c9d3448769435d8894f4700e14cf880ef6

Observation 940a74c5-8ba4-48df-a740-0a32275092f2 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.045964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.045964Z digest=sha256:708e254aa629336838a480bb736395ccbdfc10c2a0949968950290e59ea7f555

Observation 10576d6b-f8c2-49ee-b744-fed08389c296 · outbound

This paper cites Carion, F.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Carion, F

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.875882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.156760Z digest=sha256:b4dceb4baec472a6073f9675167e0b73380bed6fca38c0c72f03f1748cc804ea

Observation 3c1a8817-28ba-4336-8eb7-c38a2dcd6382 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.861327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.193224Z digest=sha256:a4635358350f0498fcc69dd7b4dd78291cbebead0fd7384441a4edeb67fcfec1

Observation 4dceec40-7eab-41d6-9659-523db1b678d9 · outbound

This paper cites The Llama 3 Herd of Models.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.199332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.199332Z digest=sha256:9d42498e92308bad6eb2350211fb056f4a2b411e9ff744fd2b8c357bef6814fd

Observation 238f5e4a-4b42-4556-a9d0-95e355648d00 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.654140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.204219Z digest=sha256:49b175c95c7ecf834ea348282a5c640552b77a1c3cf81b2a65db8f6eb7781b81

Observation 6b657000-0ba8-412d-aa29-55ab72f5b5b0 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.639103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.209906Z digest=sha256:7f7a88293ac0e9e43414aea8fe7d05a736c09f6e2fc4a86be816a71e1eba0ec2

Observation bec82237-c7d1-4272-81e2-49e398be6a08 · outbound

This paper cites Kv and A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Kv and A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.623751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.214634Z digest=sha256:c86ee881fcc5f0f00394c85cc3c1afeee66560d996fcc0711c3d5fdf8b8bf752

Observation 0977a93c-da7a-4db5-8d65-abb39d5a0576 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.609875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.219493Z digest=sha256:a67347f08161434616b08486586e8f06f20bec9a5f2b0ce12e081f839f03bf36

Observation 5d547e25-110d-46d1-9a00-7089337b122f · outbound

This paper cites Minervini, A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Minervini, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.420462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.223550Z digest=sha256:7de2c4d13e2cc567c1f2e300a035f201fa8f4b55d71cc3ce767122d888bbecc5

Observation b1114157-1104-4bb7-a156-a3ff464717d8 · outbound

This paper cites Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:38:24.771826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.228150Z digest=sha256:39f3de32e6d988f85b95020ddaf9c9d04f333e1f877e5c94ae828fbda32e1f3f

Observation d6e0c9c3-502f-4b1d-8630-292561fcc54d · outbound

This paper cites Radford, J.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Radford, J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.389601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.233127Z digest=sha256:27616a6245b72154983ac05407e8cda4552fc6398dca8ad67887d9d99723e769

Observation 4492f9b0-c03f-4c34-9738-dd928133df44 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.238672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.238672Z digest=sha256:af575fac88af924b3b7156771ae239ecf9231fe61a50396bbc8e81cfb878b0bc

Observation 235333c6-971a-48ab-9389-ae5ccdd6eb78 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.374655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.293516Z digest=sha256:74788eb08bbba3c45cdaf1ff5078d3e23dc6bf447f132798ccd49461c0956d6c

Observation c63b0f55-e4ae-4d82-a09a-99c4ce743a26 · outbound

This paper cites Salewski, A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Salewski, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.358212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.411039Z digest=sha256:720641afecb359a7879aa887afe4000841e1b62326bd16e2e034a12e0cae3780

Observation 96a05e5f-1e9b-4574-be11-c733b22f262a · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.176242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.481439Z digest=sha256:ccd4c968126332551244ea2b2f94796a0a6215ff87e48ff3a4157a958424abb6

Observation 6eef9f09-924d-460d-8b05-be5df45f10d9 · outbound

This paper cites FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.505902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.505902Z digest=sha256:62341ffe90b6178f2a489a0b8a82d52eb09c4364177cbaf053722cc100f1f10c

Observation e1cdf967-880c-4933-8791-c87a16da6ac7 · outbound

This paper cites Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.510393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.510393Z digest=sha256:b0348b0062e18bc6032f0c901063ad34342cc160467a5885bc32a9b75eb50e62

Observation d08b5476-9cec-49ff-883a-52c4d94ca8a0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.514755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.514755Z digest=sha256:7eb5eed7249d7244b4963d0fe138a32c3bd0b4d9d1cb7ac863e6c3c21c8ac5ac

Observation 0cb8647d-f12e-44d0-bd35-61754c089210 · outbound

This paper cites Qwen2 Technical Report.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen2 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.519030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.519030Z digest=sha256:8b12905d89c879565b41a4f54d1795483e1764052b3953d8281431fd5bc7e850

Observation a700a717-608f-4fbb-b18d-4bd2b1c7d44b · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.523145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.523145Z digest=sha256:87c02dbd13f973fbeef4fe56ece2d2ac990e9ab58dd511e6ab1e6a7dcbe8c24f

Observation 9eea5803-ea97-434c-ace7-79e3569581ad · outbound

This paper cites Specifically, the shape limit is increased to 5–6 shapes per image, and the occlusion limit is also raised to 5–6 shapes.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Specifically, the shape limit is increased to 5–6 shapes per image, and the occlusion limit is also raised to 5–6 shapes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.121899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.528507Z digest=sha256:7f6d435e52962d33976e83c8adb93b7779114f610a7a9609001f20f1b267bf69

Observation 06aa8683-e30e-4912-a802-ed735ad8f655 · outbound

This paper cites This setup tests the model’s ability to detect and attribute more shapes in configurations that were not present in the training data.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models This setup tests the model’s ability to detect and attribute more shapes in configurations that were not present in the training data

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.106796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.533223Z digest=sha256:54f117c81030628dfc5d11f38983884b286d6d606c056dba92ed9b08effea528

Observation 628f8bf5-7b8d-4947-b61f-b63869b57ebf · outbound

This paper cites This scenario assesses the model’s ability to accurately detect and attribute shapes under previously unseen levels of occlusion.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models This scenario assesses the model’s ability to accurately detect and attribute shapes under previously unseen levels of occlusion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.038081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.538152Z digest=sha256:0fca0e367b08956d071ac517449a8be261ecaa6da113fe39b4e39afa29e833c5

Observation c0a8e2e7-c228-495c-b296-9ff3fdc85bc8 · outbound

This paper cites The ability to generalize to these out-of-domain rotations Model OD Comp.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models The ability to generalize to these out-of-domain rotations Model OD Comp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:24.879894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.543002Z digest=sha256:c71144d12a39039510f9d1c2a409a77d0214f1931a877126989363126f8dfd50

Observation 1387a2fc-f641-4791-98f9-ac2311053ecc · outbound

This paper cites Sentence Format Tuple Format MiniCPM-V2.5 Figure S4. Sentence Vs Tuple Output Comparison for MiniCPM-V2.5 Models for Validation Dataset.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Sentence Format Tuple Format MiniCPM-V2.5 Figure S4. Sentence Vs Tuple Output Comparison for MiniCPM-V2.5 Models for Validation Dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:24.864409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:38:24.547522Z digest=sha256:bd35d6645fbaf37b67a7700b48e9e539fbc3ac758fb07cbf0a993d51fb67ba94

Pith citing papers

No inbound Pith citation observations are available.