Pith. sign in

Paper Citation Record · LEDGER

Can Visual Encoder Learn to See Arrows?

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2505.19944.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19944 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:20.058994Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39471ad9-2149-4212-ab77-4c84f03b3515 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Can Visual Encoder Learn to See Arrows? Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.749975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.749975Z digest=sha256:5479a7f69fb579876efd39795fd9dc4555778ca3f72e01a3b8b4f89144dcde5b

Observation 15425ba8-7583-45ec-bcb5-af119347e278 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Can Visual Encoder Learn to See Arrows? OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.838017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.838017Z digest=sha256:b976cb19dc8bb8a1b94c7c9930c4f9578890bbb67754f77690294573c865dff3

Observation 6840dfbf-15ea-4ce4-b255-7e5c5f94f55d · outbound

This paper cites Neural codes for image retrieval.

Can Visual Encoder Learn to See Arrows? Neural codes for image retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.472836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:17.881086Z digest=sha256:1bece83f0b6680ebf34c58e982118e59b6fa7953d524e22913c3e04f73f9130a

Observation 3c66b4bf-0e7e-434e-b2d1-9df26c205382 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Can Visual Encoder Learn to See Arrows? GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.952748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.952748Z digest=sha256:ce5080d0fb221458cfd266027b7071d5e4a25bef463bd806426f26d73f83c743

Observation a519e5e2-26d5-4070-a650-1918f41a2bd2 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InAdv.

Can Visual Encoder Learn to See Arrows? Are we on the right way for evaluating large vision-language models? InAdv

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.314526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.029280Z digest=sha256:19ec43c2eb15d4f4ae3b6547a978f5b4f770f748ca5199037ff21ef0ce22fd32

Observation 01660096-ceee-4595-967a-1b7a508dab39 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Can Visual Encoder Learn to See Arrows? PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.135269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.135269Z digest=sha256:0850f16a6c69e796517a9db9bbd2582d9d813d25b0f9691aaa5b64ff6f607a16

Observation cbf9bf9f-8a64-4cde-8b7d-f57ec038c2ad · outbound

This paper cites What you can cram into a single $&!#* vector: Probing sentence embeddings for lin- guistic properties.

Can Visual Encoder Learn to See Arrows? What you can cram into a single $&!#* vector: Probing sentence embeddings for lin- guistic properties

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.141131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.206189Z digest=sha256:ec0417fc92cd716c19f52de0c3f0bd6804aa4c5395c51d5fa16ecd2812b870f4

Observation c5f4f1d8-5e21-41dd-bd76-3b8db9d45745 · outbound

This paper cites an unresolved cited work.

Can Visual Encoder Learn to See Arrows? Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:22.993229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.275398Z digest=sha256:db93fa5663200ce789e4c1bf4b9446abf49eed66e7c60b5b819de10218bb7c35

Observation feca93f9-d9bf-46aa-bb58-2b31043ea144 · outbound

This paper cites an unresolved cited work.

Can Visual Encoder Learn to See Arrows? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:22.873637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.339017Z digest=sha256:473c8f2d460d4304256bfd239eef84c6224bd1c770146b554344cb5e0ca81236

Observation 10cd801b-d7d7-47e9-a18c-c49f6d9f1bd6 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Can Visual Encoder Learn to See Arrows? BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.386462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.386462Z digest=sha256:501b9097233e3df73842bc2e8ac9a4eedc49d47bd13cdc28cb411208113850a3

Observation 9256d60c-dabc-4af2-b317-1c2f7b9f2ee2 · outbound

This paper cites Visual Instruction Tuning.

Can Visual Encoder Learn to See Arrows? Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.468295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.468295Z digest=sha256:a4a9317f96f36e16486b2bf11ebf3b60e9744aec08c0c9934927f28d714a6288

Observation e21d1d39-45dd-4cbe-9757-dd4859b1091e · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge.https: / / llava - vl.

Can Visual Encoder Learn to See Arrows? LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge.https: / / llava - vl

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.708865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.580748Z digest=sha256:9cbce37134a3d1271ee41ecb46ac14553c5482a8727fbc6f9ae0a069e3d0dee6

Observation 3d4b7091-b4eb-404e-a5c5-cf13591b1663 · outbound

This paper cites MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Can Visual Encoder Learn to See Arrows? MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.554493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.662136Z digest=sha256:b4e6d4dae4fbdd43c5f247241984ee7f8a3ef01988db273eb8f46d2a6da18405

Observation 0ee13ad4-fb7e-4074-bc2f-54c22c545f54 · outbound

This paper cites CLIP ViT-B/32.https://huggingface.co/ openai/clip-vit-base-patch32, 2021.

Can Visual Encoder Learn to See Arrows? CLIP ViT-B/32.https://huggingface.co/ openai/clip-vit-base-patch32, 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.368834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.768103Z digest=sha256:8f444217b5d0da3dce78379d73a39ae8122fb61dc6581ddad6a3a938f4f0f973

Observation ef4f94de-b285-4dcf-b4fb-68e9e6189ad8 · outbound

This paper cites CLIP ViT-L/14-336.https://huggingface.

Can Visual Encoder Learn to See Arrows? CLIP ViT-L/14-336.https://huggingface

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.193856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.860984Z digest=sha256:d9e990b3a504ecc111a424ab5f011025db7c58ff00636badb1d503bad575175e

Observation f2afb6e4-0f69-402e-803c-94fd551c98a8 · outbound

This paper cites GPT-4 Technical Report.

Can Visual Encoder Learn to See Arrows? GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.043623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.043623Z digest=sha256:e5c79722af7b3859fed81ab5ee05294062ff271ddc819c3412de86539fa8e7b2

Observation 021ebe13-dda6-4a68-9d94-6be23c62d492 · outbound

This paper cites Hello, GPT-4o, 2023.

Can Visual Encoder Learn to See Arrows? Hello, GPT-4o, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.849237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:19.114429Z digest=sha256:5a40f5ee29019fbbc4e833e060046c220100bb369e7cbc6dfe7f04c7e47dfe4b

Observation 4a48a3ef-7915-4365-80ee-d887729d4ad4 · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019.

Can Visual Encoder Learn to See Arrows? Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.211866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.211866Z digest=sha256:7a7a14c84b505a28a710f77eef46ac94fb6c131b996feadb9ca97f45dba0bd72

Observation bae674e7-28f6-4f70-9d20-ea7ac8cf3d74 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Can Visual Encoder Learn to See Arrows? Learning transferable visual models from natural language supervi- sion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.321307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.321307Z digest=sha256:df400bc39c557274ff98aa07bf6d94b17c8d845b961a450828db984d78f18a49

Observation 583e82d8-3d67-4a2a-b75b-3dc7d83a6f03 · outbound

This paper cites Vision language models are blind.

Can Visual Encoder Learn to See Arrows? Vision language models are blind

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.645904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:19.426053Z digest=sha256:470600d11826f9461979ad386673a86e55c45490f91aff6526df8d26bd374817

Observation 269978b4-6d20-4519-b803-4644348d04fa · outbound

This paper cites CNN features off-the-shelf: an astound- ing baseline for recognition.

Can Visual Encoder Learn to See Arrows? CNN features off-the-shelf: an astound- ing baseline for recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.450556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:19.484458Z digest=sha256:7d776d081120fd64e8487f1aa57caf0a521990fdd3c6ee2a33a98b5bec137026

Observation 122c9e54-676f-46a2-b903-7e38d0211383 · outbound

This paper cites FlowVQA: Mapping multimodal logic in visual question an- swering with flowcharts.

Can Visual Encoder Learn to See Arrows? FlowVQA: Mapping multimodal logic in visual question an- swering with flowcharts

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.259613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:19.576412Z digest=sha256:d5021c3437b069e86aa06d858ae3d66d4bf63918af13c6cd938bb7a69dae090b

Observation 6b44838c-7224-42f5-87f6-6b742d9b635c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Can Visual Encoder Learn to See Arrows? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.645208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.645208Z digest=sha256:a9251dae517149c6b0db0532a1b061eb338cb0e0eb09a95ca4c723d4cc09470a

Observation da8f3efc-55c7-4383-8605-3a9269e26372 · outbound

This paper cites How well do vision models encode diagram attributes? Inthe ACL 2024 Student Research Workshop, 2024.

Can Visual Encoder Learn to See Arrows? How well do vision models encode diagram attributes? Inthe ACL 2024 Student Research Workshop, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.086930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:19.770636Z digest=sha256:44cd730106455944ab0fa2d3841765d9a08ea292c3ce1a0b5ecd8a9aa341ec28

Observation 4d5510b3-200f-4346-8ad0-b15a579ca76b · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert AGI.

Can Visual Encoder Learn to See Arrows? MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert AGI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:20.800916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:19.873890Z digest=sha256:a03fc43e5fe77b73d4b252423c13de856016178ec55bb15aebf9ded911ffa8e7

Observation 53cb6495-33d5-4ab9-892d-a7e523cbd428 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Can Visual Encoder Learn to See Arrows? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:20.608602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:19.966361Z digest=sha256:be0db56fdec048416dd9d278b2930547b8de1f885593d0354192123df4ce8020

Observation 96ce0fac-082c-43f0-8b31-9f7c8e30d11e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Can Visual Encoder Learn to See Arrows? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:20.058994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:20.058994Z digest=sha256:d8e7ea438ccf2a9fbe3132fdbdb61023af0a05caed19b54088ce96abe6900976

Observation c3a0f9ef-7804-437b-8a68-3ca0b5960f85 · outbound

This paper cites Accessed: 2025-04-12.

Can Visual Encoder Learn to See Arrows? Accessed: 2025-04-12

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.028089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:07:18.935665Z digest=sha256:8b0da4dcce1499625c3b7f5779132abf68bb84c5ee8addbc30d2f264109d7ba1

Pith citing papers

No inbound Pith citation observations are available.