Pith. sign in

Paper Citation Record · LEDGER

Can Visual Encoder Learn to See Arrows?

As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2505.19944.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19944 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:20.058994Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39471ad9-2149-4212-ab77-4c84f03b3515 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Can Visual Encoder Learn to See Arrows? Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.749975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.749975Z digest=sha256:9f37a419222b5aef366fdb9128ebf11cd8de826f27cf91390a2543c324d4ace4

Observation 15425ba8-7583-45ec-bcb5-af119347e278 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Can Visual Encoder Learn to See Arrows? OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.838017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.838017Z digest=sha256:b976cb19dc8bb8a1b94c7c9930c4f9578890bbb67754f77690294573c865dff3

Observation 6840dfbf-15ea-4ce4-b255-7e5c5f94f55d · outbound

This paper cites Neural codes for image retrieval.

Can Visual Encoder Learn to See Arrows? Neural codes for image retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.472836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:17.881086Z digest=sha256:72896df4b7aa113fae78ee0196d589994666bac176d13c4af98d765c5025705a

Observation 3c66b4bf-0e7e-434e-b2d1-9df26c205382 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Can Visual Encoder Learn to See Arrows? GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.952748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.952748Z digest=sha256:ce5080d0fb221458cfd266027b7071d5e4a25bef463bd806426f26d73f83c743

Observation a519e5e2-26d5-4070-a650-1918f41a2bd2 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InAdv.

Can Visual Encoder Learn to See Arrows? Are we on the right way for evaluating large vision-language models? InAdv

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.314526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.029280Z digest=sha256:a85dd057d13e6810df4f1e3ba7c01c9907a815fa704ce2e9445c78e84213582e

Observation 01660096-ceee-4595-967a-1b7a508dab39 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Can Visual Encoder Learn to See Arrows? PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.135269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.135269Z digest=sha256:0850f16a6c69e796517a9db9bbd2582d9d813d25b0f9691aaa5b64ff6f607a16

Observation cbf9bf9f-8a64-4cde-8b7d-f57ec038c2ad · outbound

This paper cites What you can cram into a single $&!#* vector: Probing sentence embeddings for lin- guistic properties.

Can Visual Encoder Learn to See Arrows? What you can cram into a single $&!#* vector: Probing sentence embeddings for lin- guistic properties

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.141131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.206189Z digest=sha256:efe3dbeedd30d222f7720f8c3b9414f712777a0a511112f48e1226386517588c

Observation c5f4f1d8-5e21-41dd-bd76-3b8db9d45745 · outbound

This paper cites an unresolved cited work.

Can Visual Encoder Learn to See Arrows? Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:22.993229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.275398Z digest=sha256:5eb8a58a5cbbf4c8dde1ff70839df3bd9937faae2804b78031a1f02dc97fa19f

Observation feca93f9-d9bf-46aa-bb58-2b31043ea144 · outbound

This paper cites an unresolved cited work.

Can Visual Encoder Learn to See Arrows? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:22.873637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.339017Z digest=sha256:a04bfbbc6e4746ed5633301d5995d80df34fad009579be050d321abb65fcbd33

Observation 10cd801b-d7d7-47e9-a18c-c49f6d9f1bd6 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Can Visual Encoder Learn to See Arrows? BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.386462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.386462Z digest=sha256:501b9097233e3df73842bc2e8ac9a4eedc49d47bd13cdc28cb411208113850a3

Observation 9256d60c-dabc-4af2-b317-1c2f7b9f2ee2 · outbound

This paper cites Visual Instruction Tuning.

Can Visual Encoder Learn to See Arrows? Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.468295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.468295Z digest=sha256:a4a9317f96f36e16486b2bf11ebf3b60e9744aec08c0c9934927f28d714a6288

Observation e21d1d39-45dd-4cbe-9757-dd4859b1091e · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge.https: / / llava - vl.

Can Visual Encoder Learn to See Arrows? LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge.https: / / llava - vl

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.708865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.580748Z digest=sha256:85f419456e90fd5de2fff46bed7b26f732295ae367877d17d1945c0671545cdc

Observation 3d4b7091-b4eb-404e-a5c5-cf13591b1663 · outbound

This paper cites MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Can Visual Encoder Learn to See Arrows? MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.554493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.662136Z digest=sha256:fea1acfe2e55598d5adbea49debb9b510d1b681037dc0250295d4484b1faabb8

Observation 0ee13ad4-fb7e-4074-bc2f-54c22c545f54 · outbound

This paper cites CLIP ViT-B/32.https://huggingface.co/ openai/clip-vit-base-patch32, 2021.

Can Visual Encoder Learn to See Arrows? CLIP ViT-B/32.https://huggingface.co/ openai/clip-vit-base-patch32, 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.368834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.768103Z digest=sha256:a5c61860a9edb3925551b83b9d507c3767d6db6eacac77af199ee098bd22b061

Observation ef4f94de-b285-4dcf-b4fb-68e9e6189ad8 · outbound

This paper cites CLIP ViT-L/14-336.https://huggingface.

Can Visual Encoder Learn to See Arrows? CLIP ViT-L/14-336.https://huggingface

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.193856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.860984Z digest=sha256:e8051b00eddaad6aa2ee025e74c62f300da74e2b1039d6196b44277050861c39

Observation f2afb6e4-0f69-402e-803c-94fd551c98a8 · outbound

This paper cites GPT-4 Technical Report.

Can Visual Encoder Learn to See Arrows? GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.043623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.043623Z digest=sha256:e5c79722af7b3859fed81ab5ee05294062ff271ddc819c3412de86539fa8e7b2

Observation 021ebe13-dda6-4a68-9d94-6be23c62d492 · outbound

This paper cites Hello, GPT-4o, 2023.

Can Visual Encoder Learn to See Arrows? Hello, GPT-4o, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.849237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:19.114429Z digest=sha256:534b361bc5fea738ffe41377eaaee02f20bb785b23c850e9f7501497ee59f92b

Observation 4a48a3ef-7915-4365-80ee-d887729d4ad4 · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019.

Can Visual Encoder Learn to See Arrows? Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.211866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.211866Z digest=sha256:7a7a14c84b505a28a710f77eef46ac94fb6c131b996feadb9ca97f45dba0bd72

Observation bae674e7-28f6-4f70-9d20-ea7ac8cf3d74 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Can Visual Encoder Learn to See Arrows? Learning transferable visual models from natural language supervi- sion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.321307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.321307Z digest=sha256:df400bc39c557274ff98aa07bf6d94b17c8d845b961a450828db984d78f18a49

Observation 583e82d8-3d67-4a2a-b75b-3dc7d83a6f03 · outbound

This paper cites Vision language models are blind.

Can Visual Encoder Learn to See Arrows? Vision language models are blind

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.645904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:19.426053Z digest=sha256:41055f437e4a9ec2594bfc870056de8202a5ee3edb232f3282880d3530546068

Observation 269978b4-6d20-4519-b803-4644348d04fa · outbound

This paper cites CNN features off-the-shelf: an astound- ing baseline for recognition.

Can Visual Encoder Learn to See Arrows? CNN features off-the-shelf: an astound- ing baseline for recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.450556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:19.484458Z digest=sha256:d6e9f3a032a7ec9bd49c4b4635f27ab3b7749622f7fbb477fe0ea7caf4010c00

Observation 122c9e54-676f-46a2-b903-7e38d0211383 · outbound

This paper cites FlowVQA: Mapping multimodal logic in visual question an- swering with flowcharts.

Can Visual Encoder Learn to See Arrows? FlowVQA: Mapping multimodal logic in visual question an- swering with flowcharts

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.259613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:19.576412Z digest=sha256:7963ea55b1b0348e82e2fc16b6c94ad4c1e1e70716f6789421b44327167a8b52

Observation 6b44838c-7224-42f5-87f6-6b742d9b635c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Can Visual Encoder Learn to See Arrows? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.645208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.645208Z digest=sha256:a9251dae517149c6b0db0532a1b061eb338cb0e0eb09a95ca4c723d4cc09470a

Observation da8f3efc-55c7-4383-8605-3a9269e26372 · outbound

This paper cites How well do vision models encode diagram attributes? Inthe ACL 2024 Student Research Workshop, 2024.

Can Visual Encoder Learn to See Arrows? How well do vision models encode diagram attributes? Inthe ACL 2024 Student Research Workshop, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.086930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:19.770636Z digest=sha256:58492d4f93cd634b6c5636688c7ad7cb8121ed4e5aeb158ffb68540f9e189cf8

Observation 4d5510b3-200f-4346-8ad0-b15a579ca76b · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert AGI.

Can Visual Encoder Learn to See Arrows? MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert AGI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:20.800916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:19.873890Z digest=sha256:dbdea0a03c7f8184d0abac3eec1bcca554a8070a17bce947043141855c1e2569

Observation 53cb6495-33d5-4ab9-892d-a7e523cbd428 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Can Visual Encoder Learn to See Arrows? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:20.608602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:19.966361Z digest=sha256:68f0faabf375bf885e92bbe47808444129efb78adacd74db6993b04944f528e9

Observation 96ce0fac-082c-43f0-8b31-9f7c8e30d11e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Can Visual Encoder Learn to See Arrows? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:20.058994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:20.058994Z digest=sha256:d8e7ea438ccf2a9fbe3132fdbdb61023af0a05caed19b54088ce96abe6900976

Observation c3a0f9ef-7804-437b-8a68-3ca0b5960f85 · outbound

This paper cites Accessed: 2025-04-12.

Can Visual Encoder Learn to See Arrows? Accessed: 2025-04-12

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.028089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:18.935665Z digest=sha256:fc1a5676fba2073d0860b1e10c08d0e3f0af27b598cb83da8ef962c5f9a3d71b

Pith citing papers

No inbound Pith citation observations are available.