Pith. sign in

Paper Citation Record · LEDGER

Visual Programming: Compositional visual reasoning without training

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2211.11559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.11559 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:11.242711Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T19:56:10.525466Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 42e98bda-38d9-4163-a74f-459177b6fbbd · inbound

ViperGPT: Visual Inference via Python Execution for Reasoning cites this paper.

ViperGPT: Visual Inference via Python Execution for Reasoning Visual Programming: Compositional visual reasoning without training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:15:14.565790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T18:15:14.382011Z digest=sha256:a6c24387a69c59391a8bbcfed10800b5866dcd667357fc2d59015ca612d370c6

Observation 8c83b403-160f-405f-9f48-0c91077ad0af · inbound

HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face cites this paper.

HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face Visual Programming: Compositional visual reasoning without training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:06:45.617624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T00:06:45.440071Z digest=sha256:8b2a2d3c3e36f784ea53116e0e5f3f45718288efc3dba103c9ab31ad49c90357

Observation 1f21bfbc-ee16-4bac-8f20-251339993da7 · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Visual Programming: Compositional visual reasoning without training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:22:03.757506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:45d6fd31f20ccdcf78386632b3cfb4cfbd37999ae7ed5ce4be824798174f5d3f

Observation f0a27056-2b7e-4fe8-b71c-96721284a634 · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model Visual Programming: Compositional visual reasoning without training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:41:04.873384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:a2731501772c56419edd4c1e24239ea8e2c151cdca638e8dd2571bd43e1ebeb8

Observation f8603a51-e249-4eaa-9a65-c1ee18503d9a · inbound

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction cites this paper.

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction Visual Programming: Compositional visual reasoning without training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:53:33.145856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T22:51:08.753650Z digest=sha256:caea763e01c62808bf90d7353600121d05bab4f02e9bf616f22fbf6c7f2e61a2

Observation a467ee9f-2e36-41dd-b523-95bcc042a8fd · inbound

Language Model as Visual Explainer cites this paper.

Language Model as Visual Explainer Visual Programming: Compositional visual reasoning without training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:07:35.037174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:07:35.037174Z digest=sha256:443cc3d182c8d3d14dfed50ccb01b29746784bb6614cf6adc62433eccae91168

Observation 28e812d9-a7f9-409e-9fad-6f9edc821321 · inbound

Relational Programming with Foundation Models cites this paper.

Relational Programming with Foundation Models Visual Programming: Compositional visual reasoning without training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:13:21.402018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:13:21.402018Z digest=sha256:77c0ced68bd280a99026ed02b57918593e09c5e0e9696bd6dd41b2dd3454f29f

Observation 8e207c62-c1f2-4204-a0ec-8178ec26d481 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Visual Programming: Compositional visual reasoning without training

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.771294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.771294Z digest=sha256:787e21cfb512ef6ece0cce262794044d28c09d16de0246c8793dbeb5283dd7f3

Observation e413db83-12a3-4a93-b623-a79f492e4b8f · inbound

PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction cites this paper.

PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction Visual Programming: Compositional visual reasoning without training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:17:54.047245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:17:54.047245Z digest=sha256:17ff8f4f9e0e27c2a6891d4e500b6ad020965d6ec8d019d790e90645cbf588b0

Observation ffe77042-955f-4f49-b8d4-e540e43edcf0 · inbound

Multi-Modal Language Models as Text-to-Image Model Evaluators cites this paper.

Multi-Modal Language Models as Text-to-Image Model Evaluators Visual Programming: Compositional visual reasoning without training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:11.242711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:11.242711Z digest=sha256:501acb839b9408cab5c5d4ad67e0a060c1a9ece6288a4c3a5a6d43b8e6de96de

Observation 212acf97-e428-47db-ad61-4957bd06036c · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Visual Programming: Compositional visual reasoning without training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.237592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:160b38e05c30719b7b2d9c4d0e524a356e8ad5ebdd2af71d9dcabf4463c8d24a

Observation d3a5bc0b-4cf7-4701-9138-2d6a1f413096 · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review Visual Programming: Compositional visual reasoning without training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.978926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.978926Z digest=sha256:3d58a8c2b9ee06b39aebfc9e0a4dc0458cec85c2eadb83915b7d5438e9c8f62e

Observation acbf79e9-a453-445c-81e2-2cf1ccce4be1 · inbound

Multimodal Video Emotion Recognition with Reliable Reasoning Priors cites this paper.

Multimodal Video Emotion Recognition with Reliable Reasoning Priors Visual Programming: Compositional visual reasoning without training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:30.875408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:30.875408Z digest=sha256:d8936cf06095b586a8d7f2ea756670997902517ec11ed0c77ae57f7e827e8063

Observation 843ab866-a909-45f7-825e-7ecfb8bfde60 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools Visual Programming: Compositional visual reasoning without training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.935858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.935858Z digest=sha256:67d9fc0b07088133f8719b98fe8a79c8ae40a11620ab5338fa7a912c57f19c0f

Observation 62e0ba91-b313-4084-988c-b66173ef2b45 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Visual Programming: Compositional visual reasoning without training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.916710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:d6b444d5db911ff399d76d18d1f5bb1e4e394448eba6fc2724368521a4e3e07b

Observation 3f6da490-24b7-4415-b73c-b8595729239d · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Visual Programming: Compositional visual reasoning without training

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:51:23.156390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:295b11a0193f0032092d87c535026e631ff7becf9006d9cbfa09620625decdbc

Observation 498614d5-cf49-4cb6-97b1-05fcf713af9b · inbound

VESTA: Visual Exploration with Statistical Tool Agents cites this paper.

VESTA: Visual Exploration with Statistical Tool Agents Visual Programming: Compositional visual reasoning without training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:10.526893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T21:58:11.339217Z digest=sha256:437c6ede8920a67a9f75171f8321089a9d6131131345270b2e05d5c36ec3d14f

Observation 0359ec15-96f7-433d-a80b-280fa4e0a837 · inbound

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration cites this paper.

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration Visual Programming: Compositional visual reasoning without training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T15:24:53.016199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:24:53.016199Z digest=sha256:b832c95908117403a34414d68c20bed762bc340b01c802c15812f4cf835aef87