REVIEW 4 cited by
Visual Programming: Compositional visual reasoning without training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present VISPROG, a neuro-symbolic approach to solving complex and compositional visual tasks given natural language instructions. VISPROG avoids the need for any task-specific training. Instead, it uses the in-context learning ability of large language models to generate python-like modular programs, which are then executed to get both the solution and a comprehensive and interpretable rationale. Each line of the generated program may invoke one of several off-the-shelf computer vision models, image processing routines, or python functions to produce intermediate outputs that may be consumed by subsequent parts of the program. We demonstrate the flexibility of VISPROG on 4 diverse tasks - compositional visual question answering, zero-shot reasoning on image pairs, factual knowledge object tagging, and language-guided image editing. We believe neuro-symbolic approaches like VISPROG are an exciting avenue to easily and effectively expand the scope of AI systems to serve the long tail of complex tasks that people may wish to perform.
Forward citations
Cited by 4 Pith papers
-
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
SFT+GRPO training on CanvasCraft teaches an MLLM to orchestrate heterogeneous visual tools for long-horizon image creation and editing.
-
Reinforced Visual Perception with Tools
ReVPT uses GRPO reinforcement learning with a cold-start SFT phase to make Qwen2.5-VL models call visual tools, improving perception benchmarks over SFT and text-only RL baselines.
-
Multimodal Video Emotion Recognition with Reliable Reasoning Priors
Injecting MLLM-generated multimodal reasoning traces into a fused audio-visual-text emotion model, plus a balanced dual-contrastive loss, raises average accuracy on MER2024 from 77.5 to 84.7 percent.
-
Augmented Vision-Language Models: A Systematic Review
A structured taxonomy of inference-time augmentation techniques that connect vision-language models to external symbolic systems, tools, and knowledge sources.
Discussion (0). Continue with ORCID to comment.