Pith. sign in

Paper Citation Record · LEDGER

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

As of 19 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 1 inbound Pith citation observation for arXiv:2502.00372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00372 v3

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:21:22.459302Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:09:18.020873Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T17:09:19.193011Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6639aa5-f542-4942-8658-4114c6c30a7d · outbound

This paper cites Words Aren’t Enough, Their Or- der Matters: On the Robustness of Grounding Visual Refer- ring Expressions.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Words Aren’t Enough, Their Or- der Matters: On the Robustness of Grounding Visual Refer- ring Expressions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.907645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.072967Z digest=sha256:d5a009f0633feec1d4d31464ab2820e22c4e6c1ab115950bb61c5564bbd076e0

Observation bcc7def2-ce42-4eba-ba17-dfc6f8a45a74 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikołaj Bi ´nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikołaj Bi ´nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.893800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.078694Z digest=sha256:e49822c87cf975a1719b58a35f4c48448e9223a5daf299fa56797de78b1b3c9e

Observation b6defd9b-b9dc-44e9-83a4-fbaccbb86a99 · outbound

This paper cites Neuro-Symbolic Visual Reasoning: Disentangling.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Neuro-Symbolic Visual Reasoning: Disentangling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.879277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.083345Z digest=sha256:8ba541ea3225c2e2f7ac8bdedde5c3a3ead1149c9ab453158b8fd9dde648509f

Observation 058c88b8-f54a-4f49-ae26-c407b353a43e · outbound

This paper cites Qwen2.5-VL Technical Report.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.088741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.088741Z digest=sha256:206f96f71d4abc42782ae60db249fadeb557baf58a1ef47b4c10d7bfa114d400

Observation 1fc24f05-2638-4ebb-9139-ceb4c5f19d45 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.094229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.094229Z digest=sha256:5412002e3560b9442c1ac0e0ac76d9404856efcdecf3bf7e6cecd9f9f9a990d8

Observation b7463ed9-a233-4a9a-b950-a44f74d3e711 · outbound

This paper cites Language Models are Few-Shot Learners.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Language Models are Few-Shot Learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.863541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.100399Z digest=sha256:7e3ea0505514e023614f3951a78e7ed9cae0afc3cbdaaf3571db1b17ae50f77b

Observation 31dd632b-0b6e-4719-9afc-503ba9145641 · outbound

This paper cites SpatialVLM: Endow- ing Vision-Language Models with Spatial Reasoning Capa- bilities.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning SpatialVLM: Endow- ing Vision-Language Models with Spatial Reasoning Capa- bilities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.847163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.106680Z digest=sha256:1dc7e1182e3e902de104068bf0360ad63fe5806ac4e67a0e465c3489a9eb3427

Observation 59907a97-ae24-46fb-a19a-a9470ceb1a22 · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathemat- ical Expression.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathemat- ical Expression

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.829752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.111681Z digest=sha256:f3cd2d1e0842a9f8bf99c8062adf4f229be6fdeda0d36b6cd3c13ece1c1bef8d

Observation b8063a8f-5f6e-456f-923a-12b6e79e4f1e · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.117180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.117180Z digest=sha256:5463b8f87fcb35fac99f5b0ce0af20dce88f829711cc41076a21d1b0dee2f01a

Observation c89b175b-0b3c-4705-bb6c-33462768bc20 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.813610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.123211Z digest=sha256:4e35fef32c74e80f0f1033a9e9fc53d624a77c12aba74bee221a481f9838d421

Observation 9b161f01-49b7-40e2-9338-50d4c7ec2563 · outbound

This paper cites YOLO-World: Real-Time Open-V ocabulary Object Detection.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning YOLO-World: Real-Time Open-V ocabulary Object Detection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.796691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.129737Z digest=sha256:7ced6f8bff366f1159c6a86990a79190c1028c2db44d76e3154e6ac2d82ad857

Observation 11a1c6fb-c042-4573-8190-0bd4794905c6 · outbound

This paper cites an unresolved cited work.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:21:23.780375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.134922Z digest=sha256:c5fa4d5a4d40ef40e471b6c89c9a1e30115159d9d70766dd81a0bc58ce11a631

Observation 31a371a9-7fde-4fc3-a3de-acc4ea461166 · outbound

This paper cites SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.764754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.140257Z digest=sha256:f385c9fe3ee62882006d1c292f3b113e6ac403e5ddf050c78a62cd3dd2a6de70

Observation 673e0e28-8bfe-43ea-9e43-9ae2b3e79906 · outbound

This paper cites InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.145505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.145505Z digest=sha256:36ceed7b1d67455a3851ab83c8ee546940a18d744a521a7c698b1e47bccd9f6a

Observation f28592fc-8cc2-47b2-80ae-43541ac96b98 · outbound

This paper cites ProbLog: a probabilistic prolog and its application in link discovery.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning ProbLog: a probabilistic prolog and its application in link discovery

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.737699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.157235Z digest=sha256:d76fd125eb24bde2f292f7b10261224e6dfc78845fff049281e333a863f35b9c

Observation 6d7a8aeb-fb7e-4485-b387-d0195e7c062d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.162212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.162212Z digest=sha256:640bd499de984186b96447aaaa5889058a57bb005e28a3bc0b388d6910508a36

Observation 2dd95166-f5ec-42d6-bc57-471da1c52eaf · outbound

This paper cites Vision-Language Transformer and Query Generation for Re- ferring Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Vision-Language Transformer and Query Generation for Re- ferring Segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.722690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.167727Z digest=sha256:570bc59820e95778a186d123625aeb9ff87be70958615abea2967e0198fc4819

Observation 858a74b3-8f84-4879-ba2b-612389f6eada · outbound

This paper cites The Llama 3 Herd of Models, 2024.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning The Llama 3 Herd of Models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.706289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.172924Z digest=sha256:05e7855d8ca7829f565c315f39c2d7f8cb76e419855a8ed7901f51c8d8efe9fe

Observation 7bf0417c-9e0d-41bb-9fa5-2ef9f4a6b56c · outbound

This paper cites Visual Program- ming: Compositional Visual Reasoning Without Training.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Visual Program- ming: Compositional Visual Reasoning Without Training

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.691404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.178028Z digest=sha256:91827e257a92c7b15ee1d4645fb52ac28144d8d69e0be1e4b1839036d1384bac

Observation cc9040eb-c2f4-44b0-b399-6eb28f0f04da · outbound

This paper cites Video OWL-ViT: Temporally-consistent Open-world Localization in Video.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Video OWL-ViT: Temporally-consistent Open-world Localization in Video

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.674438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.183138Z digest=sha256:227d05d8b3735dcdb75854ca78540667068a3211d15ad90dac22d8eb7d8dd523

Observation eafc050e-a398-4b8a-ad93-d9e9b54e5424 · outbound

This paper cites ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.658492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.188494Z digest=sha256:124112d2441aefb234b9c19c7e0b608b6662b37a629fff11783ed05cf3aa9cb1

Observation b7a9271c-8b80-445a-846e-cc548993ec8e · outbound

This paper cites HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.640787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.193736Z digest=sha256:fc8a1f0c7e4900d0ef78e6b081668650bb08f8eff2067f3353f669a7f3251a1d

Observation 0066eba5-fccc-4e59-a488-e583db03f4d7 · outbound

This paper cites Segment Anything.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Segment Anything

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.204372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.204372Z digest=sha256:ed99ac5286509fd311753f9e5b6209c701ed616002adf30f718330a19bec7c18

Observation a6bbe0b8-3650-4a0c-9f28-0bfdfeed0284 · outbound

This paper cites LISA: Reasoning Seg- mentation via Large Language Model.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning LISA: Reasoning Seg- mentation via Large Language Model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.607109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.209975Z digest=sha256:8086bba0d0af6af5cb69ebb09941e86a6d3a42b8439e7b90da5ad6658c6c2294

Observation 635c4a1d-3db7-4901-aa9a-a9a9be8a92fa · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.215330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.215330Z digest=sha256:67f4fc8b1b938f7860ba87b48238a848d1fd2f8e4a3ecf497278d07b04d7f3f1

Observation e2ee7a45-9a72-42d3-ad53-d822732f5b3c · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.588928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.220888Z digest=sha256:c01f2cd2c5fafbede93b1de17ce1cb85d8216528ae1c1b4b43f70c78936bab3b

Observation 4f29b02a-279f-4de1-bdd0-4a74d9057284 · outbound

This paper cites Grounded Language-Image Pre-Training.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Grounded Language-Image Pre-Training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.573052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.225228Z digest=sha256:f1771536dbe0bdb1d83f9923838dc5129dccbcf9ca39d4e575e9247066783ecf

Observation 7f14091b-929b-42ed-8698-4aa5b43c44fa · outbound

This paper cites Scallop: A Lan- guage for Neurosymbolic Programming.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Scallop: A Lan- guage for Neurosymbolic Programming

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.557753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.229629Z digest=sha256:b809555792a0525f8886ae5e89cfde898fe1205082bef9360778bf6a772d1b35

Observation 16008e4b-fcd4-43e3-9578-ba5edafa03e9 · outbound

This paper cites GRES: Gen- eralized Referring Expression Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning GRES: Gen- eralized Referring Expression Segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.542239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.233935Z digest=sha256:4a8d3c5b5c44ac392c032cee45707bfd10bc00e8197b31e463e88c1281a52e68

Observation 7219750f-7efb-4088-aa01-9cdaaea59751 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Improved Baselines with Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.238320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.238320Z digest=sha256:2fbc02caa67af244b491a9f390f5e2a28029c4b8002e971aa2592415f2877ec0

Observation b465832c-150d-45c4-8ea4-17b23d782d85 · outbound

This paper cites Visual Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.242857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.242857Z digest=sha256:ffd445ee100d7931ed6677a2277aad31b3c08bc1d5480378abaefbe63c10dd7c

Observation 28ffe688-f29a-4dde-a1db-792d5e1e5c62 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.526504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.247511Z digest=sha256:5944d750b94b035e907680eb3c0def57dadc966c850dbcc4e14a8f41f332984e

Observation 8a0f35a8-5794-414c-979a-45d5f353eff8 · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.510559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.257759Z digest=sha256:22fc13ffe9680218fd8f278f0cb79a021c568a6d307e5df73efb8fcb0ce2fec8

Observation 19f2f243-e1c1-4c2e-a7c1-cd2090574d27 · outbound

This paper cites Multi-Task Collabora- tive Network for Joint Referring Expression Comprehension and Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Multi-Task Collabora- tive Network for Joint Referring Expression Comprehension and Segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.493566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.262858Z digest=sha256:a44a7775c39c92039773f7894727a307b872677067c7eb8a7d6fab651aa6d8b3

Observation cf47f353-1114-4f8f-9ed3-e3dd2cf67b9a · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.252109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.252109Z digest=sha256:bce586fec19df898089b0e91efb3306886d93ad677e3f7485bc4112b171af4fc

Observation ef2bf3cf-038f-42e6-9119-e48b06d418ca · outbound

This paper cites GPT-4o System Card, 2024.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning GPT-4o System Card, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.459981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.273010Z digest=sha256:cd3faaa13cf7d54f6417db4c2cd2bdcbf79183364566f53f06a3c0cdea3216a0

Observation 18e9c71d-6314-422e-b5e6-5e9c7c7d05a3 · outbound

This paper cites Christiano, Jan Leike, and Ryan Lowe.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Christiano, Jan Leike, and Ryan Lowe

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.445602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.278013Z digest=sha256:8bfd62a9ffecc51e438bccd6ea1185814fc92c612ea3bec87cfca3f6cc2367ec

Observation 636e7608-8c33-4e83-b740-3b8a787115f8 · outbound

This paper cites Scaling Open-V ocabulary Object Detection.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Scaling Open-V ocabulary Object Detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.476105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.267940Z digest=sha256:dadd2ff73e8fb2103c54243a59c79283a268ad56dc242abd06c71f665851e0a1

Observation cfa9d043-281e-4300-a349-fcce7730aed4 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.288779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.288779Z digest=sha256:c5f11a126885548955d766186721adcf8996cd17f6acd8ce42871ad3fde0be86

Observation b438e301-e867-4968-9a69-cccb0519f80e · outbound

This paper cites PerceptionGPT: Effectively Fusing Visual Percep- tion into LLM.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning PerceptionGPT: Effectively Fusing Visual Percep- tion into LLM

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.415996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.294444Z digest=sha256:d7b0b78725cf81a3ad752194b098d45379177f333d2fc3ebd24e668714e7f90b

Observation ad6f9421-5caf-4284-981b-f694c4ba7320 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.431240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.283060Z digest=sha256:9e827f99bbad445b8b286a310933926509b1d570bc83a0bbfb47bb3eba0a52f0

Observation d0d9da8c-6868-4a37-8486-a8b9abca8f41 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.390115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.304626Z digest=sha256:2e918b4b96aef22211b9fd07aba03dd3c0f593e48290c7396c827a219576d57d

Observation 8ed197d9-5b23-4c00-bd0d-0a709d72d192 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.374381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.310038Z digest=sha256:defd29e198c63ab5d713f6bb85f01ea7f86e1832957b27c2f6229485ced0e20b

Observation 158df1f1-bce1-4be1-95d7-52fd3a0d73e2 · outbound

This paper cites Language models are unsuper- vised multitask learners.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Language models are unsuper- vised multitask learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.299518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.299518Z digest=sha256:1515c7c433e3ff273e6f4e6375a8a14554e01c36619eaf33d57dc8700b7126dd

Observation 801fb7c1-26c1-49a1-b9b4-00660efe8695 · outbound

This paper cites ViperGPT: Visual Inference via Python Execution for Reasoning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning ViperGPT: Visual Inference via Python Execution for Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.326159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.326159Z digest=sha256:9dafa7bae0cec39509bd9092a253061ac88e5808c47689497a54fc436a3fcb0a

Observation b313a20f-a130-4ece-aa42-1db54810ae0b · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Gemma 2: Improving Open Language Models at a Practical Size

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.331855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.331855Z digest=sha256:947c7e87f90f90c0179ad78a9a208f6006ed9464f739bba3e1856c01acae4036

Observation 303c93aa-f55b-448a-9493-d46515b391de · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.315609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.315609Z digest=sha256:c812674c47d0b8c9fa1eac0bbf515b6ead9ae0ea1350cbab9fca21266c3d06c7

Observation c8d15d09-73c9-4158-b08b-53d52a851ef5 · outbound

This paper cites Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.358955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.321165Z digest=sha256:2b2adfc2edc28a9c728f4ff410a33b08b1d15954561d5a753e573c2556884dd1

Observation 2dab5cce-0165-443d-b23a-46ecc6ae5b64 · outbound

This paper cites CRIS: CLIP- Driven Referring Image Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning CRIS: CLIP- Driven Referring Image Segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.327510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.353958Z digest=sha256:8985d7e4ff931d1f65901c06bab979f7a5b2b377185228165bad77a30e912537

Observation 0828f7ac-e6ee-4f1f-88cd-065ac54e2dbe · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.358878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.358878Z digest=sha256:13a0eb4c901ec93edade68ff1e9ebc6ec7939f89f4a4f72a7efaea0cb10df3c5

Observation 4731627e-7fa4-4e11-a8e2-dc5e07e6b0df · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.337684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.337684Z digest=sha256:08fd707a6a99be2db7d3655d6981a3ac21a3e10c1ef9b7fa3f0b363fd7bea737

Observation de374cc4-7d5c-4d3e-83f8-e18810316348 · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.343366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.343222Z digest=sha256:b72facf88c7724730e7a43ac0f6c326b7f172b1d683f23de59ca976141bf336e

Observation 6aaae9ad-0aa3-4802-992e-9e3a7360e4e9 · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.348314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.348314Z digest=sha256:d4659ac42f1f28744b2ce44292a0de390df6d7479d629fd99eb68f88aae9b5b8

Observation a96f5036-d889-42c1-984c-4acd7ad1c729 · outbound

This paper cites an unresolved cited work.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.378332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.378332Z digest=sha256:a7eee455620e0f872ea3d8e18d975bcc99d509bb0b7524f5401d75ae59c6f233

Observation aa0e58d1-885b-422d-9749-ff4462c09c4f · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.382440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.382440Z digest=sha256:d8159539755494638c0537ff20ce007e6cf98c97e2160da6b22edc74813e3b72

Observation 871c4183-e5c8-4bfa-98bc-e6d6e3c02401 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.364061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.364061Z digest=sha256:fffcb05db32e48fac304261a46626ea3f260a42b7c06d0fbe45936e8089fe1f0

Observation f52294ea-dba9-488e-a57e-552a7d76b732 · outbound

This paper cites GSV A: Generalized Segmentation via Multimodal Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning GSV A: Generalized Segmentation via Multimodal Large Language Models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.311665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.369269Z digest=sha256:1421fb1487dc8ef0cb1af2b91b402b6bbd0053b082f6fadb8b12390068e23a1a

Observation f4be7c81-dd03-4907-a27d-19caae15154b · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Va- riety of Vision Tasks.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Florence-2: Advancing a Unified Representation for a Va- riety of Vision Tasks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.190942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.373917Z digest=sha256:aa04d9d0a2eda8c50d47ef5000d38791b482283e891a8e55e7da96f896b877cf

Observation cf5111f1-ef1d-498b-90c9-e1d0bcbfcc34 · outbound

This paper cites IdealGPT: Iteratively Decomposing Vision and Lan- guage Reasoning via Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning IdealGPT: Iteratively Decomposing Vision and Lan- guage Reasoning via Large Language Models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.142095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.399485Z digest=sha256:f9bd09954c6f2a91c2e5c81b840471c87729d99c893c9f24975cac9b614030f4

Observation ee68c0b7-80ea-4f64-9e94-6db57a70e0eb · outbound

This paper cites Berg, and Tamara L.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Berg, and Tamara L

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.125998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.403779Z digest=sha256:ca6a396b841ea6256e173ee6396aa8cffd6767bc4706a0797cdefb012bad1f53

Observation 204ca23d-d05b-4f61-8bab-778710993bb4 · outbound

This paper cites Depth Anything V2.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Depth Anything V2

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.386798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.386798Z digest=sha256:03d08940b22c64c2facc9ccca0ebcbc36baac62191d65d0ebcd117b495b1126f

Observation 3f5e6032-e629-4330-afe5-ef8985c5386c · outbound

This paper cites Cross-Modal Rela- tionship Inference for Grounding Referring Expressions.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Cross-Modal Rela- tionship Inference for Grounding Referring Expressions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.174061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.391084Z digest=sha256:1e073bd4902c577a57b7f14eb5a723222aa0b00c8b6e54f48fc2d28e3e607991

Observation ea5fc35b-df3a-45ab-b4b1-d35a4ce2b789 · outbound

This paper cites an unresolved cited work.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:21:23.157455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.395185Z digest=sha256:26ecc76b67215729a095c5578c02cf6769fa2123f6fbcc6ff5b5cb55f6af91c0

Observation 0fdb9299-4307-4693-9fd3-6e763cecfb25 · outbound

This paper cites Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.109629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.408130Z digest=sha256:d31f56d240ddefde7343570bb39f302aa0c291dbc10362c1829d2b0ee61ac335

Observation c5662df4-091f-4c5e-a5f9-8449060109a6 · outbound

This paper cites PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:21:22.525450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.413504Z digest=sha256:798a4e97e2f5a9e9d4f87fd8cefbbe92e60718ca88497f71f7e224eb06975f13

Observation 6fb6f62e-3ad3-4506-8ab4-010f3d7ca1c6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.418661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.418661Z digest=sha256:2e26de7578a29fb47a807af55545deada45ed2731cef234a95f1a7d014f10fe6

Observation c660ab1a-4fdf-409a-8e52-8013b95abe95 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.424064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.424064Z digest=sha256:bc27874251accc1ca71d84c16c90872ab771a7470d84c417c18146fc8873e5aa

Observation 0ad299b5-34e6-4a09-97a9-3db4dbf26d03 · outbound

This paper cites Table 9 presents a quantitative comparison on the RefCOCO, RefCOCO+, RefCOCOg [21], and Ref- Adv [1] datasets.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Table 9 presents a quantitative comparison on the RefCOCO, RefCOCO+, RefCOCOg [21], and Ref- Adv [1] datasets

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.082917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.429171Z digest=sha256:15c8c2cdf69bb226f9ea930ee1db82c7e8e0a90bec3775f1c147e4a0b932028c

Observation 9b871822-dd7c-4635-ab03-3c39b278c443 · outbound

This paper cites For a fair comparison, we use the same foundation models and LLM for the ex- periments.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning For a fair comparison, we use the same foundation models and LLM for the ex- periments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.067105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.434128Z digest=sha256:4dc5d862c3b53fa172a36c8bb03378cf5f09097979fa7bee70728b3d2f9a5bda

Observation 781c9b0b-4f0d-4c1e-a214-78337446f75f · outbound

This paper cites The results are shown in Table 11.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning The results are shown in Table 11

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.050044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.439036Z digest=sha256:1f157a881d9463fb4af26d00db707cb01dcbdb3c4524010a79ee8635f0915afb

Observation 504a224c-bc61-4fc1-8b62-96165ff670e5 · outbound

This paper cites In this mechanism, each time the system transitions into the self-correction state (indicated by a red arrow in Figure 2), it is counted as one retry.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning In this mechanism, each time the system transitions into the self-correction state (indicated by a red arrow in Figure 2), it is counted as one retry

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.034758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.444127Z digest=sha256:9bdcf5aeef3328dce584d9af9a58c6a1cb15939213778bbd91dec4a4e2a5198e

Observation 53f673b3-c505-4da8-94d8-5b95494c63aa · outbound

This paper cites The results, shown in Figure 4, indicate that NA VER consistently reaches SoTA performance compared to all baselines regardless of query length.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning The results, shown in Figure 4, indicate that NA VER consistently reaches SoTA performance compared to all baselines regardless of query length

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.018330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.449051Z digest=sha256:20fb6cf8280f61b83c9437a83dd51e47419db958fe4ed9784b5aba45fd56d3bd

Observation 3ce02b20-ac36-4fb9-a623-5a0cfb14e61f · outbound

This paper cites An intuitive alternative is to skip captioning and ask the VLM to predict categories directly.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning An intuitive alternative is to skip captioning and ask the VLM to predict categories directly

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.002013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.454378Z digest=sha256:558796d9593dcab8c509e58a830ef895c379fe7c9a230fe8790213fd284203ec

Observation 91e6a8e7-9401-490a-b0c7-454915489a28 · outbound

This paper cites Yes” or “No.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Yes” or “No

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:22.985311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.459302Z digest=sha256:145a1bf6843fd201c8f829b4b7dbe0526bac3b7ea5e235d276f9e29d0353c0c3

Observation c2fe5022-c14a-4810-bd90-a991e5487999 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.151438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.151438Z digest=sha256:003d96ff682c0bb8199e84b54927c2a3954991a2c7e65fbfdf4cea131ccac5a3

Observation bcdb2266-fb07-4797-9585-249b93d864fc · outbound

This paper cites 2, 3, 4, 6, 7, 1.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning 2, 3, 4, 6, 7, 1

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.624357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-09T19:21:22.199123Z digest=sha256:8b9c7ed0294c0ba37f33433602c2f013c935669f9b2b6102c14633e6bb2d0fe4

Pith citing papers

Observation d45b8d54-c1a7-4f87-a5ad-e3061a8e04f4 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:09:19.196908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:09:18.020873Z digest=sha256:6cecd5415e7bd6f619e1f9507327665e0c9d354b22f10eb5b996371dc105917c