Pith. sign in

Paper Citation Record · LEDGER

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

As of 20 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 1 inbound Pith citation observation for arXiv:2502.00372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00372 v3

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:21:22.459302Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:09:18.020873Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T17:09:19.193011Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6639aa5-f542-4942-8658-4114c6c30a7d · outbound

This paper cites Words Aren’t Enough, Their Or- der Matters: On the Robustness of Grounding Visual Refer- ring Expressions.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Words Aren’t Enough, Their Or- der Matters: On the Robustness of Grounding Visual Refer- ring Expressions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.907645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.072967Z digest=sha256:8d453e7f37cc92f3c3f23aff8f76fcdf186ec98342e2b62f6a07c63798132dc6

Observation bcc7def2-ce42-4eba-ba17-dfc6f8a45a74 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikołaj Bi ´nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikołaj Bi ´nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.893800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.078694Z digest=sha256:e65a9ee676cb9f603aa07729296cef518eff51ab73a404363e1e9292022380cb

Observation b6defd9b-b9dc-44e9-83a4-fbaccbb86a99 · outbound

This paper cites Neuro-Symbolic Visual Reasoning: Disentangling.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Neuro-Symbolic Visual Reasoning: Disentangling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.879277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.083345Z digest=sha256:55801f9d9dc719d83e6f5db252bf6c351f8637d473e435f08e8c8fc94040bce1

Observation 058c88b8-f54a-4f49-ae26-c407b353a43e · outbound

This paper cites Qwen2.5-VL Technical Report.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.088741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.088741Z digest=sha256:206f96f71d4abc42782ae60db249fadeb557baf58a1ef47b4c10d7bfa114d400

Observation 1fc24f05-2638-4ebb-9139-ceb4c5f19d45 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.094229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.094229Z digest=sha256:5412002e3560b9442c1ac0e0ac76d9404856efcdecf3bf7e6cecd9f9f9a990d8

Observation b7463ed9-a233-4a9a-b950-a44f74d3e711 · outbound

This paper cites Language Models are Few-Shot Learners.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Language Models are Few-Shot Learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.863541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.100399Z digest=sha256:66cbd14abb1bd93dbf153e5e645d2b83f2d4ec71ebcc64aae7c662f82a652f64

Observation 31dd632b-0b6e-4719-9afc-503ba9145641 · outbound

This paper cites SpatialVLM: Endow- ing Vision-Language Models with Spatial Reasoning Capa- bilities.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning SpatialVLM: Endow- ing Vision-Language Models with Spatial Reasoning Capa- bilities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.847163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.106680Z digest=sha256:c8d41e949517faa82b119721c47ee2b58a6cf3c7580c910ab17ad65c4d781daa

Observation 59907a97-ae24-46fb-a19a-a9470ceb1a22 · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathemat- ical Expression.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathemat- ical Expression

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.829752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.111681Z digest=sha256:a406f1eaf7589f2f8a42c236829284b750e690aadac3f2a90fad2237612d79a7

Observation b8063a8f-5f6e-456f-923a-12b6e79e4f1e · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.117180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.117180Z digest=sha256:5463b8f87fcb35fac99f5b0ce0af20dce88f829711cc41076a21d1b0dee2f01a

Observation c89b175b-0b3c-4705-bb6c-33462768bc20 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.813610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.123211Z digest=sha256:860bf1a03b8f1d797f9a49d1c26381e2c00ad25ff8231ebc1167610bd26c440e

Observation 9b161f01-49b7-40e2-9338-50d4c7ec2563 · outbound

This paper cites YOLO-World: Real-Time Open-V ocabulary Object Detection.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning YOLO-World: Real-Time Open-V ocabulary Object Detection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.796691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.129737Z digest=sha256:3ac4a89fc7afb0a5cfe2ad2585aaddab48dfd2d88ece088427e247c032902719

Observation 11a1c6fb-c042-4573-8190-0bd4794905c6 · outbound

This paper cites an unresolved cited work.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:21:23.780375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.134922Z digest=sha256:feaa36b66d71b84924a6809b6ab38c8933a9c8277f3cc688af8ccb20b33b2a16

Observation 31a371a9-7fde-4fc3-a3de-acc4ea461166 · outbound

This paper cites SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.764754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.140257Z digest=sha256:ef54e3243c6927d1dc738990ea5540d56e35b363094d3d67f74fed27ecfdb38b

Observation 673e0e28-8bfe-43ea-9e43-9ae2b3e79906 · outbound

This paper cites InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.145505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.145505Z digest=sha256:36ceed7b1d67455a3851ab83c8ee546940a18d744a521a7c698b1e47bccd9f6a

Observation f28592fc-8cc2-47b2-80ae-43541ac96b98 · outbound

This paper cites ProbLog: a probabilistic prolog and its application in link discovery.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning ProbLog: a probabilistic prolog and its application in link discovery

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.737699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.157235Z digest=sha256:f296d6d0c56f8d87acfe4dd1ba78761e2279e9e877dae7950575cc90bb66052f

Observation 6d7a8aeb-fb7e-4485-b387-d0195e7c062d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.162212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.162212Z digest=sha256:640bd499de984186b96447aaaa5889058a57bb005e28a3bc0b388d6910508a36

Observation 2dd95166-f5ec-42d6-bc57-471da1c52eaf · outbound

This paper cites Vision-Language Transformer and Query Generation for Re- ferring Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Vision-Language Transformer and Query Generation for Re- ferring Segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.722690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.167727Z digest=sha256:5cabf3160d609cbe34b19a751e886c384bfa1adc09ade5a3a647f2ac195950c2

Observation 858a74b3-8f84-4879-ba2b-612389f6eada · outbound

This paper cites The Llama 3 Herd of Models, 2024.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning The Llama 3 Herd of Models, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.706289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.172924Z digest=sha256:1a883523b7716028d1b1fd6faa8a811f5e6ba73a7328e344c1be329603b15d6d

Observation 7bf0417c-9e0d-41bb-9fa5-2ef9f4a6b56c · outbound

This paper cites Visual Program- ming: Compositional Visual Reasoning Without Training.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Visual Program- ming: Compositional Visual Reasoning Without Training

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.691404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.178028Z digest=sha256:1b0def597d66cd94d4b849db32a0bf02d0e98db3cc44a6bfe06707134af9bad9

Observation cc9040eb-c2f4-44b0-b399-6eb28f0f04da · outbound

This paper cites Video OWL-ViT: Temporally-consistent Open-world Localization in Video.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Video OWL-ViT: Temporally-consistent Open-world Localization in Video

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.674438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.183138Z digest=sha256:ec2afe29b292c31f6fe2454cac571f015ab89d354421b71dadfd00a442fb02e5

Observation eafc050e-a398-4b8a-ad93-d9e9b54e5424 · outbound

This paper cites ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.658492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.188494Z digest=sha256:83fce8424a5c4b03778477559091f278ea4bdf4fa0162b3f28e9b8cdb94af310

Observation b7a9271c-8b80-445a-846e-cc548993ec8e · outbound

This paper cites HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.640787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.193736Z digest=sha256:2bc653c3a39e061be426e6bf43e5d1427b17c884ff628861cc74fd3c815db7cd

Observation 0066eba5-fccc-4e59-a488-e583db03f4d7 · outbound

This paper cites Segment Anything.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Segment Anything

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.204372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.204372Z digest=sha256:ed99ac5286509fd311753f9e5b6209c701ed616002adf30f718330a19bec7c18

Observation a6bbe0b8-3650-4a0c-9f28-0bfdfeed0284 · outbound

This paper cites LISA: Reasoning Seg- mentation via Large Language Model.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning LISA: Reasoning Seg- mentation via Large Language Model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.607109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.209975Z digest=sha256:899a72c99a71dd162911328012cc0b2fc28764b3af1500e7ae767d7e20ec2152

Observation 635c4a1d-3db7-4901-aa9a-a9a9be8a92fa · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.215330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.215330Z digest=sha256:67f4fc8b1b938f7860ba87b48238a848d1fd2f8e4a3ecf497278d07b04d7f3f1

Observation e2ee7a45-9a72-42d3-ad53-d822732f5b3c · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.588928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.220888Z digest=sha256:b6013d8465cd9819b4997cdadd703bd045e831301ba7f0beea3e6c6ef87f32b6

Observation 4f29b02a-279f-4de1-bdd0-4a74d9057284 · outbound

This paper cites Grounded Language-Image Pre-Training.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Grounded Language-Image Pre-Training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.573052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.225228Z digest=sha256:23d77b4be4154c0502e15da899ed492fc5e599d826064d5b642bd6c2e103711b

Observation 7f14091b-929b-42ed-8698-4aa5b43c44fa · outbound

This paper cites Scallop: A Lan- guage for Neurosymbolic Programming.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Scallop: A Lan- guage for Neurosymbolic Programming

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.557753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.229629Z digest=sha256:af2bbaef8c8224c7b793db776927047822f504a5beb44a32557008797c67ffa8

Observation 16008e4b-fcd4-43e3-9578-ba5edafa03e9 · outbound

This paper cites GRES: Gen- eralized Referring Expression Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning GRES: Gen- eralized Referring Expression Segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.542239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.233935Z digest=sha256:f11bbe36053500a866bd998e4e7ddadb9f06557d3837568079e44a7b84d9a71e

Observation 7219750f-7efb-4088-aa01-9cdaaea59751 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Improved Baselines with Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.238320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.238320Z digest=sha256:2fbc02caa67af244b491a9f390f5e2a28029c4b8002e971aa2592415f2877ec0

Observation b465832c-150d-45c4-8ea4-17b23d782d85 · outbound

This paper cites Visual Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.242857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.242857Z digest=sha256:ffd445ee100d7931ed6677a2277aad31b3c08bc1d5480378abaefbe63c10dd7c

Observation 28ffe688-f29a-4dde-a1db-792d5e1e5c62 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.526504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.247511Z digest=sha256:7637a1d64fbc9f7acab2ddf32e9e283ae911ee13772b36111bb09650485681a5

Observation 8a0f35a8-5794-414c-979a-45d5f353eff8 · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.510559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.257759Z digest=sha256:68fdd71d641c7abf31aa51308f501ff34e1872dc9ed6c233da895af5ec816fe6

Observation 19f2f243-e1c1-4c2e-a7c1-cd2090574d27 · outbound

This paper cites Multi-Task Collabora- tive Network for Joint Referring Expression Comprehension and Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Multi-Task Collabora- tive Network for Joint Referring Expression Comprehension and Segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.493566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.262858Z digest=sha256:0c1e588d34d497611698d46c055c2ae0488675385278effc7f16e77791cf2f42

Observation cf47f353-1114-4f8f-9ed3-e3dd2cf67b9a · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.252109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.252109Z digest=sha256:bce586fec19df898089b0e91efb3306886d93ad677e3f7485bc4112b171af4fc

Observation ef2bf3cf-038f-42e6-9119-e48b06d418ca · outbound

This paper cites GPT-4o System Card, 2024.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning GPT-4o System Card, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.459981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.273010Z digest=sha256:76bf899c83fd7f0082033eef90ba2b7d16ebb9169b0cf7adce128d7d3df14230

Observation 18e9c71d-6314-422e-b5e6-5e9c7c7d05a3 · outbound

This paper cites Christiano, Jan Leike, and Ryan Lowe.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Christiano, Jan Leike, and Ryan Lowe

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.445602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.278013Z digest=sha256:683b4dc3c87e763452a62e772ebe5ef744b63a04d3966e2bdd8153167a99ebc5

Observation 636e7608-8c33-4e83-b740-3b8a787115f8 · outbound

This paper cites Scaling Open-V ocabulary Object Detection.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Scaling Open-V ocabulary Object Detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.476105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.267940Z digest=sha256:d46298da41ac38a26bc34d6c4405dfe9793eef2eba7caf0b18c2434a94b6d9da

Observation cfa9d043-281e-4300-a349-fcce7730aed4 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.288779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.288779Z digest=sha256:c5f11a126885548955d766186721adcf8996cd17f6acd8ce42871ad3fde0be86

Observation b438e301-e867-4968-9a69-cccb0519f80e · outbound

This paper cites PerceptionGPT: Effectively Fusing Visual Percep- tion into LLM.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning PerceptionGPT: Effectively Fusing Visual Percep- tion into LLM

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.415996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.294444Z digest=sha256:6f8f473126f098d37909f49c1bac95c0650cc61a5925160e67feb97974648d82

Observation ad6f9421-5caf-4284-981b-f694c4ba7320 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.431240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.283060Z digest=sha256:1053b54b237fa816fadacd3c7199a25d0dc0db4cf0dfb96b966b6cf8f892899e

Observation d0d9da8c-6868-4a37-8486-a8b9abca8f41 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.390115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.304626Z digest=sha256:a20600eeb9e1e6a1e7ec1ecf964fea43ed3ecc012ea08380fd28f0acb8747438

Observation 8ed197d9-5b23-4c00-bd0d-0a709d72d192 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.374381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.310038Z digest=sha256:81642de009d49d667182db872d7742c0ff60f580026887b3ce84b25a61c06fc6

Observation 158df1f1-bce1-4be1-95d7-52fd3a0d73e2 · outbound

This paper cites Language models are unsuper- vised multitask learners.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Language models are unsuper- vised multitask learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.299518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.299518Z digest=sha256:1515c7c433e3ff273e6f4e6375a8a14554e01c36619eaf33d57dc8700b7126dd

Observation 801fb7c1-26c1-49a1-b9b4-00660efe8695 · outbound

This paper cites ViperGPT: Visual Inference via Python Execution for Reasoning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning ViperGPT: Visual Inference via Python Execution for Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.326159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.326159Z digest=sha256:9dafa7bae0cec39509bd9092a253061ac88e5808c47689497a54fc436a3fcb0a

Observation b313a20f-a130-4ece-aa42-1db54810ae0b · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Gemma 2: Improving Open Language Models at a Practical Size

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.331855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.331855Z digest=sha256:947c7e87f90f90c0179ad78a9a208f6006ed9464f739bba3e1856c01acae4036

Observation 303c93aa-f55b-448a-9493-d46515b391de · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.315609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.315609Z digest=sha256:c812674c47d0b8c9fa1eac0bbf515b6ead9ae0ea1350cbab9fca21266c3d06c7

Observation c8d15d09-73c9-4158-b08b-53d52a851ef5 · outbound

This paper cites Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.358955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.321165Z digest=sha256:25d01ed699e8a7e77603270704e77f7ac204af27322636d41844538e9ef1421f

Observation 2dab5cce-0165-443d-b23a-46ecc6ae5b64 · outbound

This paper cites CRIS: CLIP- Driven Referring Image Segmentation.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning CRIS: CLIP- Driven Referring Image Segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.327510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.353958Z digest=sha256:9dbfb3bcddc9dd3b35c0a6f73d2913d8f320fb9f72ed3eddf220e74be23cc928

Observation 0828f7ac-e6ee-4f1f-88cd-065ac54e2dbe · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.358878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.358878Z digest=sha256:13a0eb4c901ec93edade68ff1e9ebc6ec7939f89f4a4f72a7efaea0cb10df3c5

Observation 4731627e-7fa4-4e11-a8e2-dc5e07e6b0df · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.337684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.337684Z digest=sha256:08fd707a6a99be2db7d3655d6981a3ac21a3e10c1ef9b7fa3f0b363fd7bea737

Observation de374cc4-7d5c-4d3e-83f8-e18810316348 · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.343366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.343222Z digest=sha256:89b423112cfd7e033fb1b6610066556c862fa787aea15e0cdd401afdc7ed986d

Observation 6aaae9ad-0aa3-4802-992e-9e3a7360e4e9 · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.348314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.348314Z digest=sha256:d4659ac42f1f28744b2ce44292a0de390df6d7479d629fd99eb68f88aae9b5b8

Observation a96f5036-d889-42c1-984c-4acd7ad1c729 · outbound

This paper cites an unresolved cited work.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.378332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.378332Z digest=sha256:a7eee455620e0f872ea3d8e18d975bcc99d509bb0b7524f5401d75ae59c6f233

Observation aa0e58d1-885b-422d-9749-ff4462c09c4f · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.382440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.382440Z digest=sha256:d8159539755494638c0537ff20ce007e6cf98c97e2160da6b22edc74813e3b72

Observation 871c4183-e5c8-4bfa-98bc-e6d6e3c02401 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.364061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.364061Z digest=sha256:fffcb05db32e48fac304261a46626ea3f260a42b7c06d0fbe45936e8089fe1f0

Observation f52294ea-dba9-488e-a57e-552a7d76b732 · outbound

This paper cites GSV A: Generalized Segmentation via Multimodal Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning GSV A: Generalized Segmentation via Multimodal Large Language Models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.311665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.369269Z digest=sha256:4bfe56a78d81a339982ac40cd0a53779bda942755641fd539f527dc71b4f43b8

Observation f4be7c81-dd03-4907-a27d-19caae15154b · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Va- riety of Vision Tasks.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Florence-2: Advancing a Unified Representation for a Va- riety of Vision Tasks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.190942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.373917Z digest=sha256:9637a4d7112bf269577d556c4753b0ff342348880ded93ebf00e58ed0ce1a551

Observation cf5111f1-ef1d-498b-90c9-e1d0bcbfcc34 · outbound

This paper cites IdealGPT: Iteratively Decomposing Vision and Lan- guage Reasoning via Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning IdealGPT: Iteratively Decomposing Vision and Lan- guage Reasoning via Large Language Models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.142095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.399485Z digest=sha256:4555b29bed9910017552d7fd21ad96ed8ce958153ccd1e00594292fc909e15db

Observation ee68c0b7-80ea-4f64-9e94-6db57a70e0eb · outbound

This paper cites Berg, and Tamara L.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Berg, and Tamara L

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.125998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.403779Z digest=sha256:29f5fc6a829a894d3e90494681ae117d9985d0f2438574fad6554a72b098f558

Observation 204ca23d-d05b-4f61-8bab-778710993bb4 · outbound

This paper cites Depth Anything V2.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Depth Anything V2

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.386798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.386798Z digest=sha256:03d08940b22c64c2facc9ccca0ebcbc36baac62191d65d0ebcd117b495b1126f

Observation 3f5e6032-e629-4330-afe5-ef8985c5386c · outbound

This paper cites Cross-Modal Rela- tionship Inference for Grounding Referring Expressions.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Cross-Modal Rela- tionship Inference for Grounding Referring Expressions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.174061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.391084Z digest=sha256:2ae185be1fa3b396f49fb72edd416da77bbd88c76ce48969eac2bdf444dfbf2a

Observation ea5fc35b-df3a-45ab-b4b1-d35a4ce2b789 · outbound

This paper cites an unresolved cited work.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:21:23.157455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.395185Z digest=sha256:5abd515e071943a462806e7e258af4027ef3add715095fd4ad59f5329dfda582

Observation 0fdb9299-4307-4693-9fd3-6e763cecfb25 · outbound

This paper cites Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.109629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.408130Z digest=sha256:894a94afddd42e62e92577ab1535ac16db57069cda76f89e8db13b8164f33c15

Observation c5662df4-091f-4c5e-a5f9-8449060109a6 · outbound

This paper cites PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:21:22.525450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.413504Z digest=sha256:fb6fc67bfc3bbc00031cef2d8ea3d7cacb03c0c1eec36b019a470142e90cc382

Observation 6fb6f62e-3ad3-4506-8ab4-010f3d7ca1c6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models,.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.418661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.418661Z digest=sha256:2e26de7578a29fb47a807af55545deada45ed2731cef234a95f1a7d014f10fe6

Observation c660ab1a-4fdf-409a-8e52-8013b95abe95 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.424064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.424064Z digest=sha256:bc27874251accc1ca71d84c16c90872ab771a7470d84c417c18146fc8873e5aa

Observation 0ad299b5-34e6-4a09-97a9-3db4dbf26d03 · outbound

This paper cites Table 9 presents a quantitative comparison on the RefCOCO, RefCOCO+, RefCOCOg [21], and Ref- Adv [1] datasets.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Table 9 presents a quantitative comparison on the RefCOCO, RefCOCO+, RefCOCOg [21], and Ref- Adv [1] datasets

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.082917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.429171Z digest=sha256:407c20d637d08f7e7171e75197ee594b984508e67bc81030131a733f3fecf128

Observation 9b871822-dd7c-4635-ab03-3c39b278c443 · outbound

This paper cites For a fair comparison, we use the same foundation models and LLM for the ex- periments.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning For a fair comparison, we use the same foundation models and LLM for the ex- periments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.067105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.434128Z digest=sha256:864b99ee8ce05f8a06e6413831a5387ec5f9e52958c28f272a33d1ea6334fa4c

Observation 781c9b0b-4f0d-4c1e-a214-78337446f75f · outbound

This paper cites The results are shown in Table 11.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning The results are shown in Table 11

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.050044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.439036Z digest=sha256:19155f05eaad4b956fd0a3793cb4b98e4cb82c9ec307cd3bcd99fab9040e2051

Observation 504a224c-bc61-4fc1-8b62-96165ff670e5 · outbound

This paper cites In this mechanism, each time the system transitions into the self-correction state (indicated by a red arrow in Figure 2), it is counted as one retry.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning In this mechanism, each time the system transitions into the self-correction state (indicated by a red arrow in Figure 2), it is counted as one retry

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.034758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.444127Z digest=sha256:668f424ea749c4340dd7c73f397a0aef2599b4aa1b05d72595aa4b1c815e4752

Observation 53f673b3-c505-4da8-94d8-5b95494c63aa · outbound

This paper cites The results, shown in Figure 4, indicate that NA VER consistently reaches SoTA performance compared to all baselines regardless of query length.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning The results, shown in Figure 4, indicate that NA VER consistently reaches SoTA performance compared to all baselines regardless of query length

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.018330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.449051Z digest=sha256:d344a22104545547a9fb9c62f731bf2e13add212b81914ef5bf446f0babf68cf

Observation 3ce02b20-ac36-4fb9-a623-5a0cfb14e61f · outbound

This paper cites An intuitive alternative is to skip captioning and ask the VLM to predict categories directly.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning An intuitive alternative is to skip captioning and ask the VLM to predict categories directly

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.002013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.454378Z digest=sha256:609c9165b26ab4595497e3a08b496bf9266c4ee57e5f234910517890661da794

Observation 91e6a8e7-9401-490a-b0c7-454915489a28 · outbound

This paper cites Yes” or “No.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning Yes” or “No

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:22.985311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.459302Z digest=sha256:2401a9a77fa796a4e255b9c004d9fcf89342a7e33417a03d58a7eadc6584e708

Observation c2fe5022-c14a-4810-bd90-a991e5487999 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.151438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.151438Z digest=sha256:003d96ff682c0bb8199e84b54927c2a3954991a2c7e65fbfdf4cea131ccac5a3

Observation bcdb2266-fb07-4797-9585-249b93d864fc · outbound

This paper cites 2, 3, 4, 6, 7, 1.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning 2, 3, 4, 6, 7, 1

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:21:23.624357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T19:21:22.199123Z digest=sha256:46825fa3938a6d934794e50587c2f165c1a74ca25b37f6ced8fda4ca02b62f2f

Pith citing papers

Observation d45b8d54-c1a7-4f87-a5ad-e3061a8e04f4 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:09:19.196908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:09:18.020873Z digest=sha256:31a80306c533dea097fe6104c4efc7b5c99f2f081b613f77e90f68adef17a9bf