Pith. sign in

Paper Citation Record · LEDGER

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2505.12194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12194 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:42:01.350617Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b4be10b-1382-44bd-b4af-51394726bd5b · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.199814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.099614Z digest=sha256:9f3a88afd1708ee2f36b3d212acbfc5da125e6fd47e5e4197f105db88ac079c2

Observation da35ada0-b447-4d57-b7e9-3dacc66b8a36 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.104538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.104538Z digest=sha256:99cf8a4f188f5ce469a5c9eeeee0c4a7a3c1631019b8d28b0c69395640312a41

Observation ed9d6ede-b858-48e3-b7f7-dbce13430eea · outbound

This paper cites Paligemma: A versatile 3b vlm for transfer, 2024.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Paligemma: A versatile 3b vlm for transfer, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.171991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.108823Z digest=sha256:bb826f2aebbc90f6af1ea7e04a8a61fba405652067cfeee6ac8f4da71363d828

Observation 6acb85eb-f62b-4b90-b140-5ed25d4dc631 · outbound

This paper cites Meyer, Yuning Chai, and Yong Jae Lee.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Meyer, Yuning Chai, and Yong Jae Lee

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.155473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.113068Z digest=sha256:0516b0ae67c37a469e3b28c6ee3271480c6930c705558ee70a307bd6e85aae83

Observation 6d3a2c34-fd32-4b0c-a84d-2f92adad4e50 · outbound

This paper cites End-to-End Object Detection with Transformers.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding End-to-End Object Detection with Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.117530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.117530Z digest=sha256:da4bf22c7f138e40549cfbcc4ebbb6101f11dcd5826bab32debfc44c0cfb3e18

Observation 87d26fe7-bffa-416e-ab89-53b1cee56357 · outbound

This paper cites Chang, and Matthias Nießner.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Chang, and Matthias Nießner

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.140911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.122887Z digest=sha256:f9314d0af0782849f9018b9e02e0a431c2a92d0792fe123a82d7719a2d02c181

Observation f83ae1f4-d6ec-4ce9-a251-543cbc481efd · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Reproducible scaling laws for contrastive language-image learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.122862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.127302Z digest=sha256:8286f3d4884214467ab532ccaf07db0d26d4cfaf11b3970a14a7e9962424217e

Observation 39df7291-0c35-4e9b-8dad-793358ac81fd · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.131352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.131352Z digest=sha256:95b50bc962cf9645dc2c7c68d65eb7ac079caa9a133bbb0d615c3831f1cc91e5

Observation 762ea977-ff9d-4b23-8948-1f2264653156 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.135730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.135730Z digest=sha256:b5c3bd0ea45382d00417445254dacc2e3ca33e387a10ee74492d0b389b8eba45

Observation f3c3f399-48b3-460c-b636-d75fea4626d4 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Scaling Instruction-Finetuned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.140159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.140159Z digest=sha256:56f3f6ad58125254034a356522a709407f88c61f81224ac2a1e5f9eaa60406db

Observation f8130732-5a40-464b-a7ca-5f1466da59d3 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.145285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.145285Z digest=sha256:c097cf3154f56b4339b6fa3ea50f25a2a5b0d8dc72cd6441b11dbc8f69a23f85

Observation c45436c2-9cb1-4049-a484-6e16fa69a7d2 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Qlora: Efficient finetuning of quantized llms, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.079283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.149222Z digest=sha256:b27a3c1356fa07d0b502f4cd8dd35310289dbe84f6cd0da08c56bc924059587a

Observation 37ccbc78-2b94-4772-9e69-0970e51f125a · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding, 10 2018.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Bert: Pre-training of deep bidirectional transformers for language understanding, 10 2018

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.062339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.153728Z digest=sha256:0b99b1783387f362d48655f5007474ca051b9a244a5b3771dca4fce2989c827c

Observation c9f7267a-0eec-42da-bda3-570e8bd66d9e · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale, 12 2022.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Eva: Exploring the limits of masked visual representation learning at scale, 12 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.042675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.157500Z digest=sha256:ec352ba383dd2d82c21e0d22932d511bf05d0aeb6f14a1cab7e0fc0168ca47b4

Observation cf39f242-ab4c-42d8-a1da-c4bfbcc6a177 · outbound

This paper cites Sebastian Borgeaud.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sebastian Borgeaud

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.024059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.161837Z digest=sha256:31746c545412c198a9be94069a71c9b565adf827755ea68d636e77e6cd64e1f3

Observation b58ed854-ed16-400c-9f2d-071448d86b51 · outbound

This paper cites Arzen-llm: Code-switched egyptian arabic-english translation and speech recognition using llms.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Arzen-llm: Code-switched egyptian arabic-english translation and speech recognition using llms

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.007589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.165947Z digest=sha256:4903bcda69b5575dcec6d691fbfbe1037ce73f324a605aba5cc23c1811241a1b

Observation a5bc6b3b-bfbb-46e5-be73-e7cf9a4bf894 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Lawrence Zitnick, and Ross Girshick

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.984777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.170276Z digest=sha256:9042ae370f2a69aeac68352717ac130e6399592957d007dd002d10b8b7b3d937

Observation 1e6b80dc-c27f-46a2-9f2c-b95a4af42833 · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.174246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.174246Z digest=sha256:cedd1ce0cdd295570f07dd82b05ef88c299c3a11589de83ef589a172fe0d2c40

Observation cdb5dbec-3f04-4245-9e98-ac561733b61f · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Referitgame: Referring to objects in photographs of natural scenes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.178917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.178917Z digest=sha256:ee5fcd364501b00816432c2e8e7be2be3f48c3ce577ced5ecba8f96a9601bb73

Observation 3927bc2f-8dbe-4b33-a61c-38541ce2b2d6 · outbound

This paper cites Computational genera- tion of referring expressions: A survey.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Computational genera- tion of referring expressions: A survey

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.956685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.183124Z digest=sha256:858a14e4dfa58689336e7776c58dbe57eb743a5bdd1ddbed3910d4202db579ff

Observation 50d03fe8-9096-468a-9ad8-5caeda702d2e · outbound

This paper cites When to retrieve: Teaching llms to utilize information retrieval effectively, 05 2024.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding When to retrieve: Teaching llms to utilize information retrieval effectively, 05 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.940399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.187349Z digest=sha256:ff739035f6b9ae82d71d05b5aa553771421b74528696a6fb50a84bbd1e899a7d

Observation 5c85146a-578c-46bf-bcea-0d30960c5668 · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.926351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.194495Z digest=sha256:8153248b2c4e2ceabed5ab384807a4c15eb9697a492bbc804e75a831dfd466a2

Observation 27718a0e-9a3d-4103-a71c-5aeb284cd0f7 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.201033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.201033Z digest=sha256:2bb8247469aabdfc2ac6ca4dcd038a88d6aacfc70226dbaee045511609ff8eb8

Observation 5f3c152d-a8c4-4ad4-94e0-c62a001d1e4c · outbound

This paper cites Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.902440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.205674Z digest=sha256:a983cebc2876045b158d7d650244fad7411f5e2dc7a7338e0ce2296257d57b81

Observation d69a8043-cef0-40f3-93c5-14e17c9fe60a · outbound

This paper cites Lawrence Zitnick, and Piotr Doll´ar.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Lawrence Zitnick, and Piotr Doll´ar

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.884073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.210634Z digest=sha256:4edcf2b65a572432d0c628e7a4c24eab7e54410f2a1828330ffbec1e012f7fe6

Observation ff3f4944-11d7-4a4e-9e26-782abbd2c978 · outbound

This paper cites Visual spatial reasoning, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Visual spatial reasoning, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.215353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.215353Z digest=sha256:9294091a72c226309042b33905e86111d9242700dbb6c57722a2bc9928245bc7

Observation 2e54f4f4-558c-4328-af62-5476b10f019b · outbound

This paper cites Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.860306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.220484Z digest=sha256:d92f108fa6a1985c805a0b7070ce16ff8f56949c19b82659f212dd224deae53a

Observation 378b87a3-6f1a-429d-8baf-4ab029211aa2 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Improved baselines with visual instruction tuning, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.227573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.227573Z digest=sha256:5cae467dc6b506be1ac241a4fe1707740cfbd89ab3b13e1af94546babdab6516

Observation 3e2f3754-f5cc-4591-bcb2-dcc28e77cb6b · outbound

This paper cites Visual instruction tuning.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.232461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.232461Z digest=sha256:c8b6b198fb805a2944ca022b02320034a5e13fd5225d17d21366d523ecb1d98f

Observation 2ad51824-6f5f-4c72-9c8a-aec9b6d663c0 · outbound

This paper cites Clevr- ref+: Diagnosing visual reasoning with referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Clevr- ref+: Diagnosing visual reasoning with referring expressions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.822384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.237438Z digest=sha256:7fb5ce641a1fd7183baeede18719cc642d23ca2fcd37797d46b7efc5bff09b23

Observation abaaab58-51d3-4651-a77c-354ec57b547f · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach, 2019.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Roberta: A robustly optimized bert pretraining approach, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.805674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.243529Z digest=sha256:1a6dd3c32dca8b82aa84c9fdcd635cc19c4f1d4838c4e7a0fbd9a940d12295cd

Observation 7891a45e-b672-4cea-9da9-f46cbd335944 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding MMBench: Is Your Multi-modal Model an All-around Player?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.248047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.248047Z digest=sha256:ecd424dcf89c7d60648c900c657039db1f9ac595f320fad769778d860b03998d

Observation f2dd3b37-32f9-4cbf-adc4-c732769c27d4 · outbound

This paper cites An embarrassingly simple approach for llm with strong asr capacity.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding An embarrassingly simple approach for llm with strong asr capacity

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.789663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.254755Z digest=sha256:0d119abfcf8d7ab630bffa74dc4a7eab28c220060a121595f00c47cd33b40b49

Observation 14052c44-5cfb-4b4e-bfb2-aa3f316cb49a · outbound

This paper cites Generation and comprehension of unambiguous object descriptions, 2016.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Generation and comprehension of unambiguous object descriptions, 2016

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.774052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.259634Z digest=sha256:109a51f3f511ec7305dab572845cc18c9faea28346ec809785262a4ea455ef2e

Observation 8f8cf1a7-e58c-4095-8fa7-3a664ea5b488 · outbound

This paper cites Sun-spot: An rgb-d dataset with spatial referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sun-spot: An rgb-d dataset with spatial referring expressions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.759669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.263919Z digest=sha256:0d54abc1e88c3d81f4c4195b85f78acde5fcf67b8db401704b80a5cae966e4d4

Observation 6212bf1c-060d-4f35-8ddf-b9d89230f993 · outbound

This paper cites Domain terminology integration into machine translation: Leveraging large language models, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Domain terminology integration into machine translation: Leveraging large language models, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.744988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.267895Z digest=sha256:3e475b59152d2af197e69fd307357f90f44d5c70aa25e61fc35271e2c8d2c212

Observation 172f7d6a-b936-4f74-98e3-f8da32adf00f · outbound

This paper cites MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.272339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.272339Z digest=sha256:e9902100b9133c72b6baa915c7bdcfce6e59331522e3c39b809733bc09d6513e

Observation a287a7dd-1a97-4e49-97df-e94acd94d29e · outbound

This paper cites GPT-4 Technical Report.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding GPT-4 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.278138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.278138Z digest=sha256:e13b159ff6af0ba1eb83d80513b3934be0f75a316cc99c7014058e3e8e0ac6e8

Observation 9ef06a3e-801d-43e5-bbec-4cab472d661c · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environments, 2020.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Reverie: Remote embodied visual referring expression in real indoor environments, 2020

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.725428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.283316Z digest=sha256:f2462fa1344827f810d45684c442632fb914a426389ff7dc0ce14326aa8b7f30

Observation aa93edda-840a-48b5-a640-a9d3aa962d3b · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.288751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.288751Z digest=sha256:7f544806dcca794edf84ecad7fd57633c37154425dc0c07dc1a1959fc59c292f

Observation 5dac8927-2b91-471b-8687-8912fdb530b0 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.293525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.293525Z digest=sha256:05cee83f96e11d59c021cca47f0f61f75464951f19abce560cf32346cca8c05f

Observation 1e05419f-41f1-41e1-b87b-2e32960c3ae2 · outbound

This paper cites Eva- clip: Improved training techniques for clip at scale, 03 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Eva- clip: Improved training techniques for clip at scale, 03 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.693187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.298020Z digest=sha256:e466662dc39c975dd5e98852d6cd0e7fa7caa73fac74b5ffd06d5c3f3a3bef62

Observation ba500ec0-b31a-4592-9f9c-68fd8781acca · outbound

This paper cites Self-retrieval: Building an information retrieval system with one large language model.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Self-retrieval: Building an information retrieval system with one large language model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.674285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.303168Z digest=sha256:091447606a1ad1bc297ae745f267db58c7403f2b93a748f1b636c70afbe3a298

Observation 7cc07cf8-ec25-4113-a282-d956bbc8a967 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.311814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.311814Z digest=sha256:7d6997c4fc1ffe50e5351c1555006b3707086b9aa0d64952ae448801d9e0d63b

Observation f5505bb0-e851-4ef3-a301-8ecbf0ea6c15 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.654685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.318526Z digest=sha256:bd7cf05bc5b92e2c6b8c98802d3a860bd6011900cbce6d370cf7a551437cdb5f

Observation 3f3ad556-9c77-41fb-b97d-86ec82214b33 · outbound

This paper cites The generation of natural descrip- tions: corpus-based investigations of referring expressions in visual domains.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding The generation of natural descrip- tions: corpus-based investigations of referring expressions in visual domains

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.638951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.322797Z digest=sha256:dcaa6b070fa1746e63083a449d0fc97e7e74b6e6be1e50bb31aa0a21c14e3331

Observation 90b8c13a-315c-4806-982e-dacf13fa1c13 · outbound

This paper cites Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.327938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.327938Z digest=sha256:87baa605ee263ee46bdef08344662c58569784fa3ffd4179ff4a0ec89300baa3

Observation 2d377aad-26fb-44de-89cb-11dafc29afbe · outbound

This paper cites Xlnet: Generalized autoregressive pretraining for language understanding, 2019.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Xlnet: Generalized autoregressive pretraining for language understanding, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.609997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.331978Z digest=sha256:b4790f8ee6b1922f3574ce36f65de977145582f62a448c3743a6045fb199829c

Observation 200d4cde-550a-45b7-a4d3-9273d800c70f · outbound

This paper cites Modeling context in referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Modeling context in referring expressions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.595606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.336192Z digest=sha256:6c8ca88249616ceed4b2c0f5606835b73a78e261506dbb226de91042dc076c1e

Observation 0361d342-2114-4f40-aa67-82e7f06e30b6 · outbound

This paper cites A joint speaker-listener-reinforcer model for referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding A joint speaker-listener-reinforcer model for referring expressions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.579028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.340627Z digest=sha256:212fe7dca239955ce3b7b7c9289cff601f27431e783e263183f4b367c4fd245d

Observation 6ceb79ce-3570-4b12-aa65-c996ba87e7e2 · outbound

This paper cites Prompting large language model for machine translation: A case study, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Prompting large language model for machine translation: A case study, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.563240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.345926Z digest=sha256:6a36ec539435574ef83bea56b73ef47ea947cce9a230e5ef1de843713e0ddafb

Observation d2ff4c5a-c55a-4eb0-bf61-eb5f79fce442 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Toolqa: A dataset for llm question answering with external tools

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.549350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:42:01.350617Z digest=sha256:ab8d3374ba474dcfb410b8eb672f7b99263ab79fe45169d2a64197b1f91b9489

Pith citing papers

No inbound Pith citation observations are available.