Pith. sign in

Paper Citation Record · LEDGER

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding

As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2505.12194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12194 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:42:01.350617Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b4be10b-1382-44bd-b4af-51394726bd5b · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.199814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.099614Z digest=sha256:4fb29b73a41701bf13ec55f3f5ef90443c68ca03dafbdbae1505149f2831c4d9

Observation da35ada0-b447-4d57-b7e9-3dacc66b8a36 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.104538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.104538Z digest=sha256:99cf8a4f188f5ce469a5c9eeeee0c4a7a3c1631019b8d28b0c69395640312a41

Observation ed9d6ede-b858-48e3-b7f7-dbce13430eea · outbound

This paper cites Paligemma: A versatile 3b vlm for transfer, 2024.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Paligemma: A versatile 3b vlm for transfer, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.171991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.108823Z digest=sha256:6fb6a309c4de160d144e8b4e13daf43d1cc9830c0ca0a37b70a0b238a634129e

Observation 6acb85eb-f62b-4b90-b140-5ed25d4dc631 · outbound

This paper cites Meyer, Yuning Chai, and Yong Jae Lee.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Meyer, Yuning Chai, and Yong Jae Lee

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.155473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.113068Z digest=sha256:8a22eed964de7afb1f43fbbe886452cafe82269714e297d304f02cbfbf9c08e1

Observation 6d3a2c34-fd32-4b0c-a84d-2f92adad4e50 · outbound

This paper cites End-to-End Object Detection with Transformers.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding End-to-End Object Detection with Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.117530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.117530Z digest=sha256:da4bf22c7f138e40549cfbcc4ebbb6101f11dcd5826bab32debfc44c0cfb3e18

Observation 87d26fe7-bffa-416e-ab89-53b1cee56357 · outbound

This paper cites Chang, and Matthias Nießner.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Chang, and Matthias Nießner

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.140911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.122887Z digest=sha256:2de88a95920391f881db749f5d5e4fd5cdc74742c1a72ccace46aa6462376570

Observation f83ae1f4-d6ec-4ce9-a251-543cbc481efd · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Reproducible scaling laws for contrastive language-image learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.122862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.127302Z digest=sha256:ace228696bf82e9e5e4f37a446b20a5675dc7a3f72e4c0953f103bdab279a786

Observation 39df7291-0c35-4e9b-8dad-793358ac81fd · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.131352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.131352Z digest=sha256:95b50bc962cf9645dc2c7c68d65eb7ac079caa9a133bbb0d615c3831f1cc91e5

Observation 762ea977-ff9d-4b23-8948-1f2264653156 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.135730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.135730Z digest=sha256:b5c3bd0ea45382d00417445254dacc2e3ca33e387a10ee74492d0b389b8eba45

Observation f3c3f399-48b3-460c-b636-d75fea4626d4 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Scaling Instruction-Finetuned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.140159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.140159Z digest=sha256:56f3f6ad58125254034a356522a709407f88c61f81224ac2a1e5f9eaa60406db

Observation f8130732-5a40-464b-a7ca-5f1466da59d3 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.145285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.145285Z digest=sha256:c097cf3154f56b4339b6fa3ea50f25a2a5b0d8dc72cd6441b11dbc8f69a23f85

Observation c45436c2-9cb1-4049-a484-6e16fa69a7d2 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Qlora: Efficient finetuning of quantized llms, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.079283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.149222Z digest=sha256:6d99b34f002a587bf2461d76ca11cd457648d2102b01368382dfdfd987cf46a3

Observation 37ccbc78-2b94-4772-9e69-0970e51f125a · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding, 10 2018.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Bert: Pre-training of deep bidirectional transformers for language understanding, 10 2018

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.062339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.153728Z digest=sha256:c423efa5d04f8d6b8b242d6323bd09ef69289b1691d3c102421e15377a8751fc

Observation c9f7267a-0eec-42da-bda3-570e8bd66d9e · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale, 12 2022.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Eva: Exploring the limits of masked visual representation learning at scale, 12 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.042675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.157500Z digest=sha256:70822afe5eb04c97ac483eeb05c86c675fefd6b8a7fb180831eef941193eaefa

Observation cf39f242-ab4c-42d8-a1da-c4bfbcc6a177 · outbound

This paper cites Sebastian Borgeaud.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sebastian Borgeaud

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.024059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.161837Z digest=sha256:ec706c19443ce656cebd9a30bbe9b3ccf8555251187490383340b1a93ec864a0

Observation b58ed854-ed16-400c-9f2d-071448d86b51 · outbound

This paper cites Arzen-llm: Code-switched egyptian arabic-english translation and speech recognition using llms.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Arzen-llm: Code-switched egyptian arabic-english translation and speech recognition using llms

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:02.007589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.165947Z digest=sha256:0c6dd691f4b6dd3d2dd084fe7ba47b04626420b3e8d70e26ace44e76262b0fd2

Observation a5bc6b3b-bfbb-46e5-be73-e7cf9a4bf894 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Lawrence Zitnick, and Ross Girshick

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.984777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.170276Z digest=sha256:fe004da4922ffc2022cb3d823b7062b9bbf527426e3806367696bafc35db6928

Observation 1e6b80dc-c27f-46a2-9f2c-b95a4af42833 · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.174246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.174246Z digest=sha256:cedd1ce0cdd295570f07dd82b05ef88c299c3a11589de83ef589a172fe0d2c40

Observation cdb5dbec-3f04-4245-9e98-ac561733b61f · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Referitgame: Referring to objects in photographs of natural scenes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.178917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.178917Z digest=sha256:ee5fcd364501b00816432c2e8e7be2be3f48c3ce577ced5ecba8f96a9601bb73

Observation 3927bc2f-8dbe-4b33-a61c-38541ce2b2d6 · outbound

This paper cites Computational genera- tion of referring expressions: A survey.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Computational genera- tion of referring expressions: A survey

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.956685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.183124Z digest=sha256:ee5b4256aa1de05efcfb02fb0e72ca2b18fe6a8e9331edfc4c4dc6deacaf94c7

Observation 50d03fe8-9096-468a-9ad8-5caeda702d2e · outbound

This paper cites When to retrieve: Teaching llms to utilize information retrieval effectively, 05 2024.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding When to retrieve: Teaching llms to utilize information retrieval effectively, 05 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.940399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.187349Z digest=sha256:a60cebf7e0b71105c23d5c35641308fb9e4dee8be4b00408879b73eda6164a77

Observation 5c85146a-578c-46bf-bcea-0d30960c5668 · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.926351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.194495Z digest=sha256:b2340f80666bb49953a7d73fa34a12c9283a461014296955cc8db176c06dcbd0

Observation 27718a0e-9a3d-4103-a71c-5aeb284cd0f7 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.201033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.201033Z digest=sha256:2bb8247469aabdfc2ac6ca4dcd038a88d6aacfc70226dbaee045511609ff8eb8

Observation 5f3c152d-a8c4-4ad4-94e0-c62a001d1e4c · outbound

This paper cites Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.902440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.205674Z digest=sha256:adca3ffd62a662a144a4c264f8766ae5d1b95432a0b6e4f054937453d9446d1e

Observation d69a8043-cef0-40f3-93c5-14e17c9fe60a · outbound

This paper cites Lawrence Zitnick, and Piotr Doll´ar.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Lawrence Zitnick, and Piotr Doll´ar

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.884073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.210634Z digest=sha256:686d4518da4419f2798e2e179b247eab7f081780375e5ac66b6183d09a471013

Observation ff3f4944-11d7-4a4e-9e26-782abbd2c978 · outbound

This paper cites Visual spatial reasoning, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Visual spatial reasoning, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.215353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.215353Z digest=sha256:9294091a72c226309042b33905e86111d9242700dbb6c57722a2bc9928245bc7

Observation 2e54f4f4-558c-4328-af62-5476b10f019b · outbound

This paper cites Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.860306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.220484Z digest=sha256:56b7379974f023c11b7a7b89022d95552fd0536dbc200f3c0ace1f16ce811652

Observation 378b87a3-6f1a-429d-8baf-4ab029211aa2 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Improved baselines with visual instruction tuning, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.227573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.227573Z digest=sha256:5cae467dc6b506be1ac241a4fe1707740cfbd89ab3b13e1af94546babdab6516

Observation 3e2f3754-f5cc-4591-bcb2-dcc28e77cb6b · outbound

This paper cites Visual instruction tuning.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.232461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.232461Z digest=sha256:c8b6b198fb805a2944ca022b02320034a5e13fd5225d17d21366d523ecb1d98f

Observation 2ad51824-6f5f-4c72-9c8a-aec9b6d663c0 · outbound

This paper cites Clevr- ref+: Diagnosing visual reasoning with referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Clevr- ref+: Diagnosing visual reasoning with referring expressions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.822384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.237438Z digest=sha256:a14116c8aac0c099e4b42de9d7aa400e030132b19ff129806e4a75ee6060e05c

Observation abaaab58-51d3-4651-a77c-354ec57b547f · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach, 2019.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Roberta: A robustly optimized bert pretraining approach, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.805674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.243529Z digest=sha256:3e865a7f6ba7665aba87dc1d4c5213779aea1ccb4ae7950eaeed7b9cf522d78d

Observation 7891a45e-b672-4cea-9da9-f46cbd335944 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding MMBench: Is Your Multi-modal Model an All-around Player?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.248047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.248047Z digest=sha256:ecd424dcf89c7d60648c900c657039db1f9ac595f320fad769778d860b03998d

Observation f2dd3b37-32f9-4cbf-adc4-c732769c27d4 · outbound

This paper cites An embarrassingly simple approach for llm with strong asr capacity.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding An embarrassingly simple approach for llm with strong asr capacity

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.789663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.254755Z digest=sha256:2e1d5790e6fe2dff114cd14fc7f473929839535fcc42185c49c84ea86fb9cbf3

Observation 14052c44-5cfb-4b4e-bfb2-aa3f316cb49a · outbound

This paper cites Generation and comprehension of unambiguous object descriptions, 2016.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Generation and comprehension of unambiguous object descriptions, 2016

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.774052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.259634Z digest=sha256:b21a962171b40cfb0847b8c472a704897fac814c7535fd693739ff794454239e

Observation 8f8cf1a7-e58c-4095-8fa7-3a664ea5b488 · outbound

This paper cites Sun-spot: An rgb-d dataset with spatial referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sun-spot: An rgb-d dataset with spatial referring expressions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.759669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.263919Z digest=sha256:74e1689f67e55d7a9080f575d82155be28933759d5f9cd961a3cc9cf27d7bc97

Observation 6212bf1c-060d-4f35-8ddf-b9d89230f993 · outbound

This paper cites Domain terminology integration into machine translation: Leveraging large language models, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Domain terminology integration into machine translation: Leveraging large language models, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.744988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.267895Z digest=sha256:a7536abf99e46f2df4cf52c649e638bb8742a67ce20c48bb1ac53a38b74c92ef

Observation 172f7d6a-b936-4f74-98e3-f8da32adf00f · outbound

This paper cites MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.272339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.272339Z digest=sha256:e9902100b9133c72b6baa915c7bdcfce6e59331522e3c39b809733bc09d6513e

Observation a287a7dd-1a97-4e49-97df-e94acd94d29e · outbound

This paper cites GPT-4 Technical Report.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding GPT-4 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.278138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.278138Z digest=sha256:e13b159ff6af0ba1eb83d80513b3934be0f75a316cc99c7014058e3e8e0ac6e8

Observation 9ef06a3e-801d-43e5-bbec-4cab472d661c · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environments, 2020.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Reverie: Remote embodied visual referring expression in real indoor environments, 2020

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.725428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.283316Z digest=sha256:7fa4942b6a9ac2d5ab4975cfaf0054919ac09ec810e240055b488865583d7f43

Observation aa93edda-840a-48b5-a640-a9d3aa962d3b · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.288751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.288751Z digest=sha256:7f544806dcca794edf84ecad7fd57633c37154425dc0c07dc1a1959fc59c292f

Observation 5dac8927-2b91-471b-8687-8912fdb530b0 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.293525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.293525Z digest=sha256:05cee83f96e11d59c021cca47f0f61f75464951f19abce560cf32346cca8c05f

Observation 1e05419f-41f1-41e1-b87b-2e32960c3ae2 · outbound

This paper cites Eva- clip: Improved training techniques for clip at scale, 03 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Eva- clip: Improved training techniques for clip at scale, 03 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.693187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.298020Z digest=sha256:326c906f6f49a68b3e601fdc92990af9f148570693bcacc2abecf0152b8cf418

Observation ba500ec0-b31a-4592-9f9c-68fd8781acca · outbound

This paper cites Self-retrieval: Building an information retrieval system with one large language model.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Self-retrieval: Building an information retrieval system with one large language model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.674285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.303168Z digest=sha256:41595c75ca84d337c0f48dfa9ff3f68cc39b1ed71042835cf2635d1f3be1599b

Observation 7cc07cf8-ec25-4113-a282-d956bbc8a967 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.311814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.311814Z digest=sha256:7d6997c4fc1ffe50e5351c1555006b3707086b9aa0d64952ae448801d9e0d63b

Observation f5505bb0-e851-4ef3-a301-8ecbf0ea6c15 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.654685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.318526Z digest=sha256:9c3b3973bd8306ac6152245d9f7b184b32eab350d12871d889e171b08f88e277

Observation 3f3ad556-9c77-41fb-b97d-86ec82214b33 · outbound

This paper cites The generation of natural descrip- tions: corpus-based investigations of referring expressions in visual domains.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding The generation of natural descrip- tions: corpus-based investigations of referring expressions in visual domains

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.638951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.322797Z digest=sha256:89c6729bcb1e4a9821ec0a9593418a62bbe6c88c2408efe9cef84a3f04f6d08d

Observation 90b8c13a-315c-4806-982e-dacf13fa1c13 · outbound

This paper cites Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:01.327938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:01.327938Z digest=sha256:87baa605ee263ee46bdef08344662c58569784fa3ffd4179ff4a0ec89300baa3

Observation 2d377aad-26fb-44de-89cb-11dafc29afbe · outbound

This paper cites Xlnet: Generalized autoregressive pretraining for language understanding, 2019.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Xlnet: Generalized autoregressive pretraining for language understanding, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.609997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.331978Z digest=sha256:8d311d2f9a79a9f279bcc6bec7e6807bf7fdf96b9ba218cdf9f839f1cc194483

Observation 200d4cde-550a-45b7-a4d3-9273d800c70f · outbound

This paper cites Modeling context in referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Modeling context in referring expressions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.595606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.336192Z digest=sha256:902802a1ceda10a72c2110c9eafa55969ebe2bb76de0d9a138da8ecf588ee974

Observation 0361d342-2114-4f40-aa67-82e7f06e30b6 · outbound

This paper cites A joint speaker-listener-reinforcer model for referring expressions.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding A joint speaker-listener-reinforcer model for referring expressions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.579028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.340627Z digest=sha256:79039dbce6e715274477488f95c3f09e10d75db52b8c96de179a3655f6ae9986

Observation 6ceb79ce-3570-4b12-aa65-c996ba87e7e2 · outbound

This paper cites Prompting large language model for machine translation: A case study, 2023.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Prompting large language model for machine translation: A case study, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.563240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.345926Z digest=sha256:21e41838bddee275eef1d93cdcdc457414b0c1601b120fadf8124d68dd488a5a

Observation d2ff4c5a-c55a-4eb0-bf61-eb5f79fce442 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Toolqa: A dataset for llm question answering with external tools

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:42:01.549350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:42:01.350617Z digest=sha256:02e4c0a6ded37adab36d10ce32c09b8e6842b5e983ff159640f54c03437c738e

Pith citing papers

No inbound Pith citation observations are available.