Pith. sign in

Paper Citation Record · LEDGER

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

As of 19 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2507.00505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00505 v3

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:45.204350Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:34:40.626626Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:34:40.824777Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cee1f6ce-6a32-4bdf-9b3c-7352849c29ac · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.629953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:43.427456Z digest=sha256:be589c579925643585124166f5c8bf57170bd4d4570d4f26757331084e030e7b

Observation ab0cc359-29f9-4889-99a6-c5525bbf7929 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.475383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.475383Z digest=sha256:990ab5c29f84e6e6b3c4726c2338fdc02c215915a951e9e4c904879c96adee31

Observation 6f2721f9-a9a1-483f-88cb-f0159f6ec301 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.611398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:43.517516Z digest=sha256:33a7d4086d04d63d6ee9eeb39625cd48af72ab60e96182e24ae191f63a1d611c

Observation b55466f2-e18b-4f99-a487-c8b24a190331 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.593338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:43.600087Z digest=sha256:d5f3f23757440159ed896b6ab03d43ed6ddaeb4a28dab9e16ccbd3ab999aa555

Observation e98cef87-098b-4b09-9d65-a8faa89b346f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.631540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.631540Z digest=sha256:896ead34d5c3111f174b612684e380d997675b002c5aa0059a6c9c8519c3708e

Observation 0c9fe8de-7a77-45bd-baba-7d4913a8f81b · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.712897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.712897Z digest=sha256:a3c43610e547a2201d58cab06580d5b53e9f60aaa9fd03fda86a28d6e8ec6778

Observation 545f009c-c476-415b-8c41-ec954171c74e · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.915177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.915177Z digest=sha256:a6d2957e903da7aecb5c8c245fe070617bc2929e44a26e7672a30ffe96e0f697

Observation d1ada282-759c-40ec-8784-3aa7265ac3ae · outbound

This paper cites Uniter: Universal image-text representation learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Uniter: Universal image-text representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.574843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.031994Z digest=sha256:a88ac173baa62cc1d67feb528498f6ad959a39ee3ba0ad16e87ce072a0440195

Observation e9145929-a1dc-4553-a4aa-6727d9e6cca7 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.115572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.115572Z digest=sha256:38be002f2efef0d6b88c8f793cb7343661194ddfa4f5d66156d0ad20745ac6a5

Observation 6d7ef2a5-c6f2-4970-ae26-e0d07112cdc5 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.555610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.197563Z digest=sha256:497b6a03a3a57c972825a144c0af803967b16faccef9b438c2fa01f5122251cd

Observation 6e9bb0d1-a4e3-412e-b890-99f880d9b9c6 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.537078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.290139Z digest=sha256:eb607f61712f29eb53d9eff862ea23df8adc637b2a6523f7838e5ba34c95db1d

Observation 3aa9277e-10de-47ab-b680-eb9f8412d345 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.519939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.369599Z digest=sha256:989234251f46c8b9a6b80500cc87cac42dac5590fdf4e3196dce1ee416d0828c

Observation e7305b86-daf9-40fe-b4aa-958b5de44d65 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.483409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.483409Z digest=sha256:b528a074f20edeea2f8cc42caf09826eb68e4e1f71d34f4d74ec5ef6e30cda8e

Observation 68489f04-2607-4e17-b14b-c4c59d5c6de7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.585083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.585083Z digest=sha256:4e7100e6406fdba1b94be034c012f67e959b1636abf967df48eccef10805807c

Observation d73db33c-a918-4358-a229-4f0b58ebd1e4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.704161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.704161Z digest=sha256:37c5ca1ba7464391595d5be81d002aa8940bfa09e3a188744b163da36fe591e1

Observation e67fd48f-ed16-4bfd-b86f-e8dd213de02e · outbound

This paper cites Mmgpt4lf: Leverag- ing an optimized pre-trained gpt-2 model with multi-modal cross-attention for load forecasting.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mmgpt4lf: Leverag- ing an optimized pre-trained gpt-2 model with multi-modal cross-attention for load forecasting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.493689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.807213Z digest=sha256:26e7d700a83e6c75026bcdfbea3800e5daf97fb1eebf1ca60e63c61d2ae7bac8

Observation e03fe239-a1a0-40f4-8940-fbe98b498ecd · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.476285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.881633Z digest=sha256:b88c42fcdaf6f0579820408ef21b3b1a092707d1a2536e0ad4294b26314ac477

Observation 2e250998-a5bd-4a2b-846f-6d6d09f20097 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vizwiz grand challenge: Answering visual questions from blind people

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.457966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.896147Z digest=sha256:73013a0ae6de0cde3517849829c93d283b3658ca02a62abf345d3e0f7436ae42

Observation 97b4e0f1-3f29-4140-b356-8d4085e4f5d3 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.920866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.920866Z digest=sha256:a18f5266751ee5e456c6f6aea07837b0554152b7b453378f5b01cf2df6d4e17b

Observation 39b07aac-03de-48a1-9614-df0622cafa0d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.926088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.926088Z digest=sha256:148db709072397ceb324db7c9107c87893c6a29eb03e77d7600d9a7cd8d9f66c

Observation 37f40682-d614-42b4-a53a-4d2462c1cde0 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.931122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.931122Z digest=sha256:17a233531316cbbd351440a72d416245edba7db62aa0a29b5072d8725b134c10

Observation a125aabf-069c-468e-893a-6409cbfc2219 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.936712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.936712Z digest=sha256:7cc663b6ec8f2392743c5eca5fee410b0aa93fb98aa37cec2cd690805abeb65f

Observation 4c4ce55f-2ae4-497d-8c5c-be6f910a51c7 · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Dvqa: Understanding data visualizations via ques- tion answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.427769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.941653Z digest=sha256:e727997aeab5e8ac6e20afb0a1dba5e4c859856d16884a11bdd53d79803aaf41

Observation 7e2dc6ef-960f-4711-9de2-38504f54325f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.946588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.946588Z digest=sha256:c4e779a61ccd49abc8f3aeaa0ee0e8327eb9b53f29d6d644a232cd956dcb056e

Observation 1a7db1ba-c02c-4de9-9908-3c8a97f2f384 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs OtterHD: A High-Resolution Multi-modality Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.950880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.950880Z digest=sha256:eed738930b1d75f8800f2208d0ad428dc4e1b38bd2474f1ad0b71df9ed5f3405

Observation ee97ffc7-ea7d-484f-9daf-344ab183e9ce · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.411783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.956249Z digest=sha256:42b3713e2236e3c0df0b5d23fb9cf2115a49448491ab872144126507047ea0a2

Observation fec20d40-0c96-4c3c-8727-edda88af5379 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.395048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.961149Z digest=sha256:d0080306f49c7146fdf358c0f73db4844f8ebcf56b06ddf47fb6521a2b1321a1

Observation adbb46d0-1439-4dc9-ac89-57bb7b0e8210 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.965646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.965646Z digest=sha256:40096be9e0025bd7ce5c3588481984df8041236e6beab1132224dcf962f5ed26

Observation 92d33b10-36b3-44dc-8822-a3e46aab0ad4 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.970173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.970173Z digest=sha256:1a8ace628d2aa5b1336f2a962ff50f57485c78ad0c502b45d2c71d1c505ce05f

Observation a0e70f54-e15c-4080-9aa0-952c91d5f73a · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.376466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.975214Z digest=sha256:762e150a44ed39c3d15ef5b196052da6dc5691fc8ee7897c50b3a5b51735a852

Observation 359df87f-915d-4925-aed3-013de3b65b0c · outbound

This paper cites Vila: On pre-training for visual language models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vila: On pre-training for visual language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.359343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:44.979454Z digest=sha256:5d7d64a21573c5707d18588d58be7904300e830c84b19286826a49ecfe7bd654

Observation 5c43d5bd-080b-44d8-a57f-8defc6663b94 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.983616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.983616Z digest=sha256:91e2a5ace435f2ab1f07ae4eae487770a62e471bcbccb2758f4279bcaf7c0d03

Observation 9665645e-4cba-4403-8551-2dfccc20952f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Improved Baselines with Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.987878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.987878Z digest=sha256:609323d3cb09ec9695f9e07af7c815720468b91de83a3aff18c69f7861637d4d

Observation 77762b0e-6226-4098-ab4a-7f3de606e4fb · outbound

This paper cites Visual instruction tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.992726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.992726Z digest=sha256:2b30b5a770acbc12da725f0a79cef459f1c07d755fea882fb00eb93ea61d0e79

Observation 54ee26ed-be62-4a20-8a4c-3d777a7ae95f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.997621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.997621Z digest=sha256:f5eef0e9cb5c7cc607874d9efa6e473756a8ed828f5e13e3d2fcaaeb340d233b

Observation 68c43b42-5a1a-49f0-bf9d-63210c86902f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.321340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.002680Z digest=sha256:ef9413f4bb37d6caa1e1a362beada287fab2f1d14b9af840f40bd6cdc0a7eefe

Observation 2edb4b9e-99df-45ef-adb5-42ba2f571515 · outbound

This paper cites A convnet for the 2020s.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs A convnet for the 2020s

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.303530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.008240Z digest=sha256:4c66cd24f25cae3346923c7c30e1a6a33da77280c28608dde3c46fbbed94e0db

Observation d0ef25ee-4b4b-4d37-aa7c-ec6ec9e42007 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.287850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.013156Z digest=sha256:5e568bf1eb3d326099cc6708ab51bd651ed507a533c9e4cac51940dde7183990

Observation 0870f533-a5ce-4bec-b390-43a3a9d4e9ab · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.018551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.018551Z digest=sha256:c176693e9e19517aac27baae0f1112bee089fc54b5955af464414d46ae2da234

Observation 2f1baa1c-0008-485a-91cf-4a9730393dd8 · outbound

This paper cites Unirgb-ir: A unified frame- work for rgb-infrared semantic tasks via adapter tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unirgb-ir: A unified frame- work for rgb-infrared semantic tasks via adapter tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.024693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.024693Z digest=sha256:3b464519a9ba778db026278575d0286df604d3b83020047af4b2dd5d4e4a63b4

Observation 4c90dc6b-1bd6-49ab-b71e-951cbf27c62f · outbound

This paper cites Gpt-4v(ision) system card, 2023.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gpt-4v(ision) system card, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.029990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.029990Z digest=sha256:9ee6af282ccefef08e8138a92f8bf53cc4e502e5175e6eeb1f19bc369dbd0a69

Observation 10f53c0c-c4a9-4b9e-8b31-45b6b3d077d9 · outbound

This paper cites an unresolved cited work.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.035094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.035094Z digest=sha256:eeae20b2689cf9e2acab87a3b2f13d9403b85c4e77ec7aa1428ab4810f2b9e5f

Observation 61817f82-debb-4b6f-9f30-5a9c7a65cf73 · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Im2text: Describing images using 1 million captioned pho- tographs

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.245228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.040957Z digest=sha256:2937acc3d7f836b94e1a80bd74aa86f416d1d4af44e441eb63d4ef308b934b6d

Observation 86824750-a82c-400a-90c0-6109603038fc · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Learn- ing transferable visual models from natural language super- vision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.229834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.045946Z digest=sha256:01ffaa6fe65ad65c02109415c417202906db0fbb2ba848b95fb9cda81d623edb

Observation 42a8aa21-46f4-4598-a6f7-94c37158d116 · outbound

This paper cites Introducing our multimodal models, 2023.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Introducing our multimodal models, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.213938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.051203Z digest=sha256:683e2810fe8a337bde20b2df9fbfd79abf9fea4551204817d9406dc686e22ef9

Observation b1a630e2-8570-472b-a26a-bea2785aed4c · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.056171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.056171Z digest=sha256:313edfef24cf5deeb277129e37904dedc2bb1f0b35fcf99ce4b085af341753e5

Observation 0a7bcfc6-44a7-4e98-ac0f-e6e35e9920ad · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.062070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.062070Z digest=sha256:46f21aca8e155b45398928bf3cd35b8cc69a7fa20d0b1c048df4cce2e2a0da85

Observation 301a441b-d411-4a07-a674-d5140293ea0d · outbound

This paper cites Towards vqa models that can read.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Towards vqa models that can read

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.196495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.067048Z digest=sha256:c47bb79f35cc300b16e82e8835ad98edad2ba778410311cb41f0ef8743075614

Observation 77389585-b3f9-4a47-ab97-aa12d7948fa4 · outbound

This paper cites Towards vqa models that can read.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Towards vqa models that can read

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.178948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.073006Z digest=sha256:c81393ee557b53f97511b783498c808b6e705b112793686225a85760da84f3a7

Observation 3b8bd4df-b998-43a8-a499-7dc662571fe8 · outbound

This paper cites Adapool: Expo- nential adaptive pooling for information-retaining downsam- pling.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Adapool: Expo- nential adaptive pooling for information-retaining downsam- pling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.079975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.079975Z digest=sha256:4aea86033fb13666fa5e180705657c2eb3f404fe28596533e9334bb73efe5208

Observation bbc6f390-ff3c-4a79-84d7-ea45e7c07116 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.086031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.086031Z digest=sha256:206fb5f2aaa7c90e989cd3c8dcdfa56c29aa0f6818eda5037d0672b7b950067f

Observation 72462e34-fbb2-48d0-9d92-2f6d9989153a · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.091175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.091175Z digest=sha256:07df0215cbad40bd5b894fdf998e723bdf934dd556aa0f3977e82691c5a16180

Observation 44557196-22e8-4c08-8f0d-2e3ff4fb4a5b · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.149443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.097816Z digest=sha256:4076b155eede00e4cce9bde1e946531d0ff9332e698628d310060e191444f5dd

Observation 38bd5eed-883d-4378-aabc-c748c6c41fcc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.104103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.104103Z digest=sha256:e5525e6d4ae9350052595948984a41a28c45965860c0c1632dcdf005e343a012

Observation d99a8b76-1fc4-4252-8b4a-f455e17a2046 · outbound

This paper cites Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.132019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.109266Z digest=sha256:e4e75ca475f49a57e26389a564534bdc8cc83df6dbba90566b2430670726c670

Observation 04f360b7-f5ba-4a69-b6cf-8a1077010caa · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.114851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.114851Z digest=sha256:928132e538ef4c083a324a0e145cee3ed4dd70337a8edf91dbeb4fafd0d148be

Observation fca24d84-cd1c-417a-bb59-ff56e4f0d028 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs xgen-mm (blip-3): A family of open large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.121380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.121380Z digest=sha256:1f4fb8c0c11675d90a8880872cb607d34c126cb706fc839365cb60a6d2f49f67

Observation b53642ed-20c3-4215-833c-5837dc8c421f · outbound

This paper cites Dense Connector for MLLMs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Dense Connector for MLLMs

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:45.429257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.127437Z digest=sha256:532b2cab111eefe32159f0110d1eb354bcf6be3951f16dc675dff56f405d7805

Observation 5be07de4-f1f7-4d71-8631-80c31b637cec · outbound

This paper cites Image difference cap- tioning with pre-training and contrastive learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Image difference cap- tioning with pre-training and contrastive learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.115108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.133736Z digest=sha256:bd6771ad66c8c954b24f32b07a62ca9739f53ed2cf7b7d28b19b91c695143143

Observation 34b26c7e-af54-4113-a51c-b0978ddb19d5 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.139225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.139225Z digest=sha256:b5dade7c7e6f87dc1f53735a89103d69139534d939ed72930fd92ed641efea2a

Observation 6b84d7fd-9ebe-4e04-ad98-20fcbb1b05a8 · outbound

This paper cites Modeling context in referring expres- sions.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Modeling context in referring expres- sions

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.097275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.145435Z digest=sha256:c74e4c52f0268655af3bd505ac3ebc64f782335b834620d6b773fc45671174be

Observation 4591d180-2c6b-4798-99d0-f88e1091ee2a · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.150548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.150548Z digest=sha256:e5800a6a87b57a4095948053e5ac21fc0bb30064e251da910fdace0359b0a200

Observation 8041b73d-0350-4191-a166-d342e42d7523 · outbound

This paper cites Improving rgb-infrared object detection with cascade alignment-guided transformer.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Improving rgb-infrared object detection with cascade alignment-guided transformer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.079474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.155968Z digest=sha256:1587a0c3a207d7e6004fac313a130b6db8ac8b51e6fb0e8d3da3b37f1afc073d

Observation cd0d889e-0e3b-40f4-bc79-bbe21341c7ae · outbound

This paper cites Lane detection transformer based on multi- frame horizontal and vertical attention and visual trans- former module.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Lane detection transformer based on multi- frame horizontal and vertical attention and visual trans- former module

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.062261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.161184Z digest=sha256:48a5cbf2fa78e2e44def253e578146555323cbe91f4173414ae9eb1e8fbe6192

Observation b7cadc4d-47b2-4570-921e-e97df1d5db9c · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.166758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.166758Z digest=sha256:f36f1ce24f777841b98aa89381e7b02afe84c1a104f9be308f0c3371c9f74e33

Observation 22f5ab8f-4ecb-46d3-b237-dc67a5255b2c · outbound

This paper cites Removal and selection: Improving rgb- infrared object detection via coarse-to-fine fusion.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Removal and selection: Improving rgb- infrared object detection via coarse-to-fine fusion

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.173798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.173798Z digest=sha256:f4b007e248bdc1605378f6913eee6001655587e0ac08b4b679983f06492cc4e8

Observation b728e6cb-ce85-4178-9119-4c47f55440c3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.179796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.179796Z digest=sha256:afcf7a6474df9ba7809f14f83dfc63f7d09d219f6d1ee4af518b31f73d7855fd

Observation 504e8ba0-fff7-4e4e-b664-abed4832f663 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.185864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.185864Z digest=sha256:e45e1f5c1275612c0c93dd7e49c8d88a508979e17432ffba6116560762dba4d1

Observation 4fa3c775-5141-4b1b-a531-3693d22ca1c3 · outbound

This paper cites Future work will involve experiments on various LLMs such as Qwen2.5, Mistral, and LLaMA3.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Future work will involve experiments on various LLMs such as Qwen2.5, Mistral, and LLaMA3

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.027455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.198962Z digest=sha256:56627c8479d660cfed6ee511f33559db165e1ec0f582b15dec200e8ffd45d137

Observation a08b0375-8b33-4c30-a8b9-e9dcc87f04fc · outbound

This paper cites The input and output channels of the convolutional kernels are 1024 (equal to the visual feature dimension output by ViT) and 512, respectively.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs The input and output channels of the convolutional kernels are 1024 (equal to the visual feature dimension output by ViT) and 512, respectively

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.010281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.204350Z digest=sha256:0bf5741f645dba161c088c28e2ada0b478bac51608aebb70247a3a59c965823b

Observation 04e26855-5c68-4bb6-8c77-55131c8235a2 · outbound

This paper cites CLIP-bind pairs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs CLIP-bind pairs

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.044601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T21:19:45.192425Z digest=sha256:d9f57faefa59d73d53f4f6bd5efae8eeed1a95c821a55bc20c8b42886865e4f4

Pith citing papers

Observation 321d912f-f4ff-479b-98ee-ac67f5f417cd · inbound

STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding cites this paper.

STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:34:40.830381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T00:34:40.626626Z digest=sha256:982b9450d9c02f0885a7bba97bd6048041b0f54998b2f7e63194dd08730f4fe2