Pith. sign in

Paper Citation Record · LEDGER

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2507.00505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00505 v3

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:45.204350Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cee1f6ce-6a32-4bdf-9b3c-7352849c29ac · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.629953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:43.427456Z digest=sha256:1c073dc6246b2bea24a1a36c53c17f49ed26be4d8dd373482a8038bdc68fad37

Observation ab0cc359-29f9-4889-99a6-c5525bbf7929 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.475383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.475383Z digest=sha256:2c0149a83b6ede82f2e15902b8734a109968f3d5bf96adc7769ba79bc2ca0818

Observation 6f2721f9-a9a1-483f-88cb-f0159f6ec301 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.611398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:43.517516Z digest=sha256:6a78cd83396458b6ff64578de7bfef392ed05524de9fc4990613a0c6d1c5549d

Observation b55466f2-e18b-4f99-a487-c8b24a190331 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.593338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:43.600087Z digest=sha256:b96b9f7b2053d7f8196f4c9c219d2c8ee0527291e475b6117ce026ef5b9f7d3b

Observation e98cef87-098b-4b09-9d65-a8faa89b346f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.631540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.631540Z digest=sha256:bdcbd80036880607cc0d433ef0c5536b0be8997391c8eca5508fda7bbdb40166

Observation 0c9fe8de-7a77-45bd-baba-7d4913a8f81b · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.712897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.712897Z digest=sha256:ca595fd18e8ca801ed6e37e7d10d8cbc0c630544158e1261bdebe656c422be7e

Observation 545f009c-c476-415b-8c41-ec954171c74e · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:43.915177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:43.915177Z digest=sha256:a7afad6981450b3e1285c3e175702de253cc35b716c037ff94d4364a64f42268

Observation d1ada282-759c-40ec-8784-3aa7265ac3ae · outbound

This paper cites Uniter: Universal image-text representation learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Uniter: Universal image-text representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.574843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.031994Z digest=sha256:46295a70a092bda17d0d31c4d65ec6c8204b654444e212e33cb48d36b8da6513

Observation e9145929-a1dc-4553-a4aa-6727d9e6cca7 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.115572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.115572Z digest=sha256:1d4cc09271d8f370568826ed87bef140f7c8d50de3c4304aa55a8ec2e94e1d78

Observation 6d7ef2a5-c6f2-4970-ae26-e0d07112cdc5 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.555610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.197563Z digest=sha256:da04127492236398df86687ac18f91c93898ee845200b12f5c0b60275d442566

Observation 6e9bb0d1-a4e3-412e-b890-99f880d9b9c6 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.537078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.290139Z digest=sha256:f452735b482001becd9064bb58bf9454f82feab9a83d1a3efcbf143faf67ada1

Observation 3aa9277e-10de-47ab-b680-eb9f8412d345 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.519939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.369599Z digest=sha256:736b619dc86518380453c2883b547145e8b196e2630ac1a8abe760382f60684b

Observation e7305b86-daf9-40fe-b4aa-958b5de44d65 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.483409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.483409Z digest=sha256:c87f2fd34a4adad2ca67bd6423e053812d3a6294bdf795bae17ce53e86b7ee8d

Observation 68489f04-2607-4e17-b14b-c4c59d5c6de7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.585083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.585083Z digest=sha256:ed0bb0c67e9e6cbab6b747fdf002261c7506cf066370075d9ebbc989d57e8bbe

Observation d73db33c-a918-4358-a229-4f0b58ebd1e4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.704161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.704161Z digest=sha256:1bbf0737892300b82c424582c07a50513a21f7f45a3634f87433dc74a870bfba

Observation e67fd48f-ed16-4bfd-b86f-e8dd213de02e · outbound

This paper cites Mmgpt4lf: Leverag- ing an optimized pre-trained gpt-2 model with multi-modal cross-attention for load forecasting.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mmgpt4lf: Leverag- ing an optimized pre-trained gpt-2 model with multi-modal cross-attention for load forecasting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.493689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.807213Z digest=sha256:8f88a319dfd0cc25674061660e78bac4925423cef46793af6c83dbaeac3bc7cd

Observation e03fe239-a1a0-40f4-8940-fbe98b498ecd · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.476285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.881633Z digest=sha256:9c956c136bc297b8de9343da252a45551a9ebd6b766af8a6ae6716c07fe73334

Observation 2e250998-a5bd-4a2b-846f-6d6d09f20097 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vizwiz grand challenge: Answering visual questions from blind people

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.457966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.896147Z digest=sha256:2b2268b1308038bb205988242c6e7afb04a185b415a715dfc74dbd30a5e5a708

Observation 97b4e0f1-3f29-4140-b356-8d4085e4f5d3 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.920866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.920866Z digest=sha256:b4267f40f2acd3148f2328f96d3ecaba503a5ac19a3e15f4618c799060ab0728

Observation 39b07aac-03de-48a1-9614-df0622cafa0d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.926088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.926088Z digest=sha256:4735c939b82d1772a19eaa18b53815c0c04f2c600325fe03f090abf49cabda9a

Observation 37f40682-d614-42b4-a53a-4d2462c1cde0 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.931122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.931122Z digest=sha256:aedbfbb82a6736fa41a8008796b1fbccc7cac585190f96e205c51e90e4da0fef

Observation a125aabf-069c-468e-893a-6409cbfc2219 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.936712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.936712Z digest=sha256:fb160fbc6331ce2d7b8749b0958cfa785edb5c155d7a9492bdc0936484894f62

Observation 4c4ce55f-2ae4-497d-8c5c-be6f910a51c7 · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Dvqa: Understanding data visualizations via ques- tion answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.427769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.941653Z digest=sha256:b7b6f5685824534c8dfb4b0685fb7f53445400792f7b284097ae225656f128a7

Observation 7e2dc6ef-960f-4711-9de2-38504f54325f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.946588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.946588Z digest=sha256:d2b1d4745e35a3d9067fedf5ccfb787a2db7a04db8178f19af1c617c0a8ebcb3

Observation 1a7db1ba-c02c-4de9-9908-3c8a97f2f384 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs OtterHD: A High-Resolution Multi-modality Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.950880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.950880Z digest=sha256:902230dfed5372aedb774ce02380df8913c178babe2dbcddaad646ff6c75110e

Observation ee97ffc7-ea7d-484f-9daf-344ab183e9ce · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.411783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.956249Z digest=sha256:8045fcede825b81bbe3164d6de56726057e70eed442bace9567233d9fd0898ff

Observation fec20d40-0c96-4c3c-8727-edda88af5379 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.395048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.961149Z digest=sha256:62fb0ecc2d88d647e97839926c32c3ac3b74501b47e27d9890de58c6a4222fad

Observation adbb46d0-1439-4dc9-ac89-57bb7b0e8210 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.965646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.965646Z digest=sha256:45891e588916c8a69f5a62036e899e57fb4318e196f4ce5c7ae7e21a726aad36

Observation 92d33b10-36b3-44dc-8822-a3e46aab0ad4 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.970173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.970173Z digest=sha256:feb17906d12385d1152a7662580434c19c53fe2fb28c068cb46509c44d45a170

Observation a0e70f54-e15c-4080-9aa0-952c91d5f73a · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.376466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.975214Z digest=sha256:de12e17e9d84a60e5a132ebf15804139d255255b5fc35b1d0b718c4d2b913541

Observation 359df87f-915d-4925-aed3-013de3b65b0c · outbound

This paper cites Vila: On pre-training for visual language models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vila: On pre-training for visual language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.359343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:44.979454Z digest=sha256:edf501d77f4b3e6af33c87287e72c361d197344f178fb07ccacccbc541e5a66c

Observation 5c43d5bd-080b-44d8-a57f-8defc6663b94 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.983616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.983616Z digest=sha256:883525ac781aaec094557fa5c584b405996dd8a8cc2eb764d49440596d21e3a6

Observation 9665645e-4cba-4403-8551-2dfccc20952f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Improved Baselines with Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.987878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.987878Z digest=sha256:e0025844901b4d50649fc7c63773141731e244da1279d7635a04a10d2d2402df

Observation 77762b0e-6226-4098-ab4a-7f3de606e4fb · outbound

This paper cites Visual instruction tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.992726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.992726Z digest=sha256:46c38e7c19b628dc2a5689a487d89ba83aa724747218288d64ee954fc24c2a55

Observation 54ee26ed-be62-4a20-8a4c-3d777a7ae95f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.997621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.997621Z digest=sha256:2f097fe71f09f1425e12c109da5596c9d9502f47d69c21d949a6b66516d78c76

Observation 68c43b42-5a1a-49f0-bf9d-63210c86902f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.321340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.002680Z digest=sha256:b376096e1e4f71e4e2a3789bb5962cd19493a666beaad80abcfa13584896cb55

Observation 2edb4b9e-99df-45ef-adb5-42ba2f571515 · outbound

This paper cites A convnet for the 2020s.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs A convnet for the 2020s

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.303530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.008240Z digest=sha256:7df5fe76521174e2fa6482265c35b2d2860f03a14500d104d23ed7bcd96317b8

Observation d0ef25ee-4b4b-4d37-aa7c-ec6ec9e42007 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.287850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.013156Z digest=sha256:539717ef0d2dcaad72c451f651b1195694f8307026ca787013f9279736447e14

Observation 0870f533-a5ce-4bec-b390-43a3a9d4e9ab · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.018551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.018551Z digest=sha256:dc8a8fe7dfe6162602ecbd70d7e0462715819ad82d6842eba47c8b671960eb20

Observation 2f1baa1c-0008-485a-91cf-4a9730393dd8 · outbound

This paper cites Unirgb-ir: A unified frame- work for rgb-infrared semantic tasks via adapter tuning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unirgb-ir: A unified frame- work for rgb-infrared semantic tasks via adapter tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.024693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.024693Z digest=sha256:355d4c86658e584a8ca0f13c4c12d2ceccc033d68b8f13284b79c783e0c5f9f3

Observation 4c90dc6b-1bd6-49ab-b71e-951cbf27c62f · outbound

This paper cites Gpt-4v(ision) system card, 2023.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gpt-4v(ision) system card, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.029990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.029990Z digest=sha256:0bdc7908fe2b7a08bb7eb018b9461574d3c80f2419abf57d492449074e46a15e

Observation 10f53c0c-c4a9-4b9e-8b31-45b6b3d077d9 · outbound

This paper cites an unresolved cited work.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.035094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.035094Z digest=sha256:b253970719a7217072c49a7af8d54f7e86ac0eb5b728ed486ded7041ff9b3529

Observation 61817f82-debb-4b6f-9f30-5a9c7a65cf73 · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Im2text: Describing images using 1 million captioned pho- tographs

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.245228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.040957Z digest=sha256:9d23c6c05033b148f88cad7d2dd172fc6868240ff644b1a236b27a33f6442634

Observation 86824750-a82c-400a-90c0-6109603038fc · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Learn- ing transferable visual models from natural language super- vision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.229834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.045946Z digest=sha256:bceef66c630dbe1ae559ed4c342967662d411a627f421bbc1de9425ddd10e165

Observation 42a8aa21-46f4-4598-a6f7-94c37158d116 · outbound

This paper cites Introducing our multimodal models, 2023.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Introducing our multimodal models, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.213938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.051203Z digest=sha256:9919977baf96ebe179c8a68098648d84bc4fb52bc508611693506fd10c602ea9

Observation b1a630e2-8570-472b-a26a-bea2785aed4c · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.056171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.056171Z digest=sha256:c17f50319735985baf360489482d095406e9c8c528dd81a38fa6f971c47f4288

Observation 0a7bcfc6-44a7-4e98-ac0f-e6e35e9920ad · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.062070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.062070Z digest=sha256:6f838b831848c4bb04f30dbe4ad4d1837acb9f709856e5bcc14afa1d802c52da

Observation 301a441b-d411-4a07-a674-d5140293ea0d · outbound

This paper cites Towards vqa models that can read.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Towards vqa models that can read

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.196495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.067048Z digest=sha256:2ac6268becc2a1427c8e9b75d9ceb06115c4ba653284854af120b86c834657a7

Observation 77389585-b3f9-4a47-ab97-aa12d7948fa4 · outbound

This paper cites Towards vqa models that can read.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Towards vqa models that can read

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.178948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.073006Z digest=sha256:95dbe9a5f3e64d7dc22572f2a7e689ccbbbce2bf53a9dabd8ec12945858a3297

Observation 3b8bd4df-b998-43a8-a499-7dc662571fe8 · outbound

This paper cites Adapool: Expo- nential adaptive pooling for information-retaining downsam- pling.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Adapool: Expo- nential adaptive pooling for information-retaining downsam- pling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.079975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.079975Z digest=sha256:eab7bd40c74a800d05e8be8e288dba4200716bac34266ca5f024b3c56ad4d7e9

Observation bbc6f390-ff3c-4a79-84d7-ea45e7c07116 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.086031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.086031Z digest=sha256:880fdd8f28b8e290fa3d2bb8be12963777a1cbde74d113d645f6004fc1776790

Observation 72462e34-fbb2-48d0-9d92-2f6d9989153a · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.091175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.091175Z digest=sha256:2993260f2094eddaec4c6881ec275c5ea8c9eaec9f32fad811e11bdb719f114e

Observation 44557196-22e8-4c08-8f0d-2e3ff4fb4a5b · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.149443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.097816Z digest=sha256:969f09b83f51a1d3303ceb32af974de5830ada551b8307a4058fc2e2941387ee

Observation 38bd5eed-883d-4378-aabc-c748c6c41fcc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.104103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.104103Z digest=sha256:66856f71620e774fc2c3b18dc9251b6c8b5a011f402c793f2e692c17aef9b1f1

Observation d99a8b76-1fc4-4252-8b4a-f455e17a2046 · outbound

This paper cites Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.132019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.109266Z digest=sha256:64134935c5cc25eab01d0c81e817acb73ada4bb5e9a306a507ed1e74c0a08778

Observation 04f360b7-f5ba-4a69-b6cf-8a1077010caa · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.114851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.114851Z digest=sha256:e03d4b7d1dbf94b1922f06c08da5b2958428ef82b1928ec9bae53b1ecc9e01fe

Observation fca24d84-cd1c-417a-bb59-ff56e4f0d028 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs xgen-mm (blip-3): A family of open large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.121380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.121380Z digest=sha256:cda478b3a3a8d765bd10bd29744e80aedc0e5fb8e5afa4d79b99047a17b8c797

Observation b53642ed-20c3-4215-833c-5837dc8c421f · outbound

This paper cites Dense Connector for MLLMs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Dense Connector for MLLMs

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:45.429257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.127437Z digest=sha256:2c7c88e2cff7d70858a9257325ee1dc7904d0a9492d739af07ad71464d35208e

Observation 5be07de4-f1f7-4d71-8631-80c31b637cec · outbound

This paper cites Image difference cap- tioning with pre-training and contrastive learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Image difference cap- tioning with pre-training and contrastive learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.115108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.133736Z digest=sha256:0e7b6e023167caf009b1c9ed098f293978b13f5e346b8306877c5c56e3bc8f0b

Observation 34b26c7e-af54-4113-a51c-b0978ddb19d5 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.139225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.139225Z digest=sha256:b41180059efb4d20621898b1dcd0173f636a2efaf30c4dcb022f95135a5e922a

Observation 6b84d7fd-9ebe-4e04-ad98-20fcbb1b05a8 · outbound

This paper cites Modeling context in referring expres- sions.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Modeling context in referring expres- sions

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.097275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.145435Z digest=sha256:c827604f81997668d89f50d3a33c1ef86a3afc145d78704af36b08c91e03f919

Observation 4591d180-2c6b-4798-99d0-f88e1091ee2a · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.150548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.150548Z digest=sha256:dc28590422f637383fecb6b45ce3dec2b7e8b4d3ae1d0864096b70122e64171f

Observation 8041b73d-0350-4191-a166-d342e42d7523 · outbound

This paper cites Improving rgb-infrared object detection with cascade alignment-guided transformer.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Improving rgb-infrared object detection with cascade alignment-guided transformer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.079474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.155968Z digest=sha256:d6284df8bacd374f025af6845a38a205c9d69b6d57a282e51fd64164062c1ca2

Observation cd0d889e-0e3b-40f4-bc79-bbe21341c7ae · outbound

This paper cites Lane detection transformer based on multi- frame horizontal and vertical attention and visual trans- former module.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Lane detection transformer based on multi- frame horizontal and vertical attention and visual trans- former module

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.062261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.161184Z digest=sha256:849a7a42f771ad059807007e806e33f78c22cbcb5f867c5f2e75524dd4bedc3c

Observation b7cadc4d-47b2-4570-921e-e97df1d5db9c · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.166758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.166758Z digest=sha256:be101ed03a853981f509902b2fcbbd2cfdeebbe2e2e81247a5df90ea54f783aa

Observation 22f5ab8f-4ecb-46d3-b237-dc67a5255b2c · outbound

This paper cites Removal and selection: Improving rgb- infrared object detection via coarse-to-fine fusion.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Removal and selection: Improving rgb- infrared object detection via coarse-to-fine fusion

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.173798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.173798Z digest=sha256:c141eccf77ba683dad9dd0584445d3d439810dd9bc590fdb3ad777303fe35b83

Observation b728e6cb-ce85-4178-9119-4c47f55440c3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.179796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.179796Z digest=sha256:b5f7ebbe6cb3215cd7f7c208278dc75c5bf4482194a95efabe34c4c263c8283c

Observation 504e8ba0-fff7-4e4e-b664-abed4832f663 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.185864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.185864Z digest=sha256:c7be48ae55ca8c9529385f57132ca19228205695d56b3d666c328c0ad3c08b82

Observation 4fa3c775-5141-4b1b-a531-3693d22ca1c3 · outbound

This paper cites Future work will involve experiments on various LLMs such as Qwen2.5, Mistral, and LLaMA3.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Future work will involve experiments on various LLMs such as Qwen2.5, Mistral, and LLaMA3

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.027455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.198962Z digest=sha256:5f5cbb22652ade89f7d0d032dbb53951b527bfe4a675d4971ef08880f5180a1e

Observation a08b0375-8b33-4c30-a8b9-e9dcc87f04fc · outbound

This paper cites The input and output channels of the convolutional kernels are 1024 (equal to the visual feature dimension output by ViT) and 512, respectively.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs The input and output channels of the convolutional kernels are 1024 (equal to the visual feature dimension output by ViT) and 512, respectively

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.010281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.204350Z digest=sha256:27d5572c880c6423a53a895910469c282e42af24f80d1bc84bd4850118a727ab

Observation 04e26855-5c68-4bb6-8c77-55131c8235a2 · outbound

This paper cites CLIP-bind pairs.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs CLIP-bind pairs

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:46.044601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:19:45.192425Z digest=sha256:d50bf164223c1b6cee5e0a39d291a2006317adbc5ceb679e90ca660404656a22

Pith citing papers

No inbound Pith citation observations are available.