Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:45.204350Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2507.00505.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:45.204350Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cee1f6ce-6a32-4bdf-9b3c-7352849c29ac · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab0cc359-29f9-4889-99a6-c5525bbf7929 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2721f9-a9a1-483f-88cb-f0159f6ec301 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Honeybee: Locality-enhanced projector for multimodal llm
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b55466f2-e18b-4f99-a487-c8b24a190331 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e98cef87-098b-4b09-9d65-a8faa89b346f · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9fe8de-7a77-45bd-baba-7d4913a8f81b · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545f009c-c476-415b-8c41-ec954171c74e · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ada282-759c-40ec-8784-3aa7265ac3ae · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Uniter: Universal image-text representation learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9145929-a1dc-4553-a4aa-6727d9e6cca7 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7ef2a5-c6f2-4970-ae26-e0d07112cdc5 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e9bb0d1-a4e3-412e-b890-99f880d9b9c6 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3aa9277e-10de-47ab-b680-eb9f8412d345 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7305b86-daf9-40fe-b4aa-958b5de44d65 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unveiling Encoder-Free Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68489f04-2607-4e17-b14b-c4c59d5c6de7 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d73db33c-a918-4358-a229-4f0b58ebd1e4 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e67fd48f-ed16-4bfd-b86f-e8dd213de02e · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mmgpt4lf: Leverag- ing an optimized pre-trained gpt-2 model with multi-modal cross-attention for load forecasting
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e03fe239-a1a0-40f4-8940-fbe98b498ecd · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e250998-a5bd-4a2b-846f-6d6d09f20097 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vizwiz grand challenge: Answering visual questions from blind people
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97b4e0f1-3f29-4140-b356-8d4085e4f5d3 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Efficient Multimodal Learning from Data-centric Perspective
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b07aac-03de-48a1-9614-df0622cafa0d · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LoRA: Low-Rank Adaptation of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f40682-d614-42b4-a53a-4d2462c1cde0 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a125aabf-069c-468e-893a-6409cbfc2219 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c4ce55f-2ae4-497d-8c5c-be6f910a51c7 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Dvqa: Understanding data visualizations via ques- tion answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e2dc6ef-960f-4711-9de2-38504f54325f · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7db1ba-c02c-4de9-9908-3c8a97f2f384 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs OtterHD: A High-Resolution Multi-modality Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee97ffc7-ea7d-484f-9daf-344ab183e9ce · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fec20d40-0c96-4c3c-8727-edda88af5379 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adbb46d0-1439-4dc9-ac89-57bb7b0e8210 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Evaluating Object Hallucination in Large Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d33b10-36b3-44dc-8822-a3e46aab0ad4 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e70f54-e15c-4080-9aa0-952c91d5f73a · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 359df87f-915d-4925-aed3-013de3b65b0c · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vila: On pre-training for visual language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c43d5bd-080b-44d8-a57f-8defc6663b94 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9665645e-4cba-4403-8551-2dfccc20952f · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Improved Baselines with Visual Instruction Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77762b0e-6226-4098-ab4a-7f3de606e4fb · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Visual instruction tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ee26ed-be62-4a20-8a4c-3d777a7ae95f · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c43b42-5a1a-49f0-bf9d-63210c86902f · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2edb4b9e-99df-45ef-adb5-42ba2f571515 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs A convnet for the 2020s
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0ef25ee-4b4b-4d37-aa7c-ec6ec9e42007 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0870f533-a5ce-4bec-b390-43a3a9d4e9ab · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1baa1c-0008-485a-91cf-4a9730393dd8 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unirgb-ir: A unified frame- work for rgb-infrared semantic tasks via adapter tuning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c90dc6b-1bd6-49ab-b71e-951cbf27c62f · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gpt-4v(ision) system card, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f53c0c-c4a9-4b9e-8b31-45b6b3d077d9 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61817f82-debb-4b6f-9f30-5a9c7a65cf73 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Im2text: Describing images using 1 million captioned pho- tographs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86824750-a82c-400a-90c0-6109603038fc · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Learn- ing transferable visual models from natural language super- vision
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42a8aa21-46f4-4598-a6f7-94c37158d116 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Introducing our multimodal models, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1a630e2-8570-472b-a26a-bea2785aed4c · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a7bcfc6-44a7-4e98-ac0f-e6e35e9920ad · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301a441b-d411-4a07-a674-d5140293ea0d · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Towards vqa models that can read
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77389585-b3f9-4a47-ab97-aa12d7948fa4 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Towards vqa models that can read
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b8bd4df-b998-43a8-a499-7dc662571fe8 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Adapool: Expo- nential adaptive pooling for information-retaining downsam- pling
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc6f390-ff3c-4a79-84d7-ea45e7c07116 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72462e34-fbb2-48d0-9d92-2f6d9989153a · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44557196-22e8-4c08-8f0d-2e3ff4fb4a5b · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38bd5eed-883d-4378-aabc-c748c6c41fcc · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99a8b76-1fc4-4252-8b4a-f455e17a2046 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04f360b7-f5ba-4a69-b6cf-8a1077010caa · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca24d84-cd1c-417a-bb59-ff56e4f0d028 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs xgen-mm (blip-3): A family of open large multimodal models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b53642ed-20c3-4215-833c-5837dc8c421f · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Dense Connector for MLLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5be07de4-f1f7-4d71-8631-80c31b637cec · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Image difference cap- tioning with pre-training and contrastive learning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34b26c7e-af54-4113-a51c-b0978ddb19d5 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b84d7fd-9ebe-4e04-ad98-20fcbb1b05a8 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Modeling context in referring expres- sions
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4591d180-2c6b-4798-99d0-f88e1091ee2a · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8041b73d-0350-4191-a166-d342e42d7523 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Improving rgb-infrared object detection with cascade alignment-guided transformer
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd0d889e-0e3b-40f4-bc79-bbe21341c7ae · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Lane detection transformer based on multi- frame horizontal and vertical attention and visual trans- former module
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7cadc4d-47b2-4570-921e-e97df1d5db9c · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f5ab8f-4ecb-46d3-b237-dc67a5255b2c · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Removal and selection: Improving rgb- infrared object detection via coarse-to-fine fusion
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b728e6cb-ce85-4178-9119-4c47f55440c3 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504e8ba0-fff7-4e4e-b664-abed4832f663 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Deformable DETR: Deformable Transformers for End-to-End Object Detection
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa3c775-5141-4b1b-a531-3693d22ca1c3 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Future work will involve experiments on various LLMs such as Qwen2.5, Mistral, and LLaMA3
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a08b0375-8b33-4c30-a8b9-e9dcc87f04fc · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs The input and output channels of the convolutional kernels are 1024 (equal to the visual feature dimension output by ViT) and 512, respectively
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04e26855-5c68-4bb6-8c77-55131c8235a2 · outbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs CLIP-bind pairs
Reference 512
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.