Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2312.12436.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:02:30.892972Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T15:27:51.967982Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation dd7f3554-44ee-42ba-b06b-467407e2c8e1 · inbound
A Survey on Multimodal Large Language Models A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f15eaf2c-2b1c-491b-8f51-e9039fa916b6 · inbound
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 737d77a7-40dc-464e-9c30-c372d0fa1ddf · inbound
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f54563f4-19de-49c9-bf52-f9b7adf70306 · inbound
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 026ec8cd-05d3-41ab-9ba2-5bd5371627b9 · inbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f13f8c6f-5713-46a2-adce-0cef7d4feb27 · inbound
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a484bc68-6ffa-4826-a853-fe2fb956504c · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 19aa1d17-3199-40e6-8c75-6f59a9889bd4 · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe3e5d8f-105c-4c0f-95b9-272f90625caa · inbound
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40d57bb-ec1e-4364-b7a9-13a64f2269e7 · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a134b349-2279-4f04-b1dd-cc3d1a1ed253 · inbound
RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.