Pith. sign in

Paper Citation Record · LEDGER

A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2312.12436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12436 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:02:30.892972Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:27:51.967982Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd7f3554-44ee-42ba-b06b-467407e2c8e1 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:41.910944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:b0becedf395ce1a5d288c67424b37402301f6b7efe4e1359ec406e0d9ef9ff23

Observation f15eaf2c-2b1c-491b-8f51-e9039fa916b6 · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:08:01.435214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:b596650be2553da9bd9039bb40cf781acd89952a97a0956a62384ea07a6bd7e6

Observation 737d77a7-40dc-464e-9c30-c372d0fa1ddf · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.122002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:d6fbc0a90197b44984d6935a6a4d66849041226d0c1d8b406fa027c8cbd3f650

Observation f54563f4-19de-49c9-bf52-f9b7adf70306 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.803927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:432ff3e33788df7901c5b5150037f5f31bab51f19eadfadd232d142c62aa2711

Observation 026ec8cd-05d3-41ab-9ba2-5bd5371627b9 · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:51.971525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:b8047042aac29060960de540c472c4e9e77f4ece0c2cae0c56f523a84055495c

Observation f13f8c6f-5713-46a2-adce-0cef7d4feb27 · inbound

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? cites this paper.

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:29:30.088181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T01:29:30.032408Z digest=sha256:fb3039b5bd649b3457cd84bc560aa0f44562aaa2f732cd1085d3ec21700496c5

Observation a484bc68-6ffa-4826-a853-fe2fb956504c · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.643912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:b4a46e27711ac4c2e0c48e8b69a31cf88c3e7e3d502eb4a823830c86620f53d6

Observation 19aa1d17-3199-40e6-8c75-6f59a9889bd4 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.696358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:e945f0e38f47d558dce9700bcbca8719eb1cdaa2b3d26bf943f6ffc115572d8a

Observation fe3e5d8f-105c-4c0f-95b9-272f90625caa · inbound

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey cites this paper.

Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:30.892972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:02:30.892972Z digest=sha256:1c80d4fe6fd8dbb9ae6c24730d2f30f2f3abdc70fb718b695c0380d9f60f60ff

Observation b40d57bb-ec1e-4364-b7a9-13a64f2269e7 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.121595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.121595Z digest=sha256:446832992699f383b834cb2f1062beaf37ff40d90724f7c8cee159b66e790332

Observation a134b349-2279-4f04-b1dd-cc3d1a1ed253 · inbound

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting cites this paper.

RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:04.231658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:29:04.231658Z digest=sha256:7d2795353c5849d25806e7a7c26da429a3182adee88c8e793cedfdd2a4b8e54e