{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:O6DMJXICWZC2YBYRO45GWTNLLU","short_pith_number":"pith:O6DMJXIC","schema_version":"1.0","canonical_sha256":"7786c4dd02b645ac0711773a6b4dab5d048b6cdc220754a21c0af1a28ef4a652","source":{"kind":"arxiv","id":"2312.12436","version":2},"attestation_state":"computed","paper":{"title":"A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.MM"],"primary_cat":"cs.CV","authors_text":"Chaoyou Fu, Deqiang Jiang, Di Yin, Gaoxiang Ye, Hongsheng Li, Ke Li, Longtian Qiu, Mengdan Zhang, Peixian Chen, Peng Gao, Renrui Zhang, Shaohui Lin, Sirui Zhao, Xing Sun, Yubo Huang, Yunhang Shen, Zhengye Zhang, Zihan Wang","submitted_at":"2023-12-19T18:59:22Z","abstract_excerpt":"The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in visual understanding, enabling them to tackle diverse multi-modal tasks. Very recently, Google released Gemini, its newest and most capable MLLM built from the ground up for multi-modality. In light of the superior reasoning capabilities, can Gemini challenge GPT-4V's leading position in multi-modal learning? In this paper, we present a preliminary exploration"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2312.12436","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2023-12-19T18:59:22Z","cross_cats_sorted":["cs.AI","cs.CL","cs.MM"],"title_canon_sha256":"e745297725411448a3abdfa571f63de8bc86516c74d589749ee8adf571514060","abstract_canon_sha256":"e45debbbcae28087843f1a14b9c5fa683dbba5cf493999e9717b1d9e03120c55"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:26:20.910784Z","signature_b64":"Vsr0c9Y8tiCm/t2/kJU0VVtRGGUYVA4A/DcDiYozQie2Jo1dMSrtWsFRnSjC1WuWjOtvT2PurtD3nonLgVXXCQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"7786c4dd02b645ac0711773a6b4dab5d048b6cdc220754a21c0af1a28ef4a652","last_reissued_at":"2026-07-05T07:26:20.910284Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:26:20.910284Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.MM"],"primary_cat":"cs.CV","authors_text":"Chaoyou Fu, Deqiang Jiang, Di Yin, Gaoxiang Ye, Hongsheng Li, Ke Li, Longtian Qiu, Mengdan Zhang, Peixian Chen, Peng Gao, Renrui Zhang, Shaohui Lin, Sirui Zhao, Xing Sun, Yubo Huang, Yunhang Shen, Zhengye Zhang, Zihan Wang","submitted_at":"2023-12-19T18:59:22Z","abstract_excerpt":"The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in visual understanding, enabling them to tackle diverse multi-modal tasks. Very recently, Google released Gemini, its newest and most capable MLLM built from the ground up for multi-modality. In light of the superior reasoning capabilities, can Gemini challenge GPT-4V's leading position in multi-modal learning? In this paper, we present a preliminary exploration"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2312.12436","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2312.12436/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2312.12436","created_at":"2026-07-05T07:26:20.910353+00:00"},{"alias_kind":"arxiv_version","alias_value":"2312.12436v2","created_at":"2026-07-05T07:26:20.910353+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2312.12436","created_at":"2026-07-05T07:26:20.910353+00:00"},{"alias_kind":"pith_short_12","alias_value":"O6DMJXICWZC2","created_at":"2026-07-05T07:26:20.910353+00:00"},{"alias_kind":"pith_short_16","alias_value":"O6DMJXICWZC2YBYR","created_at":"2026-07-05T07:26:20.910353+00:00"},{"alias_kind":"pith_short_8","alias_value":"O6DMJXIC","created_at":"2026-07-05T07:26:20.910353+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2402.03766","citing_title":"MobileVLM V2: Faster and Stronger Baseline for Vision Language Model","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2407.03320","citing_title":"InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2401.16420","citing_title":"InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2403.14624","citing_title":"MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2312.16886","citing_title":"MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2408.13257","citing_title":"MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2306.13549","citing_title":"A Survey on Multimodal Large Language Models","ref_index":139,"is_internal_anchor":false},{"citing_arxiv_id":"2311.10122","citing_title":"Video-LLaVA: Learning United Visual Representation by Alignment Before Projection","ref_index":121,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU","json":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU.json","graph_json":"https://pith.science/api/pith-number/O6DMJXICWZC2YBYRO45GWTNLLU/graph.json","events_json":"https://pith.science/api/pith-number/O6DMJXICWZC2YBYRO45GWTNLLU/events.json","paper":"https://pith.science/paper/O6DMJXIC"},"agent_actions":{"view_html":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU","download_json":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU.json","view_paper":"https://pith.science/paper/O6DMJXIC","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2312.12436&json=true","fetch_graph":"https://pith.science/api/pith-number/O6DMJXICWZC2YBYRO45GWTNLLU/graph.json","fetch_events":"https://pith.science/api/pith-number/O6DMJXICWZC2YBYRO45GWTNLLU/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU/action/timestamp_anchor","attest_storage":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU/action/storage_attestation","attest_author":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU/action/author_attestation","sign_citation":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU/action/citation_signature","submit_replication":"https://pith.science/pith/O6DMJXICWZC2YBYRO45GWTNLLU/action/replication_record"}},"created_at":"2026-07-05T07:26:20.910353+00:00","updated_at":"2026-07-05T07:26:20.910353+00:00"}