{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:B6C5TDIKF4IDBRJ42RLPIIWMDT","short_pith_number":"pith:B6C5TDIK","schema_version":"1.0","canonical_sha256":"0f85d98d0a2f1030c53cd456f422cc1cf49ec5ab5c84c8545adc03f62b4e96d5","source":{"kind":"arxiv","id":"2412.18619","version":2},"attestation_state":"computed","paper":{"title":"Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.CV","cs.LG","cs.MM","eess.AS"],"primary_cat":"cs.CL","authors_text":"Aaron Yee, Andreas Vlachos, Baobao Chang, Ge Zhang, Haozhe Zhao, Hongcheng Guo, Jian Yang, Junyang Lin, Lei Li, Lei Zhang, Liang Chen, Lingwei Meng, Minjia Zhang, Qingxiu Dong, Ruoyu Wu, Shuai Bai, Shuhuai Ren, Shujie Hu, Tianyu Liu, Wen Xiao, Xu Tan, Yichi Zhang, Yizhe Xiong, Yulong Chen, Yunshui Li, Zefan Cai, Zekun Wang","submitted_at":"2024-12-16T05:02:25Z","abstract_excerpt":"Building on the foundations of language modeling in natural language processing, Next Token Prediction (NTP) has evolved into a versatile training objective for machine learning tasks across various modalities, achieving considerable success. As Large Language Models (LLMs) have advanced to unify understanding and generation tasks within the textual modality, recent research has shown that tasks from different modalities can also be effectively encapsulated within the NTP framework, transforming the multimodal information into tokens and predict the next one given the context. This survey intr"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2412.18619","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.CL","submitted_at":"2024-12-16T05:02:25Z","cross_cats_sorted":["cs.AI","cs.CV","cs.LG","cs.MM","eess.AS"],"title_canon_sha256":"7a33ce22a3a0ba067522dc31cd596abbcc21149c5e0cf32ba53d5bd4b5eb5a02","abstract_canon_sha256":"e9d0a749b9502bcebb45b8f702c6ace26f97e4c570629a25ea5a4f7140ce1bd2"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:55:15.887619Z","signature_b64":"erzuePW16MCwCpf/ToQ3kA+4RlJRS2g4Wp3NSSVdNhS/DoWkmewgg8cKN2BzVZ3nsqK/qkV+EGn6R1XApC25AQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"0f85d98d0a2f1030c53cd456f422cc1cf49ec5ab5c84c8545adc03f62b4e96d5","last_reissued_at":"2026-07-05T09:55:15.887154Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:55:15.887154Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.CV","cs.LG","cs.MM","eess.AS"],"primary_cat":"cs.CL","authors_text":"Aaron Yee, Andreas Vlachos, Baobao Chang, Ge Zhang, Haozhe Zhao, Hongcheng Guo, Jian Yang, Junyang Lin, Lei Li, Lei Zhang, Liang Chen, Lingwei Meng, Minjia Zhang, Qingxiu Dong, Ruoyu Wu, Shuai Bai, Shuhuai Ren, Shujie Hu, Tianyu Liu, Wen Xiao, Xu Tan, Yichi Zhang, Yizhe Xiong, Yulong Chen, Yunshui Li, Zefan Cai, Zekun Wang","submitted_at":"2024-12-16T05:02:25Z","abstract_excerpt":"Building on the foundations of language modeling in natural language processing, Next Token Prediction (NTP) has evolved into a versatile training objective for machine learning tasks across various modalities, achieving considerable success. As Large Language Models (LLMs) have advanced to unify understanding and generation tasks within the textual modality, recent research has shown that tasks from different modalities can also be effectively encapsulated within the NTP framework, transforming the multimodal information into tokens and predict the next one given the context. This survey intr"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2412.18619","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2412.18619/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2412.18619","created_at":"2026-07-05T09:55:15.887209+00:00"},{"alias_kind":"arxiv_version","alias_value":"2412.18619v2","created_at":"2026-07-05T09:55:15.887209+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2412.18619","created_at":"2026-07-05T09:55:15.887209+00:00"},{"alias_kind":"pith_short_12","alias_value":"B6C5TDIKF4ID","created_at":"2026-07-05T09:55:15.887209+00:00"},{"alias_kind":"pith_short_16","alias_value":"B6C5TDIKF4IDBRJ4","created_at":"2026-07-05T09:55:15.887209+00:00"},{"alias_kind":"pith_short_8","alias_value":"B6C5TDIK","created_at":"2026-07-05T09:55:15.887209+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.24956","citing_title":"NITP: Next Implicit Token Prediction for LLM Pre-training","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18060","citing_title":"PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience","ref_index":87,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01543","citing_title":"PathAR: Structure-First Autoregressive Synthesis of Multimodal Pathology Images","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24956","citing_title":"NITP: Next Implicit Token Prediction for LLM Pre-training","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25343","citing_title":"Toward Native Multimodal Modeling: A Roadmap","ref_index":224,"is_internal_anchor":false},{"citing_arxiv_id":"2504.08528","citing_title":"On The Landscape of Spoken Language Models: A Comprehensive Survey","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2503.07265","citing_title":"WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08809","citing_title":"SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization","ref_index":1,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT","json":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT.json","graph_json":"https://pith.science/api/pith-number/B6C5TDIKF4IDBRJ42RLPIIWMDT/graph.json","events_json":"https://pith.science/api/pith-number/B6C5TDIKF4IDBRJ42RLPIIWMDT/events.json","paper":"https://pith.science/paper/B6C5TDIK"},"agent_actions":{"view_html":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT","download_json":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT.json","view_paper":"https://pith.science/paper/B6C5TDIK","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2412.18619&json=true","fetch_graph":"https://pith.science/api/pith-number/B6C5TDIKF4IDBRJ42RLPIIWMDT/graph.json","fetch_events":"https://pith.science/api/pith-number/B6C5TDIKF4IDBRJ42RLPIIWMDT/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT/action/timestamp_anchor","attest_storage":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT/action/storage_attestation","attest_author":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT/action/author_attestation","sign_citation":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT/action/citation_signature","submit_replication":"https://pith.science/pith/B6C5TDIKF4IDBRJ42RLPIIWMDT/action/replication_record"}},"created_at":"2026-07-05T09:55:15.887209+00:00","updated_at":"2026-07-05T09:55:15.887209+00:00"}