Pith. sign in

REVIEW 6 cited by

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18619 v2 pith:B6C5TDIK submitted 2024-12-16 cs.CL cs.AIcs.CVcs.LGcs.MMeess.AS

classification cs.CLcs.AIcs.CVcs.LGcs.MMeess.AS
keywords multimodallanguagenexttaskstaxonomycomprehensivegenerationgithub
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Building on the foundations of language modeling in natural language processing, Next Token Prediction (NTP) has evolved into a versatile training objective for machine learning tasks across various modalities, achieving considerable success. As Large Language Models (LLMs) have advanced to unify understanding and generation tasks within the textual modality, recent research has shown that tasks from different modalities can also be effectively encapsulated within the NTP framework, transforming the multimodal information into tokens and predict the next one given the context. This survey introduces a comprehensive taxonomy that unifies both understanding and generation within multimodal learning through the lens of NTP. The proposed taxonomy covers five key aspects: Multimodal tokenization, MMNTP model architectures, unified task representation, datasets \& evaluation, and open challenges. This new taxonomy aims to aid researchers in their exploration of multimodal intelligence. An associated GitHub repository collecting the latest papers and repos is available at https://github.com/LMM101/Awesome-Multimodal-Next-Token-Prediction

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

    cs.CV 2025-05 conditional novelty 7.0 of 10

    CGC+VTD identifies co-occurring image token clusters as a source of hallucinated objects in discrete-token LVLMs and suppresses clusters' absent-token signals in latent space, cutting hallucination rates across Chamel...

  2. NITP: Next Implicit Token Prediction for LLM Pre-training

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    NITP augments standard next-token prediction with implicit semantic prediction in representation space using shallow-layer self-supervision, reporting consistent downstream gains on 0.5B-9B models including 5.7% on MM...

  3. A Unified Low-level Foundation Model for Enhancing Pathology Image Quality

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A prompt-guided diffusion model pretrained on 190 million pathology patches outperforms task-specific models across most restoration and virtual staining benchmarks.

  4. MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors introduce a 3D CT-based visual question answering benchmark with six error types and three task levels, and show that current 3D medical MLLMs perform poorly on it.

  5. Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Citss combines sentence-level cropping and keyphrase perturbation with contrastive learning to fine-tune both encoder and decoder language models for citation classification.

  6. Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.

Pith tools