Pith. sign in

REVIEW 2 major objections 1 minor 1 cited by

From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards

T0 review · 2 major / 1 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Multimodal models for ID card presentation attack detection generalize after fine-tuning but fail zero-shot.

desk verdict The paper shows their compact multimodal model for ID card PAD generalizes after fine-tuning but fails zero-shot, yet this pattern is demonstrated on one custom architecture rather than multimodal methods broadly. read the letter →

arxiv 2606.06966 v1 pith:LFKVPODP submitted 2026-06-05 cs.CV

classification cs.CV
keywords presentationattackdetectionmultimodalfusionIDcardscross-domaingeneralizationsyntheticdatabiometricsecurity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces a compact multimodal model that fuses visual and textual features through new generative and discriminative blocks to detect presentation attacks on ID cards despite domain shifts. After supervised fine-tuning the model handles cross-domain cases reliably, yet it shows poor results in zero-shot settings with no task-specific training. This leads to the conclusion that adequate model capacity and real-world training data are required for dependable performance, while existing synthetic datasets fail to represent actual challenges. The authors therefore advocate re-evaluating synthetic data as a benchmark and developing more realistic datasets.

What carries the argument

Compact multimodal model with generative and discriminative blocks that fuse visual and textual data for PAD on ID cards.

What would settle it

A direct comparison in which the multimodal model with the new blocks shows no measurable gain over simple unimodal baselines on held-out real cross-domain ID card data would falsify the claimed benefit of the fusion mechanism.

Watch

Extended reading notes

Core claim

A compact multimodal model using generative and discriminative blocks to combine visual and textual data for presentation attack detection on genuine and synthetic ID images achieves strong cross-domain generalization after supervised fine-tuning but fails in zero-shot settings, showing that model capacity and real-world data are essential while synthetic datasets may not reflect real-world challenges.

Load-bearing premise

The generative and discriminative blocks produce an effective fusion of visual and textual data that delivers cross-domain robustness.

Editorial extensions

If this is right

  • Supervised fine-tuning allows multimodal PAD models to generalize across domains for ID card attacks.
  • The same models exhibit unreliable performance when applied without any task-specific training.
  • Larger model capacity improves the reliability of PAD results.
  • Current synthetic datasets do not capture the difficulties of real PAD scenarios.
  • Advancing the field requires more realistic and diverse real-world datasets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Scaling model size further could reduce dependence on large amounts of labeled real data for zero-shot transfer.
  • Testing the same fusion blocks on other biometric modalities might reveal whether the approach generalizes beyond ID cards.
  • Creating improved synthetic data generators that better mimic real capture conditions could serve as an interim solution until larger real datasets become available.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper proposes a compact multimodal architecture for presentation attack detection (PAD) on ID cards that fuses visual and textual features via novel generative and discriminative blocks. It reports that the model generalizes well across domains after supervised fine-tuning but fails in zero-shot settings, concluding that model capacity and real-world data are essential while existing synthetic datasets are inadequate for benchmarking.

Significance. If the empirical findings hold under broader validation, the work would usefully highlight limitations of synthetic data for cross-domain PAD and the practical value of multimodal fusion under supervised regimes. The emphasis on real-world data needs is timely given privacy constraints in ID document datasets.

major comments (2)
  1. [Abstract, §3] Abstract and §3 (model description): The central claim that 'multimodal models exhibit strong generalisation after supervised fine-tuning' while failing zero-shot is framed as a property of the multimodal paradigm, yet the experiments appear limited to a single custom architecture with the proposed generative and discriminative blocks. This prevents attribution of the observed pattern to multimodal models in general rather than to the specific fusion design, capacity, or training recipe.
  2. [Abstract] Abstract: The statement that 'existing synthetic datasets may not reflect real-world challenges' is load-bearing for the recommendation to re-evaluate synthetic benchmarks, but no quantitative comparison (e.g., domain-shift metrics or cross-dataset performance tables) is referenced in the provided abstract to support the claim.
minor comments (1)
  1. [Abstract] Abstract lacks any mention of datasets, metrics, or experimental protocol, making it impossible to assess the strength of the reported generalization results from the summary alone.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the insightful comments, which help us clarify the scope of our claims and strengthen the presentation of our findings. We address each major comment point by point below.

read point-by-point responses
  1. Referee: [Abstract, §3] Abstract and §3 (model description): The central claim that 'multimodal models exhibit strong generalisation after supervised fine-tuning' while failing zero-shot is framed as a property of the multimodal paradigm, yet the experiments appear limited to a single custom architecture with the proposed generative and discriminative blocks. This prevents attribution of the observed pattern to multimodal models in general rather than to the specific fusion design, capacity, or training recipe.

    Authors: We acknowledge that our experiments are performed using the proposed compact multimodal architecture with the novel generative and discriminative blocks. The observed generalization after fine-tuning and failure in zero-shot are specific to this model and training setup. We will revise the abstract and §3 to replace the general phrasing 'multimodal models' with 'our multimodal model' to accurately reflect the scope of the results. Additionally, we will include a discussion note suggesting that future work could validate these patterns across a wider range of multimodal architectures. revision: yes

  2. Referee: [Abstract] Abstract: The statement that 'existing synthetic datasets may not reflect real-world challenges' is load-bearing for the recommendation to re-evaluate synthetic benchmarks, but no quantitative comparison (e.g., domain-shift metrics or cross-dataset performance tables) is referenced in the provided abstract to support the claim.

    Authors: The abstract is a concise summary, with the supporting quantitative results (cross-domain performance after fine-tuning versus zero-shot) presented in the main body of the paper. To address the concern, we will revise the abstract to briefly reference the key empirical observations, such as the performance differences that indicate limitations of synthetic data, thereby making the claim more directly supported within the abstract itself. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical model evaluation with no derivations or self-referential claims

full rationale

The paper proposes a compact multimodal architecture for cross-domain PAD and reports experimental results on generalization after fine-tuning versus zero-shot failure. No equations, derivations, or first-principles claims appear in the abstract or described content. The central findings rest on supervised training and testing of the introduced model rather than any reduction of predictions to fitted inputs or self-citation chains. The architecture is presented as novel (generative and discriminative blocks), with performance claims tied directly to its implementation and data, without renaming known results or smuggling ansatzes via prior self-citations. This is a standard empirical CV contribution whose validity can be assessed against external benchmarks.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

No free parameters, axioms, or invented entities are described or invoked in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards." pith.science (2026). https://pith.science/paper/LFKVPODP

@misc{pith2026260606966,
  author       = {Pith},
  title        = {Pith review of: From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LFKVPODP}},
  note         = {Machine review of arXiv:2606.06966}
}
read the original abstract

Cross-domain shifts challenge Presentation Attack Detection (PAD) on ID Cards, given the restricted data available due to privacy concerns. This work proposes a compact multimodal model, based on new generative and discriminative blocks, which combines visual and textual data for PAD on genuine and synthetic ID images. While multimodal models exhibit strong generalisation after supervised fine-tuning, they fail in zero-shot settings. Our findings underscore that model capacity and real-world data are essential for reliable PAD, while existing synthetic datasets may not reflect real-world challenges. We argue for a re-evaluation of synthetic data as a benchmark and emphasise the need for more realistic, diverse datasets to advance PAD research.

Figures

Figures reproduced from arXiv: 2606.06966 by the authors.

Figure 1
Figure 1. Examples of four different attack types. From top to bottom: Chile and Mexico ID Card datasets. From left to right: [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Examples of four different attack types. From top to bottom: Poland, Portugal, and Spain ID Card datasets. From left [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The two different structures of SmolVLM2 in PAD on ID Cards. special token [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection

    cs.CR 2026-07 unverdicted novelty 7.0 of 10

    A systematic survey unifies presentation, digital injection, and GenAI synthesis attacks on identity documents, audits datasets for a reality gap, identifies SDGI in multimodal models, and reports APCER above 25% for ...

Reference graph

Works this paper leans on

34 extracted references · 9 canonical work pages · cited by 1 Pith paper

  1. [1]

    Foundation models defining a new era in vision: A survey and outlook,

    M. Awais, M. Naseer, S. Khan, R. M. Anwer, H. Cholakkal, M. Shah, M.-H. Yang, and F. S. Khan, “Foundation models defining a new era in vision: A survey and outlook,”IEEE Trans. on Pattern Analysis and Machine Intelligence, vol. 47, no. 4, pp. 2245–2264, 2025

  2. [2]

    Can foundation models generalise the presentation attack detection capabilities on id cards?

    J. E. Tapia and C. Busch, “Can foundation models generalise the presentation attack detection capabilities on id cards?” 2025. [Online]. Available: https://arxiv.org/abs/2506.05263

  3. [3]

    Explainability and vision foundation models: A survey,

    R. Kazmierczak, E. Berthier, G. Frehse, and G. Franchi, “Explainability and vision foundation models: A survey,”Information Fusion, vol. 122, p. 103184, 2025. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S156625352500257X

  4. [4]

    Identity card presentation attack detection: A systematic review,

    E. M. Ruiz, J. E. Tapia, R. T. Soto, and C. Busch, “Identity card presentation attack detection: A systematic review,” 2025. [Online]. Available: https://arxiv.org/abs/2511.06056

  5. [5]

    Forged presentation attack detection for ID cards on remote verification systems,

    S. Gonzalez and J. E. Tapia, “Forged presentation attack detection for ID cards on remote verification systems,”Pattern Recognition, vol. 162, p. 111352, Jun. 2025. [Online]. Available: https://www.sciencedirect. com/science/article/pii/S0031320325000123

  6. [6]

    Hybrid Two-Stage Architecture for Tampering Detection of Chipless ID Cards,

    S. Gonzalez, A. Valenzuela, and J. Tapia, “Hybrid Two-Stage Architecture for Tampering Detection of Chipless ID Cards,”Trans. on Biometrics, Behavior, and Identity Science, vol. 3, no. 1, pp. 89–100, Jan

  7. [7]

    Available: https://ieeexplore.ieee.org/document/9197632

    [Online]. Available: https://ieeexplore.ieee.org/document/9197632

  8. [8]

    Open-Set: ID Card Presentation Attack Detection Using Neural Style Transfer,

    R. P. Markham, J. M. E. L ´opez, M. Nieto-Hidalgo, and J. E. Tapia, “Open-Set: ID Card Presentation Attack Detection Using Neural Style Transfer,”IEEE Access, vol. 12, pp. 68 573–68 585, 2024. [Online]. Available: https://ieeexplore.ieee.org/document/10520890

Show all 34 references
  1. [9]

    Image-to-image translation with conditional adversarial networks,

    P. Isola, J.-Y . Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,”CVPR, 2017

  2. [10]

    Synthetic ID Card Image Generation for Improving Presentation Attack Detection,

    D. Benalcazar, J. E. Tapia, S. Gonzalez, and C. Busch, “Synthetic ID Card Image Generation for Improving Presentation Attack Detection,” Trans. on Information Forensics and Security, vol. 18, pp. 1814–1824,

  3. [11]

    Available: https://ieeexplore.ieee.org/abstract/document/ 10065533

    [Online]. Available: https://ieeexplore.ieee.org/abstract/document/ 10065533

  4. [12]

    Idnet: A novel identity document dataset via few-shot and quality-driven synthetic data generation,

    L. Xie, Y . Wang, H. Guan, S. Nag, R. Goel, N. Swamy, Y . Yang, C. Xiao, J. Prisby, R. Maciejewski, and J. Zou, “Idnet: A novel identity document dataset via few-shot and quality-driven synthetic data generation,” inIntl. Conf. on Big Data (BigData), 2024, pp. 2244–2253

  5. [13]

    Fan- tasyID: A dataset for detecting digital manipulations in ID-documents,

    P. Korshunov, A. Mohammadi, Vidit, C. Ecabert, and S. Marcel, “Fan- tasyID: A dataset for detecting digital manipulations in ID-documents,” inIEEE Intl. Joint Conf. on Biometrics (IJCB), 2025, pp. 1–9

  6. [14]

    First competition on presentation attack detection on ID card,

    J. E. Tapia, N. Damer, C. Busch, J. M. Espin, J. Barrachina, A. S. Rocamora, K. Ocvirk, L. Alessio, B. Batagelj, S. Patwardhan, R. Ra- machandra, R. Mudgalgundurao, K. Raja, D. Schulz, and C. Aravena, “First competition on presentation attack detection on ID card,” inIntl. Joi...

  7. [15]

    Second competition on presentation attack detection on ID card,

    J. E. Tapia, M. Nieto, J. M. Espin, A. S. Rocamora, J. Barrachina, N. Damer, C. Busch, M. Ivanovska, L. Todorov, R. Khizbullin, L. Lazarevich, A. Grishin, D. Schulz, S. Gonzalez, A. Mohammadi, K. Kotwal, S. Marcel, R. Mudgalgundurao, K. Raja, P. Schuch, S. Pat- wardhan, R. Ram...

  8. [16]

    Contrastive localized language-image pre-training,

    H.-Y . Chen, Z. Lai, H. Zhang, X. Wang, M. Eichner, K. You, M. Cao, B. Zhang, Y . Yang, and Z. Gan, “Contrastive localized language-image pre-training,” inForty-second Intl. Conf.on Machine Learning, 2025. [Online]. Available: https://openreview.net/forum?id=sGQEOXlezg

  9. [17]

    DINOv2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut,...

  10. [18]

    Fakeidet: Exploring patches for privacy-preserving fake id detection,

    J. Mu ˜noz-Haro, R. Tolosana, R. Vera-Rodriguez, A. Morales, and J. Fierrez, “Fakeidet: Exploring patches for privacy-preserving fake id detection,” inIEEE Intl. Joint Conf. on Biometrics (IJCB), 2025, pp. 1–9

  11. [19]

    Syn- idpass: Passport synthetic dataset for presentation attack detection,

    J. E. Tapia, F. Stockhardt, L. J. Gonz ´alez-Soler, and C. Busch, “Syn- idpass: Passport synthetic dataset for presentation attack detection,” in IEEE Intl. Joint Conf. on Biometrics (IJCB), 2025, pp. 1–9

  12. [20]

    Densely connected convolutional networks,

    G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” inProceedings of the IEEE confer- ence on computer vision and pattern recognition, 2017, pp. 4700–4708

  13. [21]

    Pixel-wise supervision for presentation attack detection on identity document cards,

    R. Mudgalgundurao, P. Schuch, K. Raja, R. Ramachandra, and N. Damer, “Pixel-wise supervision for presentation attack detection on identity document cards,”IET biometrics, vol. 11, no. 5, pp. 383–395, 2022

  14. [22]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009, pp. 248–255

  15. [23]

    Sigmoid loss for language image pre-training,

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inProceedings of the IEEE/CVF Intl. Conf. on computer vision, 2023, pp. 11 975–11 986

  16. [24]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020

  17. [25]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017

  18. [26]

    Gaussian error linear units (gelus),

    D. Hendrycks, “Gaussian error linear units (gelus),”arXiv preprint arXiv:1606.08415, 2016

  19. [27]

    Smolvlm: Redefining small and efficient multimodal models,

    A. Marafioti, O. Zohar, M. Farr ´e, M. Noyan, E. Bakouch, P. Cuenca, C. Zakka, L. B. Allal, A. Lozhkov, N. Taziet al., “Smolvlm: Redefining small and efficient multimodal models,”arXiv preprint arXiv:2504.05299, 2025

  20. [28]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017

  21. [29]

    Sgdr: Stochastic gradient descent with warm restarts,

    ——, “Sgdr: Stochastic gradient descent with warm restarts,”arXiv preprint arXiv:1608.03983, 2016

  22. [30]

    Why warmup the learning rate? under- lying mechanisms and improvements,

    D. S. Kalra and M. Barkeshli, “Why warmup the learning rate? under- lying mechanisms and improvements,”Advances in Neural Information Processing Systems, vol. 37, pp. 111 760–111 801, 2024

  23. [31]

    Optuna: A next- generation hyperparameter optimization framework,

    T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next- generation hyperparameter optimization framework,” inProceedings of the 25th ACM SIGKDD intl. conf. on knowledge discovery & data mining, 2019, pp. 2623–2631

  24. [32]

    C. M. Bishop and N. M. Nasrabadi,Pattern recognition and machine learning. Springer, 2006, vol. 4, no. 4

  25. [33]

    Feature selection, l 1 vs. l 2 regularization, and rotational invariance,

    A. Y . Ng, “Feature selection, l 1 vs. l 2 regularization, and rotational invariance,” inProceedings of the twenty-first intl. conf. on Machine learning, 2004, p. 78

  26. [34]

    Lora: Low-rank adaptation of large language models

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chenet al., “Lora: Low-rank adaptation of large language models.” ICLR, vol. 1, no. 2, p. 3, 2022. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12 Qingweng Zengreceived a B.Sc. degree in Dat...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.