{"id":"978121e3-6e62-4812-8ac0-04e4ed76286c","arxiv_id":"2502.08181","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that maps few-shot class incremental learning into five technical approaches and five settings, with performance comparisons and open-problem analysis.","lead":"This paper surveys few-shot class incremental learning (FSCIL), the problem of teaching a model new classes from very few examples without forgetting older classes. It sorts recent methods into five technical approaches and five problem settings, compares their reported accuracy, and lists open challenges.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1's PEFT-superiority trend is confounded: the up-to-20% gap in §4.6 compares ImageNet-pretrained ViT-B/16 PEFT against ResNet18 methods mostly trained from scratch, so the gap need not reflect the PEFT approach.","rationale":"The reader's weakest assumption accurately identifies Table 1 comparability as the main risk. My stress pass sharpens this into a single structural confound: PEFT rows and non-PEFT rows in Table 1 differ systematically in backbone (ViT-B/16 vs ResNet18) and in pretraining availability, and the paper itself provides evidence for this confound in §4.1 and §4.6. A matched-backbone experiment would settle whether the 'up to 20%' gap is attributable to PEFT as a method family or to the underlying pretrained ViT; absent that experiment, the survey's main comparative conclusion is tentative. Secondary issues—SSFSCIL is defined but no SSFSCIL method appears in Table 1 or Table 3, the prototype-tuning objective is cross-referenced as eq. 2 instead of eq. 4 in §3.3, and no systematic inclusion criteria are given—also support a conditional rather than unconditional acceptance. None of these problems requires rejection; the survey remains a useful taxonomy if its performance claims are read as preliminary and its tables are corrected. The reader's CONDITIONAL verdict is therefore appropriate and unchanged.","tokens_in":18804,"tokens_out":6859,"duration_ms":56271,"concrete_test":"Run a matched-backbone ablation on CIFAR-100 (or CUB): take one representative PEFT method (e.g., ASP or a minimal prompt-based PEFT) and the best-performing non-PEFT method (e.g., NC-FSCIL or CLOSER), and evaluate both under two conditions: (a) both using ResNet18 initialized from ImageNet, and (b) both using ViT-B/16 initialized from ImageNet, with identical base/few-shot session schedules, shots per class, and classifier type. If the accuracy gap between PEFT and non-PEFT remains above 20% within at least one matched-backbone condition, the §4.6 claim is supported; if the gap collapses or reverses when backbone and pretraining are held fixed, the reported PEFT superiority is an artifact of ViT/pretraining rather than the approach.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central practical message is that PEFT is superior, stated in §4.6 as 'PEFT methods e.g. ASP and Privilege confirm the superiority of the PEFT approach to the other 4 approaches with a significant gap i.e. up to 20%.' This conclusion is load-bearing but not supported by Table 1 because the method family is confounded with backbone and pretraining. Section 4.1 reports that the four non-PEFT approaches mostly use ResNet18 and 'utilize PTM in the CUB dataset' only, whereas PEFT methods use ViT-B/16 (86.57M parameters) and a PTM 'for all scenarios.' The survey itself says the >90% MiniImageNet results are 'common-sense results as the methods utilize pre-trained ViT on ImageNet (the superset of MiniImageNet).' Thus the accuracy gap between ASP/Privilege and methods such as NC-FSCIL or CLOSER on MiniImageNet and CIFAR100 can be explained by pretraining and backbone capacity rather than by the PEFT approach. Additionally, some 'PEFT' rows (L2P, DualP, CodaP) are rehearsal-free CIL methods, not native FSCIL methods; Table 2 shows 0.00% novel-class accuracy for them, so their inclusion in the same comparison further weakens the 'up to 20%' claim. Because the survey's headline trend depends on this comparison, the claim needs a matched evaluation before it can be accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of few-shot class incremental learning (FSCIL). It proposes a five-way taxonomy of approaches (backbone tuning, meta-learning, prototype tuning, dynamic architecture, and parameter-efficient fine-tuning), formal objectives for each approach, and a taxonomy of five settings (supervised, semi-supervised, unsupervised, complete, and cross-domain FSCIL). The paper compiles a large comparison table (Table 1) covering backbone, learnable parameters, pre-trained model usage, classifier type, augmentation, prototype rectification, language guidance, and accuracy metrics (AA, PD, AHM), and it includes a separate application table (Table 3) for medical, remote sensing, NLP, and graph domains. It also discusses prototype rectification mechanisms, stability-plasticity trade-offs, open challenges, and future directions. The paper's central practical claim is that PEFT methods outperform the other four approaches by a significant margin, up to 20%.","tokens_in":19051,"tokens_out":5801,"duration_ms":45347,"significance":"If the comparison were methodologically matched, the survey would be a useful map of the FSCIL field: Table 1 is unusually detailed, the formal objective statements for each approach are a useful reference, and the application table covers areas that are often omitted from FSCIL surveys. The paper does not provide machine-checked proofs or reproducible code, but its contribution is descriptive and organizational. However, the headline trend is currently confounded: the PEFT rows use a different backbone and pretraining regime from the other approaches, so the reported accuracy gap cannot be attributed to the PEFT approach alone. With a corrected comparison and more cautious claims, the survey could serve as a valuable reference; in its present form, its main conclusion is not supported by the evidence it presents.","major_comments":[{"comment":"The claim that 'PEFT methods e.g. ASP and Privilege confirm the superiority of the PEFT approach to the other 4 approaches with a significant gap i.e. up to 20%' is confounded. In Table 1, all PEFT rows use ViT-B/16 (86.57M parameters) with a pre-trained model on all datasets, whereas the other four approaches use ResNet18 (11.67M parameters) and a PTM only on CUB. The paper itself states in §4.6 that the >90% MiniImageNet results are 'common-sense results as the methods utilize pre-trained ViT on ImageNet (the superset of MiniImageNet)'. Therefore the accuracy gap may reflect backbone capacity and pretraining rather than the PEFT approach itself. The authors should provide a matched comparison with the same backbone, pretraining, sessions, and protocol, or substantially soften the causal wording of the PEFT-superiority conclusion.","section":"§4.6, Table 1"},{"comment":"The text says 'The objective of the prototype tuning approach is defined in eq. 2,' but Eq. (2) is the backbone-tuning objective; the prototype-tuning objective is Eq. (4). This is not merely a typo: the formal objective is one of the paper's stated contributions, and the incorrect pointer makes the section difficult to verify.","section":"§3.3, Eq. (4)"},{"comment":"Section 4.7 ends mid-sentence ('indicating that the') before Table 2, and the resumed text ('methods rely too much on base class accuracy') appears only after the table. The stability-plasticity analysis is a central part of the survey, and this incomplete sentence disrupts the presentation. The passage should be rewritten as complete prose so that Table 2 is cited rather than inserted into a sentence.","section":"§4.7, Table 2"},{"comment":"The taxonomy in §2.2 lists semi-supervised FSCIL (SSFSCIL) as one of the sub-settings, and §1 claims coverage of SSFSCIL, but no surveyed method in Table 1 or Table 3 actually uses the semi-supervised setting. In §4.8, the only non-supervised settings mentioned are UFSCIL, CFSCIL, and CDFSCIL. The authors should either include SSFSCIL methods in the survey or explicitly state that the setting is defined for completeness but currently has no surveyed implementations.","section":"§2.2, §4.8"},{"comment":"L2P, DualP, and CodaP are presented as PEFT FSCIL methods, but Table 2 reports 0.00% novel-class accuracy for them on both CIFAR100 and CUB, which suggests they are not performing FSCIL as defined in §2.1. Their inclusion in the PEFT comparison, and their role as the base for the L2P+, DualP+, and CodaP+ variants, needs explicit justification. As presented, these rows weaken the comparison and the conclusion that PEFT is superior for FSCIL.","section":"§3.5, Table 1, Table 2"}],"minor_comments":[{"comment":"The heading 'Types of FSCIL Setings' contains a typo; it should be 'Settings'.","section":"§2.2 heading"},{"comment":"The word 'aporoaches' should be 'approaches', and 'discrimivative' elsewhere should be 'discriminative'.","section":"§3.1"},{"comment":"The term 'SSFCIL' is used in the introduction but 'SSFSCIL' is used elsewhere; the abbreviation should be consistent.","section":"§1"},{"comment":"The label 'CFSCIL' is used both for the 'Complete FSCIL' setting in §2.2(d) and for the method 'C-FSCIL' in Table 1. This ambiguity should be resolved, for example by renaming the setting or explicitly distinguishing the method name from the setting name.","section":"§2.2(d), Table 1"},{"comment":"The table reports accuracy numbers from different papers without variance or statistical significance information, and the 'Learn. Param.' column contains dashes for several PEFT methods. At minimum, the caption should state that the values are copied from the original papers and are not directly comparable; ideally, the authors would add citations or footnotes for each row.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The survey is positioned as a comprehensive reference, and the 'up to 20%' sentence in §4.6 is likely to be quoted by future readers. I would ask the editor to ensure that the revision addresses the confounded comparison directly rather than merely adding a caveat. Also, several future-direction items are tied to the authors' own in-press works (FFSCIL, vision-language synergy); this is not improper, but the writing should make clear that these are emerging proposals, not established subfields."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Roughly put: the survey is a useful map, and the taxonomy is worth having, but the headline performance claim about PEFT superiority is not supported by the table it is based on. PEFT methods run on ImageNet-pretrained ViT-B/16 while the other four approaches mostly use ResNet18 trained from scratch (except on CUB), so the up-to-20% gap is almost certainly confounded with backbone and pretraining. The paper even says this about the MiniImageNet numbers, then draws the contrary conclusion.\n\nWhat is genuinely good: the five-approach taxonomy (backbone tuning, meta-learning, prototype tuning, dynamic architecture, PEFT) and the five-settings classification (supervised, semi-supervised, unsupervised, complete, cross-domain) organize the literature better than the earlier Tian et al. survey. The prototype-rectification categorization into loss, pseudo-prototype, training-free calibration, and trainable network is a nice lens. The coverage of recent PEFT and language-guided methods (ASP, PriViLege, FineFMPL, ApproxFSCIL) is timely, and the attention to base vs novel class accuracy and AHM is a real service to the community.\n\nSoft spots, in order of weight. First, the PEFT superiority claim in Section 4.6 is load-bearing and confounded. The table includes L2P, DualP, and CodaP—rehearsal-free CIL methods, not FSCIL methods—and Table 2 shows 0% novel-class accuracy for them. That weakens the comparison further. Second, the survey has no methodology section; there are no inclusion/exclusion criteria, so 'comprehensive' is unverifiable. Third, the claimed coverage of SSFSCIL is not backed: it is defined in Section 2 but no SSFSCIL method appears in Tables 1 or 3. Fourth, there are small internal errors: a cross-reference to Eq. 2 when Eq. 4 is meant, a sentence cut off at the end of Section 4.7, and some garbled phrases. Table 1 has blanks and no variance/significance information, which the authors should acknowledge.\n\nThe formal objectives (Eqs. 2-5) are near-restatements of Eq. 1 with different parameter sets; that is fine as a pedagogical device but not a mathematical contribution. The self-citations to the authors' own UFSCIL and FFSCIL work are not a problem per se, but they do shape the 'future directions' section; a reader should know those are not established subfields yet.\n\nWho is this for? Someone new to FSCIL who wants a structured map and a recent pointer list. A serious referee should engage with it, mostly to force the performance claims to be re-analyzed and the missing methodology made explicit. I would not desk-reject, but I would not accept as is.","headline":"A genuinely useful organizing topology for FSCIL, but the PEFT-superiority headline is confounded and the survey needs methodological tightening.","tokens_in":19630,"tokens_out":4320,"would_cite":true,"duration_ms":29931,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey structures few-shot class-incremental learning into five approaches and five data settings, and argues that prompt-based parameter-efficient tuning on pre-trained models now leads the field.","keywords":["few-shot class-incremental learning","catastrophic forgetting","prototype rectification","parameter-efficient fine-tuning","prompt learning","pre-trained vision-language model","continual learning survey","stability-plasticity"],"falsifier":"A controlled re-benchmarking study that fixes the backbone (e.g., ViT-B/16 for all methods), fixes the pre-trained model, and uses the same session splits and evaluation code for at least one representative method from each of the five families would settle whether the up-to-20% PEFT advantage holds; if the gap shrinks to near zero, the survey's main ranking claim collapses.","tokens_in":18543,"feed_emoji":"🧠","tokens_out":2426,"duration_ms":21588,"temperature":0.7,"pith_summary":"The paper sets out to organize the sprawling few-shot class-incremental learning (FSCIL) literature into a clear topology: five method families (backbone tuning, meta-learning, prototype tuning, dynamic architecture, and parameter-efficient fine-tuning) and five problem settings (supervised, semi-supervised, unsupervised, complete, and cross-domain). It gives each family a formal objective and compares methods not only on average accuracy but also on stability-plasticity balance and base-versus-novel class performance. The survey's main empirical claim is that PEFT methods built on frozen pre-trained vision transformers outperform the other four families by a margin of up to 20% average accuracy, while using far fewer trainable parameters. It also argues that prototype rectification is a central and under-appreciated mechanism for countering data scarcity. A reliable map of this kind matters because it tells researchers which directions are paying off and where the open problems actually lie.","feed_headline":"Survey: prompt-tuned pre-trained models lead few-shot class-incremental learning","feed_subtitle":"Five method families, five data settings, and a key role for prototype rectification in beating catastrophic forgetting.","key_machinery":"The central organizing device is a five-way taxonomy of FSCIL approaches, each with an explicit objective function (equations 2 through 7), paired with a classification of problem settings (fully supervised, semi-supervised, unsupervised, complete, and cross-domain). Within the taxonomy, the load-bearing concept is the prototype-based classifier, which stores one vector per class and compares new inputs against those prototypes; unlike network classifiers, prototypes are not updated in later tasks, so they are naturally resistant to forgetting. The paper's second machinery is a topology of prototype rectification, distinguishing loss-driven updates, pseudo-prototype generation, training-free calibration, and trainable calibrator networks, and it uses this taxonomy to explain why PEFT-plus-prototype hybrids such as L2P+, DualP+, and CodaP+ markedly outperform their original prompt-only versions.","core_discovery":"The paper argues that the FSCIL field has moved through four earlier approaches that tune or extend a backbone trained on a large base task, and into a fifth, PEFT-based approach that keeps a pre-trained transformer frozen and learns only small prompt parameters. It finds that PEFT methods not only reduce trainable parameters from tens of millions to under a million but also lift average accuracy substantially, with the best results exceeding 90% on MiniImageNet and 80% on CUB, because the pre-trained model supplies the general knowledge that the few-shot tasks cannot provide. At the same time, the survey shows that almost all methods, including the best PEFT ones, still achieve far lower harmonic mean accuracy than average accuracy, indicating that plasticity on novel classes remains weak and that stability-plasticity is still an open problem. It further identifies prototype rectification—correcting biased class prototypes estimated from very few samples—as a key mechanism, and classifies existing rectification strategies into four types: loss-based, pseudo-prototype, training-free calibration, and trainable-network calibration.","pith_inferences":["The claimed up-to-20% gap between PEFT and the other approaches is likely confounded by base model scale and pre-training data: PEFT methods use ViT-B/16, which was itself pre-trained on ImageNet, the superset of the MiniImageNet test set, so part of the gain may be transfer leakage rather than the prompting mechanism per se.","A controlled comparison that keeps backbone, pre-training, and protocol identical across the five families would be the natural next experiment, and the survey's own table does not currently enable that because all non-PEFT methods use ResNet18.","The observation that almost no applied FSCIL system uses PEFT or language guidance suggests that the technique gap in real-world domains such as medical imaging and audio may be larger than the benchmark gap, and that transferring the PEFT-plus-prototype recipe to those domains is a testable extension.","The four-way rectification topology could be reused as an evaluation axis for any few-shot or continual-learning method, not only FSCIL, since biased prototypes are a general consequence of data scarcity."],"forward_implications":["If the PEFT advantage is real, future FSCIL work should default to frozen pre-trained backbones with prompt learning and a prototype-based classifier, rather than tuning full backbones from scratch.","Prototype rectification becomes a priority design choice, with training-free calibration offering a low-cost baseline that many current PEFT methods do not yet exploit.","Reporting average accuracy alone is misleading; the survey's stability-plasticity analysis implies that base-versus-novel class accuracy and harmonic mean accuracy should become standard reporting metrics.","The five-setting taxonomy gives practitioners a way to match a method to their actual data regime, e.g., choosing supervised FSCIL when labels are complete and cross-domain or semi-supervised variants when they are not.","Open challenges identified in the survey—federated FSCIL, online data streams, open-world detection, and class imbalance—point to concrete next problems that the current methods are not yet built to handle."],"supporting_citations":[{"why":"Introduces the FSCIL problem and defines the standard supervised setting that the whole survey builds on.","marker":"[Tao et al., 2020]"},{"why":"Proposes L2P, the prompt-learning method that anchors the PEFT approach and defines pool-based prompt structure.","marker":"[Wang et al., 2022c]"},{"why":"Establishes the prototype-bias problem and motivates the rectification taxonomy that the survey centers on.","marker":"[Liu et al., 2020]"},{"why":"Introduces language-guided learning in continual learning, which the survey identifies as the second major advancement.","marker":"[Khan et al., 2023]"},{"why":"Provides the CLIP vision-language model that PEFT and language-guided FSCIL methods use for their representations.","marker":"[Radford et al., 2021]"},{"why":"Defines constrained FSCIL, the basis for the complete and cross-domain settings that the survey extends.","marker":"[Hersche et al., 2022]"},{"why":"The previous FSCIL survey that the paper positions itself against, lacking PEFT, sub-settings, and prototype rectification coverage.","marker":"[Tian et al., 2024]"}],"fun_headline_variants":["Prompt-tuning leads FSCIL, but novel-class plasticity still weak","Survey: PEFT dominates FSCIL accuracy, yet stability gap persists","FSCIL: pre-trained prompt models beat finetuning, but open challenges remain","Prototype rectification key in FSCIL; PEFT now the leading paradigm","Few-shot class incremental learning: PEFT methods show top accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's central comparisons assume that the accuracy numbers collected from different papers are directly comparable even though the methods use different backbones, different pre-trained models or none, and possibly different evaluation protocols.","fun_headline_variants_meta":{"raw":{"variants":["Prompt-tuning leads FSCIL, but novel-class plasticity still weak","Survey: PEFT dominates FSCIL accuracy, yet stability gap persists","FSCIL: pre-trained prompt models beat finetuning, but open challenges remain","Prototype rectification key in FSCIL; PEFT now the leading paradigm","Few-shot class incremental learning: PEFT methods show top accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000386,"raw_usage":{"total_tokens":2014,"prompt_tokens":893,"completion_tokens":1121,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":1023}},"tokens_in":509,"tokens_out":1121,"duration_ms":9403,"temperature":1.0,"reasoning_tokens":1023,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T10:07:06.702527+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled re-benchmarking study that fixes the backbone (e.g., ViT-B/16 for all methods), fixes the pre-trained model, and uses the same session splits and evaluation code for at least one representative method from each of the five families would settle whether the up-to-20% PEFT advantage holds; if the gap shrinks to near zero, the survey's main ranking claim collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the FSCIL problem and defines the standard supervised setting that the whole survey builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prototype-bias problem and motivates the rectification taxonomy that the survey centers on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces language-guided learning in continual learning, which the survey identifies as the second major advancement."},{"cited_title":"Radford, J","cited_arxiv_id":null,"evidence_quote":"Provides the CLIP vision-language model that PEFT and language-guided FSCIL methods use for their representations."},{"cited_title":"Hersche, G","cited_arxiv_id":null,"evidence_quote":"Defines constrained FSCIL, the basis for the complete and cross-domain settings that the survey extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The previous FSCIL survey that the paper positions itself against, lacking PEFT, sub-settings, and prototype rectification coverage."}],"review_version":1}