{"id":"d1a077e2-f507-4f04-879d-1daffb26c840","arxiv_id":"2506.12413","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic survey that categorizes domain-generalizable person re-identification methods and compares their cross-domain performance.","lead":"This survey maps the young field of domain-generalizable person re-identification, sorting methods into seven families and comparing their reported accuracy on standard benchmarks. It is the first dedicated synthesis of this subfield according to the authors, making it a potential entry point for newcomers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 3 does not implement Protocol-1: per-row source sets differ (M+D, C2+C3+CS, M+C2+C3+CS), so ReNorm/BAU's reported lead is not a controlled comparison.","rationale":"The reader's weakest assumption points at transcription accuracy and comparability of the three evaluation protocols, and identifies that Protocol-1 mixes different source-domain sets across methods. My stress-test concurs and sharpens the issue: the problem is visible in the table itself, not merely a transcription risk. Table 3 lists per-row source sets that contradict the Section 5.3 definition of Protocol-1, including a row with the withdrawn Duke dataset (M+D) despite Section 5.4 saying Duke is excluded. Because the survey's main empirical contribution is the state-of-the-art ranking in Section 5.4, this is the single most load-bearing concern. The survey's organizational content—taxonomy, dataset descriptions, and future-directions discussion—remains useful and does not depend on the contested numbers, so the appropriate disposition is conditional acceptance pending table reconstruction rather than rejection. No ad hominem is intended; the issue is an evidentiary one in the comparison tables. I recommend no verdict change from the reader's CONDITIONAL, since the concern confirms rather than alters that conditional stance.","tokens_in":32892,"tokens_out":5191,"duration_ms":63029,"concrete_test":"Reconstruct Table 3 from the primary sources: for every listed method, record the exact source-domain set used in the original paper and the reported PRID, GRID, VIPeR, and iLIDS mAP/R1 values. Restrict the comparison to methods trained on the exact Protocol-1 source set (M+C2+C3+CS, with Duke excluded) and recompute the averages; also spot-check five randomly selected cells directly against the PDFs of the cited papers. Separately, check M3L's original paper for Protocol-2/3 results and resolve the contradiction between Section 5.4 and Table 4. If ReNorm still beats BAU within the common source set after these checks, the ranking stands; otherwise the Section 5.4 conclusions must be qualified as comparing incomparable protocols.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claims—ReNorm and BAU lead the field, performance is approaching saturation—rest on Table 3 and Table 4. But Table 3 is not a Protocol-1 comparison as Protocol-1 is defined in Section 5.3. That section fixes the training set to Full-(M+C2+C3+CS), yet Table 3 lists different source combinations per row: 'M+D' for SNR, 'C2+C3+CS' for DMG-Net, 'M+C2+C3+CS' for M3L, and so on. Some rows even include DukeMTMC ('D'), which Section 5.4 explicitly says is excluded from training. Training on three, four, or five source datasets, with or without the withdrawn Duke set, is a different task; the claimed 1.0% mAP and 2.7% Rank-1 advantage of ReNorm over BAU could therefore be a source-mix artifact rather than evidence of method superiority. The same uncontrolled comparison is used to support the 'approaching saturation' claim. In addition, Section 5.4 states that M3L is excluded from the Protocol-2/3 comparison, but Table 4 includes M3L rows under both Protocol-2 and Protocol-3. This internal inconsistency indicates that the performance tables were assembled without a consistent provenance audit, which is exactly the condition needed for the survey's headline conclusions to be reliable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a survey of domain-generalizable person re-identification (DG-ReID), a setting where models are trained on multiple source domains and evaluated on unseen target domains without any target-domain data. The authors review the standard architecture components (backbones, multi-source input configurations), propose a taxonomy of DG-ReID methods (normalization-based, mixture-of-experts, memory-based, meta-learning, data-driven, CLIP-based, and others), provide a case study on visible-infrared ReID, and assemble performance comparisons under three evaluation protocols. The paper claims to be the first systematic survey dedicated to DG-ReID and concludes that ReNorm and BAU are state-of-the-art, with performance approaching saturation on some benchmarks.","tokens_in":33142,"tokens_out":5958,"duration_ms":65246,"significance":"If the empirical comparisons are reliable, the survey provides a useful and timely map of a rapidly growing field. The taxonomy is broadly sensible, the attention-map analysis and the VI-ReID case study are valuable additions, and the authors ship a curated resource list on GitHub. However, the central empirical claims rest on performance tables that are not internally consistent: Table 3 does not implement the Protocol-1 defined in Section 5.3, and Section 5.4 contradicts Table 4 regarding M3L. These issues undermine confidence in the headline conclusions about which methods lead the field. The survey's taxonomic and qualitative contributions are significant, but the performance analysis requires major revision.","major_comments":[{"comment":"Table 3 does not implement Protocol-1 as defined in Section 5.3 and Table 2. The protocol fixes the training set to Full-(M+C2+C3+CS), but individual rows in Table 3 list different source mixes: 'M+D' for SNR, 'C2+C3+CS' for DMG-Net/RaMoE/MDA/DTIN-Net, and 'M+C2+C3+CS' for M3L/META/MetaBIN/ACL/ISR/PAOA/BAU/ReNorm, with DIMN lacking a source entry entirely. Training on three, four, or five source datasets, and with or without the withdrawn DukeMTMC set, constitutes a different task; the claimed 1.0% mAP and 2.7% Rank-1 advantage of ReNorm over BAU in Section 5.4 could therefore be a source-mix artifact rather than evidence of method superiority. Please either restrict Table 3 to rows that actually use the Protocol-1 training set or explicitly state that the comparison is not controlled and adjust the claims accordingly.","section":"§5.3 and Table 3"},{"comment":"The text in Section 5.4 states: 'DIMN, DMG-Net, M3L, RaMoE, DTIN-Net and MDA are excluded from this comparison due to no results reported for Protocol-2 and -3.' However, Table 4 includes M3L rows under both Protocol-2 and Protocol-3. This is a direct internal inconsistency. It indicates that the performance tables were assembled without a consistent provenance audit, which is exactly the condition needed for the survey's headline performance conclusions to be reliable. Please correct the text or the table and re-verify each entry against the cited source.","section":"§5.4 and Table 4"}],"minor_comments":[{"comment":"Reference [28] misspells 'CLIP-FGDI' as 'CILP-FGDI'; please correct the title.","section":"Reference [28]"},{"comment":"The iLIDS dataset is attributed to reference [88] with venue 'ECCV'18', but that reference is a paper on unsupervised tracklet association, not the original source of the iLIDS dataset, and the venue does not match the original iLIDS publication. Please cite the correct dataset reference and update the venue in Table 1.","section":"Table 1 / Reference [88]"},{"comment":"ReNorm is listed as 'ECCV'24' in Tables 3 and 4 but reference [31] says ECCV 2025; BAU appears as both 'NeurIPS'24' and 'NIPS'24' across the tables. Please make venue and year entries consistent between the tables and the bibliography.","section":"Tables 3 and 4 vs. bibliography"},{"comment":"Section 3 says methods are categorized into 'eight methodological aspects,' then lists seven categories; Figure 2 and Section 7 consistently refer to seven modules. Please align the count.","section":"§3, §7, and Figure 2"},{"comment":"The text says 'As visualized in Fig. 9,' but the corresponding figure is labeled 'Figure 8: Visualization of frequency decomposition results.' Please correct the cross-reference.","section":"§6.3, Figure 8"},{"comment":"The shared column headers 'M+MS+CS→C3' etc. obscure the difference between Protocol-2 (only training subsets) and Protocol-3 (full source domains); please add explicit phrasing in the caption, such as 'Full-' for Protocol-3, to avoid confusion.","section":"Table 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The taxonomic and qualitative parts of the survey are valuable and likely within the scope of the venue. The main concern is the reliability of the performance tables, which are not internally consistent and do not support the controlled comparison claimed. I would advise the editor to ask for a careful re-audit of Tables 3 and 4 before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuine first: a dedicated survey of domain-generalizable person re-identification. The module-based taxonomy (normalization, mixture-of-experts, memory, meta-learning, data-driven, CLIP, others) is reasonable, and the overview of backbones, multi-source configurations, and datasets gives newcomers a solid entry point. The VI-ReID case study is a useful cross-task addition, and the curated GitHub list is a plus. The writing is clear, and the methodological descriptions are mostly faithful to the cited works.\n\nThe soft spots are concentrated in the performance comparison sections, and they are not minor. Protocol-1 is defined in Section 5.3 as training on Full-(M+C2+C3+CS), but Table 3 lists different source mixtures per row (M+D for SNR, C2+C3+CS for DMG-Net, M+C2+C3+CS for M3L, and so on), with some rows including DukeMTMC even though Section 5.4 says it is excluded. The claimed lead of ReNorm over BAU (1.0% mAP, 2.7% Rank-1) could therefore be a training-data artifact rather than a method advantage. Similarly, Section 5.4 says M3L is excluded from Protocol-2/3 because no results were reported, yet Table 4 includes M3L rows under both protocols. These inconsistencies directly undermine the headline empirical claims, including the saturation argument. There are also mechanical citation errors: iLIDS is attributed to an ECCV'18 tracklet-association paper rather than the actual dataset source, reference [28] spells the method name as 'CILP-FGDI', and ReNorm's venue/year are inconsistent between the bibliography (ECCV 2025) and the tables (ECCV'24). A careful provenance pass on the tables and references is needed.\n\nThat said, the taxonomy and protocol definitions are the real contribution, and the paper does not need a new method to be useful. The qualitative attention maps in Figure 7 are clearly illustrative, not a new scientific result, so they do not bother me.\n\nI would send this to peer review with a request for major revision on the tables and references. It deserves a serious referee; the errors are fixable, and the surveyed content is otherwise solid. I would not cite the performance tables until they are corrected, but I would point students to the taxonomy and the GitHub list.","headline":"Useful first survey of DG-ReID with a sensible taxonomy, but the headline performance tables don't implement the stated protocols, so the 'ReNorm leads' conclusion is not supported.","tokens_in":33674,"tokens_out":2148,"would_cite":false,"duration_ms":25179,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents the first systematic survey of domain-generalizable person re-identification, organizing the field into seven method families and comparing their performance.","keywords":["person re-identification","domain generalization","domain-invariant representation learning","image-based retrieval","normalization-based methods","meta-learning","CLIP-based methods","evaluation protocols"],"falsifier":"Re-run every method in Tables 3 and 4 from its released code under a single fixed source-domain mix for each protocol and compare the reproduced mAP and Rank-1 values; if the top performers change or the average numbers differ materially from the table, the survey's leaderboard conclusions (ReNorm and BAU lead; performance is approaching saturation) would be overturned.","tokens_in":32677,"feed_emoji":"🕵️","tokens_out":5932,"duration_ms":66366,"temperature":0.7,"pith_summary":"This paper tries to establish a reliable map of Domain-Generalizable Person Re-identification (DG-ReID), the setting where a model learns to match people using several labeled source domains and is then tested on an unseen domain with no target data. It claims to be the first full survey dedicated to DG-ReID, grouping methods into seven families and comparing them under three standard evaluation protocols. If the map is correct, researchers gain a shared vocabulary, a baseline leaderboard, and evidence about which designs actually generalize. The survey also argues that current leaders are approaching saturation on small test sets, while harder targets such as MSMT17 still leave room for progress.","feed_headline":"First survey maps domain-generalizable person re-ID","feed_subtitle":"Seven method families, three evaluation protocols, and a clear view of which approaches currently lead.","key_machinery":"The load-bearing object is the formal DG-ReID setting itself: $K$ source domains $D_S = \\{D_1, \\ldots, D_K\\}$ with disjoint identity label spaces ($Y_i \\cap Y_j = \\emptyset$), which rules out direct cross-domain identity matching and forces the model to learn features that are both domain-invariant and identity-discriminative. Around this setting the survey builds two analytical instruments: a seven-category, module-centric taxonomy of methods, and a three-protocol evaluation scheme (Protocol-1 tests on four small unseen datasets; Protocol-2 and Protocol-3 leave out one large dataset as target). These instruments do the work of converting a scattered literature into a leaderboard, with Tables 3 and 4 carrying the comparative evidence.","core_discovery":"The paper's central claim is that existing DG-ReID research can be organized around a single pipeline—a backbone (usually ResNet-50, MobileNetV2, or ViT) trained on multiple source domains—plus one or more generalization modules, and that those modules fall into seven categories: normalization-based, mixture-of-experts-based, memory-based, meta-learning-based, data-driven, CLIP-based, and others. It backs this taxonomy with a performance comparison under three protocols, finding that normalization-focused ReNorm leads on Protocol-1 and Protocol-3, BAU and ReNorm lead on Protocol-2, and that MSMT17 remains the hardest target. The paper also treats DG-ReID as a heterogeneous domain-generalization problem with disjoint identity label spaces across source domains, and claims that its techniques carry over to the related VI-ReID task.","pith_inferences":["Editorial inference: the seven categories are not mutually exclusive—several leading methods combine normalization with meta-learning or mixture-of-experts structures—so a future survey might reorganize the field by training objective (invariance, diversity, alignment) rather than by module type.","Editorial inference: because Protocol-1 mixes different source-domain sets across rows of Table 3, the leaderboard may compare methods trained on different amounts of data; a common-source re-run would be needed to fully trust the rankings.","Editorial inference: the attention-map comparison suggests a measurable diagnostic—person-centric attention share—that could be added to DG-ReID evaluation beyond the qualitative example shown.","Editorial inference: the low MSMT17 mAP hints that target-domain difficulty, not source diversity alone, sets the ceiling for DG-ReID, which would make harder benchmark design as important as new modules."],"forward_implications":["New DG-ReID papers can position their contribution against a named category (normalization, mixture-of-experts, memory, meta-learning, data-driven, CLIP, or other) and compare with the same standard protocols.","Normalization-based designs, especially ReNorm's remix and emulation normalization, currently represent the strongest surveyed approach on the average of Protocol-1 and Protocol-3.","The three evaluation protocols give the field a common yardstick, with Protocol-2 the most severe domain shift and Protocol-3 the most realistic deployment scenario.","MSMT17 remains a bottleneck: even the best Protocol-3 method reaches only 27.8% mAP there, so claims of saturation need qualification.","The VI-ReID case study suggests that DG-ReID modules (normalization, memory, gradient reversal, graph matching) transfer across tasks that share a domain-gap structure."],"supporting_citations":[{"why":"Defines the SNR baseline and the Protocol-1 evaluation practice that the survey's comparisons follow.","marker":"[16]"},{"why":"MetaBIN is the reference learnable BN/IN combination and appears throughout the normalization and meta-learning discussion.","marker":"[17]"},{"why":"ReNorm is the top performer in the survey's Tables 3 and 4, carrying the claim that normalization-based methods lead the field.","marker":"[31]"},{"why":"BAU is the second-best method overall and the strongest non-normalization approach, supporting the saturation and leaderboard conclusions.","marker":"[18]"},{"why":"ACL is a fixed-combination normalization baseline and the method used in the attention-map comparison on MSMT17.","marker":"[30]"},{"why":"M3L provides the memory-based meta-learning baseline that anchors the meta-learning and Protocol-2/3 comparisons.","marker":"[21]"},{"why":"META is the shared-expert mixture-of-experts method, load-bearing for the MoE category and the Protocol comparisons.","marker":"[32]"},{"why":"RaMoE is the independent-expert MoE method that grounds the expert-category discussion.","marker":"[34]"}],"fun_headline_variants":["First systematic survey on domain-generalizable person re-ID","Re-ID without target data: survey maps 7 method families","Domain-agnostic ReID: first survey compares 3 protocols","Survey reveals leading methods for unseen-domain ReID","Seven ways to generalize ReID: a comprehensive survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the numbers in the comparison tables were copied correctly from the original papers and that the three evaluation formats are comparable across methods, even though the rows of Table 3 use different training-dataset combinations while the caption presents them as one common protocol.","fun_headline_variants_meta":{"raw":{"variants":["First systematic survey on domain-generalizable person re-ID","Re-ID without target data: survey maps 7 method families","Domain-agnostic ReID: first survey compares 3 protocols","Survey reveals leading methods for unseen-domain ReID","Seven ways to generalize ReID: a comprehensive survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00039,"raw_usage":{"total_tokens":2080,"prompt_tokens":998,"completion_tokens":1082,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":1002}},"tokens_in":614,"tokens_out":1082,"duration_ms":11862,"temperature":1.0,"reasoning_tokens":1002,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:50:50.793029+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run every method in Tables 3 and 4 from its released code under a single fixed source-domain mix for each protocol and compare the reproduced mAP and Rank-1 values; if the top performers change or the average numbers differ materially from the table, the survey's leaderboard conclusions (ReNorm and BAU lead; performance is approaching saturation) would be overturned.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ReNorm is the top performer in the survey's Tables 3 and 4, carrying the claim that normalization-based methods lead the field."},{"cited_title":"Generalizable Person Re-identification via Balancing Alignment and Uniformity","cited_arxiv_id":"2411.11471","evidence_quote":"BAU is the second-best method overall and the strongest non-normalization approach, supporting the saturation and leaderboard conclusions."},{"cited_title":"Zhang, S","cited_arxiv_id":null,"evidence_quote":"ACL is a fixed-combination normalization baseline and the method used in the attention-map comparison on MSMT17."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"META is the shared-expert mixture-of-experts method, load-bearing for the MoE category and the Protocol comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"RaMoE is the independent-expert MoE method that grounds the expert-category discussion."}],"review_version":1}