REVIEW 4 major objections 5 minor 77 references
Cerberus: Attribute-based person re-identification using semantic IDs
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Attribute labels, compiled into semantic IDs, let a single model outperform prior attribute-based person re-identification and also recognize attributes and search by attribute queries.
desk verdict Solid attribute-based reID paper with a genuinely useful unified framework; the SOTA claim needs a small correction, and the single-image label assumption deserves a robustness check, but the core idea holds up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the semantic ID (SID): every combination of attribute labels inside one of five groups—head, upper body, lower body, identity, carryings—so, for example, 'short red dress' is one SID. For each group the model learns a separate 512-dimensional person representation and one prototype vector per SID. The semantic guidance loss $\mathcal{L}_{sem}$ pulls each partial representation toward its SID prototype until similarity exceeds $1 - m^G_g$, where the margin $m^G_g = \log(\alpha N^G_g/N + \beta)$ grows with the number of people sharing that SID; the identification term $\mathcal{L}_{id}$ then separates same-SID identities. A regularization loss constrains prototype differences to match learned residual vectors $\mathbf{r}_{m,n} = \sum_l \mathbf{v}_l(A^G_m(l)-A^G_n(l))$, so pairs of prototypes that differ by the same attribute labels share the same residual, which lets the model place unseen SIDs such as 'old female' next to their nearest seen relatives.
What would settle it
Take a trained Cerberus model and build a test set in which the same identities appear in two outfits, with attribute labels copied from the first outfit. If reID mAP and the by-product PAR accuracy degrade sharply on the changed-outfit images compared with labels annotated per image, the single-image label assumption is the bottleneck. The same test can be run by flipping one attribute label per identity (e.g., 'backpack' to 'no backpack') and measuring the drop in all three tasks.
Extended reading notes
Core claim
The paper's central claim is that the long-standing conflict between person re-identification and person attribute recognition can be resolved by making attribute combinations themselves the units of embedding. Each attribute group defines a set of SIDs, and the model learns a prototype vector for every SID; a semantic guidance loss pulls that group's partial representation toward the correct prototype only up to an adaptive margin, while an identification loss continues to separate same-SID different-identity representations. The reported results are 89.8% mAP and 96.1% rank-1 on Market-1501, and 80.7% mAP and 91.3% rank-1 on DukeMTMC-reID, exceeding prior attribute-based reID methods. The same prototypes act as nearest-neighbor classifiers for attribute recognition (91.1% mean accuracy on Market-1501 among attribute-based reID methods) and as text-like queries for attribute-based person search, with partial attribute queries supported because each partial representation is aligned in its own embedding space.
Load-bearing premise
The assumption that carries the method is that a person's attribute labels, annotated once from a single image, stay correct for every image of that person across cameras; if an outfit or carried item changes, the semantic guidance loss pulls that image's representation to the wrong SID prototype.
Editorial extensions
If this is right
- If the reported numbers hold, attribute labels—which the paper notes are cheaply annotated from one image per person—are enough to match or beat methods that require pose estimators, body-parsing masks, or extensive extra supervision.
- A single trained model covers reID, PAR, and APS, so deployment for a surveillance pipeline would not need a separate attribute classifier or text-query network.
- Partial attribute queries, such as 'man carrying a backpack' without clothing details, become possible because each group's representation and prototype live in an independent embedding space.
- Prototype regularization specifically helps when test images contain SIDs absent from training, shrinking the zero-shot gap that hurts prior attribute-based models.
- The adaptive margin allocates more identification pressure to SIDs with many members, so common looks receive extra focus on subtle details like logos and pocket shapes.
Reading between the lines
- The paper leaves implicit that the regularization term turns the prototype space into a nearly linear attribute-compositional space; if that is true, unseen attribute combinations could be generated by adding residual vectors rather than retraining, a testable extension.
- Because SID assignment rests on one annotated image per identity, an obvious failure mode is clothing or accessory change across cameras; an experiment that corrupts attribute labels with realistic outfit-change noise would quantify how much of the reported gain depends on label stability.
- The same grouping-into-SIDs recipe could transfer to other attribute-structured retrieval domains, such as vehicle make/color search or clothing retrieval, where attributes are few but combinatorial.
- The paper's hyperparameters were selected on a Market-1501 validation split, so a concrete pressure-test is to train on Market-1501 and evaluate on DukeMTMC-reID without retuning to see how generic the margin formula is.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Cerberus, an attribute-based person re-identification framework. It groups attribute labels into five categories (head, upper body, lower body, identity, carrying), defines semantic IDs (SIDs) as combinations of attributes within each group, and learns a prototype vector per SID. A semantic guidance loss (Eq. 3) aligns partial person representations with the corresponding prototypes, an identification loss (Eq. 5) preserves ID-level discrimination, and a regularization loss (Eqs. 6-8) constrains relations among prototypes to handle unseen SIDs. At inference, reID compares five partial representations between query and gallery; APS replaces text attributes with SID prototypes; PAR uses the prototypes as nearest-neighbor classifiers. Experiments on Market-1501 and DukeMTMC-reID report reID, PAR, and APS numbers, with ablations and hyperparameter sensitivity studies on Market-1501.
Significance. If the results hold, the main contribution is a single framework that uses cheap attribute labels to improve image-based person reID while producing by-product PAR and APS capabilities with negligible extra parameters and without task-specific fine-tuning. The core ablations in Table 6 support the individual contributions of the semantic guidance term, the regularization term, and the alignment module, and the efficiency comparison against MGN and GPS is useful. The by-product tasks are substantively genuine in that SID prototypes are not used for reID at test time. However, the paper overstates its benchmark status: Table 2 shows CLIP-ReID exceeds Cerberus on Duke mAP (82.5 vs. 80.7), so the "new state of the art" claim is not supportable over all listed methods. In addition, the partial-APS capability, which is a stated novelty, has no quantitative evaluation. The method is a sound and reasonably well-ablated engineering contribution, but the writing and evaluation need tightening before the central claims are fully supported.
major comments (4)
- [Section 4.2.1 / Table 2 / Abstract] The paper repeatedly states that Cerberus "sets a new state of the art" and quotes 80.7% mAP and 91.3% rank-1 on DukeMTMC-reID. Table 2 lists Cerberus's Duke rank-1 as 91.1, and the general reID method CLIP-ReID in the same table achieves Duke mAP 82.5, which is higher than Cerberus's 80.7. Thus the SOTA claim is false over all methods in the table and should be restricted to “state of the art among attribute-based reID methods.” The text's statement that Cerberus outperforms the second-best mAP by 0.7% is also contradicted by Table 2, where the mAP difference against CLIP-ReID is -1.8. Please correct the quoted numbers and the scope of the SOTA claim.
- [Section 4.1.1 / Eq. (3)] The framework depends on the assumption that the attribute labels annotated on one image per identity are valid for all images of that identity. The paper states in Section 4.1.1 that Lin et al. annotate a single image per person and assume personal traits do not vary significantly across cameras, but this assumption is load-bearing for Cerberus: Eq. (3) pulls every image's partial representation toward the prototype of the SID derived from those single-image labels. If clothing or carried items change across views, many images receive incorrect SID supervision, and the prototype itself is trained on mixed evidence. Since PAR (Eq. 9) and APS use the same prototypes, the by-product claims inherit this risk. The paper provides no label-consistency statistics on these benchmarks and no robustness experiment such as label flipping or removal of time-varying attributes. Please add such an analysis or clearly discuss the limitation and its expected effect on the reported numbers.
- [Section 4.2.2 / Table 5] Partial APS is presented as a novel capability ("we can even search persons with partial text queries") and as part of the unified-model contribution. However, the only evidence is qualitative in Figure 6; Table 5 evaluates full attribute queries only. A quantitative protocol is needed, for example reporting rank-1/mAP over random subsets of attribute groups, or ablating which groups are retained. Without this, the partial-query claim is not supported by measurement and should either be added to the evaluation or substantially softened.
- [Section 3.2.2 / Eq. (8)] The regularization loss assumes that the difference between two SID prototypes is a linear combination of per-attribute residual vectors v_l that is shared across all attribute differences. This factorization is an ad hoc modeling choice, and the only validation is the aggregate ablation in Table 6 and Figure 7(b). The paper does not report whether the residual assumption holds per group (e.g., carrying vs. identity) or whether the benefit is concentrated in particular attribute groups. Since unseen-SID generalization is one of the two stated technical contributions, please provide per-group regularization ablations or a direct check of residual consistency across prototype pairs.
minor comments (5)
- [Section 4.2.1 / Table 2] The abstract and Section 4.2.1 quote Duke rank-1 as 91.3, but Table 2 reports 91.1 for Cerberus; please harmonize all occurrences of the Duke numbers.
- [Section 1] The claim "this is the first model that can perform reID, PAR, and APS tasks without fine tuning for each task" should be qualified in light of UPAR (Specker et al., 2023), which the paper itself cites and compares against and which jointly addresses PAR and person retrieval; the distinction (image-query reID versus attribute-query retrieval) should be stated explicitly.
- [Figure 7(b)] The plot showing the regularization effect as a function of the number of unseen SIDs lacks axis labels and units; as printed it is difficult to read the quantitative trend. Please add axis labels and a legend.
- [Table 6] The ablation table is not fully self-explanatory: the first row already includes the identification loss, and the checkmarks are added cumulatively. Please state explicitly in the caption that all rows include L_id and that each subsequent row adds the indicated component.
- [Section 4.3.1] The sentence "Linet.al. assume that personal traits would not significantly vary across cameras" contains a typographical error and should read "Lin et al."; please fix this and similar spacing issues throughout.
Circularity Check
No significant circularity: Cerberus's reID, PAR, and APS results are genuine held-out evaluations of a supervised model, not reductions to its inputs.
full rationale
The central derivation chain is not circular. Cerberus defines SIDs as combinations of attribute labels and trains SID prototypes with a semantic guidance loss (Eq. 3) alongside an identification loss (Eq. 5); at test time the reID similarity uses only person representations, and the paper states that 'we do not use them [SID prototypes] at test time' (Section 3.3). Because the prototypes are learnable parameters supervised by attribute labels, using them for PAR and APS is a standard nearest-neighbor or retrieval procedure on held-out test identities, not a fitted input renamed as a prediction. The paper explicitly tests unseen-SID generalization by holding out SIDs during training (Fig. 7b), so the regularization claim is evaluated rather than assumed. The only self-citation, ISGAN (Eom & Ham, 2019), appears as a comparison baseline in Table 2 and is not used to justify any load-bearing premise. The single-image-per-person attribute annotation assumption noted in Section 4.1.1 is a robustness weakness, not a circularity; it concerns possible label noise across cameras, not an equivalence between outputs and inputs.
Assumptions & free parameters
free parameters (6)
- alpha (boundary margin slope) =
0.4
- beta (boundary margin bias) =
1.8
- sigma (alignment threshold) =
5
- lambda_sem (semantic loss weight) =
5
- lambda_reg (regularization loss weight) =
0.001
- lambda_id (identification loss weight) =
1
assumptions (4)
- domain assumption Person attribute labels for each identity are consistent across all images of that identity, even though they are annotated from a single image (Section 4.1.1).
- domain assumption For local representations, the person occupies a contiguous vertical region and the feature magnitude is higher on the body than on the background (Section 3.1.1, alignment module).
- ad hoc to paper The difference between two SID prototypes is a linear combination of per-attribute residual vectors v_l that is shared across attribute differences (Eq. 8, Section 3.2.2).
- ad hoc to paper The boundary margin formula m = log(alpha * N_g/N + beta) correctly balances semantic grouping and identity discrimination (Eq. 4).
invented entities (1)
-
Semantic IDs (SIDs) and their prototypes
independent evidence
Cite this review
Pith. "Pith review of Cerberus: Attribute-based person re-identification using semantic IDs." pith.science (2026). https://pith.science/paper/TLNT3DP6
@misc{pith2026241201048,
author = {Pith},
title = {Pith review of: Cerberus: Attribute-based person re-identification using semantic IDs},
year = {2026},
howpublished = {\url{https://pith.science/paper/TLNT3DP6}},
note = {Machine review of arXiv:2412.01048}
}
read the original abstract
We introduce a new framework, dubbed Cerberus, for attribute-based person re-identification (reID). Our approach leverages person attribute labels to learn local and global person representations that encode specific traits, such as gender and clothing style. To achieve this, we define semantic IDs (SIDs) by combining attribute labels, and use a semantic guidance loss to align the person representations with the prototypical features of corresponding SIDs, encouraging the representations to encode the relevant semantics. Simultaneously, we enforce the representations of the same person to be embedded closely, enabling recognizing subtle differences in appearance to discriminate persons sharing the same attribute labels. To increase the generalization ability on unseen data, we also propose a regularization method that takes advantage of the relationships between SID prototypes. Our framework performs individual comparisons of local and global person representations between query and gallery images for attribute-based reID. By exploiting the SID prototypes aligned with the corresponding representations, it can also perform person attribute recognition (PAR) and attribute-based person search (APS) without bells and whistles. Experimental results on standard benchmarks on attribute-based person reID, Market-1501 and DukeMTMC, demonstrate the superiority of our model compared to the state of the art.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
Cao, Y .-T., Wang, J., & Tao, D. (2020). Symbiotic adversarial learning for attribute-based person search. In Proc. Eur. Conf. Comput. Vis.(pp. 230–247)
work page 2020
-
[3]
Chen, D., Li, H., Liu, X., Shen, Y ., Shao, J., Yuan, Z., & Wang, X. (2018). Improving deep visual representation for person re-identification by global and local image-language association. In Proc. Eur. Conf. Comput. Vis. (pp. 54–70)
work page 2018
-
[4]
Chen, X., Fu, C., Zhao, Y ., Zheng, F., Song, J., Ji, R., & Yang, Y . (2020). Salience-guided cascaded suppression network for person re- identification. In Proc. IEEE Conf. Comput. Vis. Pattern Recog. (pp. 3300–3310)
work page 2020
-
[5]
Chen, Y ., Wang, H., Sun, X., Fan, B., Tang, C., & Zeng, H. (2022). Deep attention aware feature learning for person re-identification. Pattern Recognit., 126, 108567
work page 2022
-
[6]
Deng, Y ., Luo, P., Loy, C. C., & Tang, X. (2014). Pedestrian attribute recognition at far distance. In Proc. ACM Int. Conf. on Multimedia (pp. 789–792)
work page 2014
-
[7]
Dong, Q., Gong, S., & Zhu, X. (2019). Person search by text attribute query as zero-shot learning. In Proc. IEEE Int. Conf. Comput. Vis.(pp. 3652–3661)
work page 2019
-
[8]
Du, G., Gong, T., & Zhang, L. (2024). Contrastive completing learning for practical text-image person reid: Robuster and cheaper. Expert Syst. with Appl., (pp. 123399)
work page 2024
Show all 77 references
-
[9]
& Ham, B
Eom, C. & Ham, B. (2019). Learning disentangled representation for robust person re-identification. In Proc. Int. Conf. Neural Inf. Process. Syst.(pp. 5297–5308)
2019
-
[10]
Felzenszwalb, P., McAllester, D., & Ramanan, D. (2008). A discriminatively trained, multiscale, deformable part model. In Proc. IEEE Conf. Comput. Vis. Pattern Recog.(pp. 1–8)
2008
-
[11]
Fu, H., Zhang, K., & Wang, J. (2024). An adaptive self-correction joint training framework for person re-identification with noisy labels. Expert Syst. with Appl., 238, 121771
2024
-
[12]
& Bengio, Y
Glorot, X. & Bengio, Y . (2010). Understanding the difficulty of training deep feedforward neural networks. In Proc. Int. Conf. on Artif. Intell. and Stat. (pp. 249–256)
2010
-
[13]
Guo, J., Yuan, Y ., Huang, L., Zhang, C., Yao, J.-G., & Han, K. (2019). Beyond human parts: Dual part-aligned representations for person re- identification. In Proc. IEEE Int. Conf. Comput. Vis.(pp. 3642–3651)
2019
-
[14]
Han, K., Guo, J., Zhang, C., & Zhu, M. (2018). Attribute-aware attention model for fine-grained representation learning. In Proc. ACM Int. Conf. on Multimedia (pp. 2040–2048)
2018
-
[15]
He, K., Zhang, X., Ren, S., & Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proc. IEEE Int. Conf. Comput. Vis. (pp. 1026–1034)
2015
-
[16]
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proc. IEEE Conf. Comput. Vis. Pattern Recog.(pp. 770–778)
2016
-
[17]
He, L., Liao, X., Liu, W., Liu, X., Cheng, P., & Mei, T. (2020). FastReID: A pytorch toolbox for general instance re-identification. arXiv preprint arXiv:2006.02631
2020 arXiv
-
[18]
Hermans, A., Beyer, L., & Leibe, B. (2017). In defense of the triplet loss for person re-identification. arXiv preprint arXiv:1703.07737
2017 arXiv
-
[19]
& Schmidhuber, J
Hochreiter, S. & Schmidhuber, J. (1997). Long short-term memory. Neural computation, 9(8), 1735–1780
1997
-
[20]
& Szegedy, C
Ioffe, S. & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proc. Int. Conf. Mach. Learn. (pp. 448–456)
2015
-
[21]
Jaderberg, M., Simonyan, K., Zisserman, A., et al. (2015). Spatial transformer networks. In Proc. Int. Conf. Adv. Neural Inf. Process. Syst. (pp. 2017–2025)
2015
-
[22]
Jeong, B., Park, J., & Kwak, S. (2021). Asmr: Learning attribute-based person search with adaptive semantic margin regularizer. In Proc. IEEE Int. Conf. Comput. Vis. (pp. 12016–12025)
2021
-
[23]
Jia, M., Cheng, X., Lu, S., & Zhang, J. (2022). Learning disentangled repre- sentation implicitly via transformer for occluded person re-identification. IEEE Trans. on Multimedia, 25, 1294–1305
2022
-
[24]
Jiang, B., Wang, X., & Tang, J. (2019). Attkgcn: Attribute knowledge graph convolutional network for person re-identification. arXiv preprint arXiv:1911.10544
2019 arXiv
-
[25]
Kingma, D. P. & Ba, J. (2015). Adam: A method for stochastic optimization. In Proc. Int. Conf. Learn. Representations
2015
-
[26]
Kipf, T. N. & Welling, M. (2017). Semi-supervised classification with graph convolutional networks. In Proc. Int. Conf. Learn. Representations
2017
-
[27]
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. In Proc. Int. Conf. Neural Inf. Process. Syst. (pp. 1097–1105)
2012
-
[28]
Li, D., Chen, X., & Huang, K. (2015). Multi-attribute learning for pedestrian attribute recognition in surveillance scenarios. In Proc. Asian Conf. on Pattern Recognit. (pp. 111–115)
2015
-
[29]
Li, D., Chen, X., Zhang, Z., & Huang, K. (2017). Learning deep context- aware features over body and latent parts for person re-identification. In Proc. IEEE Conf. Comput. Vis. Pattern Recog.(pp. 384–393)
2017
-
[30]
Li, D., Zhang, Z., Chen, X., Ling, H., & Huang, K. (2016). A richly annotated dataset for pedestrian attribute recognition. arXiv preprint arXiv:1603.07054
2016 arXiv
-
[31]
Li, Q., Zhao, X., He, R., & Huang, K. (2019). Pedestrian attribute recognition by joint visual-semantic reasoning and knowledge distillation. In Proc. Int. Joint Conf. on Artificial Intelligence (pp. 833–839)
2019
-
[32]
Li, S., Sun, L., & Li, Q. (2023). Clip-reid: exploiting vision-language model for image re-identification without concrete text labels. In Proc. AAAI. Conf. Artif. Intell. (pp. 1405–1413)
2023
-
[33]
Li, S., Yu, H., & Hu, R. (2020). Attributes-aided part detection and refinement for person re-identification. Pattern Recognit., 97, 107016
2020
-
[34]
Li, W., Zhu, X., & Gong, S. (2018). Harmonious attention network for person re-identification. In Proc. IEEE Conf. Comput. Vis. Pattern Recog. (pp. 2285–2294)
2018
-
[35]
Li, Y ., He, J., Zhang, T., Liu, X., Zhang, Y ., & Wu, F. (2021). Diverse part discovery: Occluded person re-identification with part-aware transformer. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(pp. 2898–2907)
2021
-
[36]
Liang, X., Gong, K., Shen, X., & Lin, L. (2018). Look into person: Joint body parsing & pose estimation network and a new benchmark. IEEE Trans. Pattern Anal. Mach. Intell., 41(4), 871–885. Eom et al.: Preprint submitted to Elsevier Page 15 of 17
2018
-
[37]
Lin, T.-Y ., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(pp. 2117–2125)
2017
-
[38]
Lin, Y ., Zheng, L., Zheng, Z., Wu, Y ., Hu, Z., Yan, C., & Yang, Y . (2019). Improving person re-identification by attribute and identity learning. Pattern Recognit., 95, 151–161
2019
-
[39]
Liu, H., Wu, J., Jiang, J., Qi, M., & Ren, B. (2018a). Sequence-based person attribute recognition with joint ctc-attention model.arXiv preprint arXiv:1811.08115
2018 arXiv
-
[40]
Liu, X., Zhao, H., Tian, M., Sheng, L., Shao, J., Yi, S., Yan, J., & Wang, X. (2017). HydraPlus-Net: Attentive deep features for pedestrian analysis. In Proc. IEEE Int. Conf. Comput. Vis.(pp. 350–359)
2017
-
[41]
& Hutter, F
Loshchilov, I. & Hutter, F. (2016). Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983
2016 arXiv
-
[42]
Luo, H., Gu, Y ., Liao, X., Lai, S., & Jiang, W. (2019). Bag of tricks and a strong baseline for deep person re-identification. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshop(pp. 0–0)
2019
-
[43]
Maaten, L. v. d. & Hinton, G. (2008). Visualizing data using t-sne. J. Mach. Learn. Res., 9(Nov), 2579–2605
2008
-
[44]
X., Nguyen, B
Nguyen, B. X., Nguyen, B. D., Do, T., Tjiputra, E., Tran, Q. D., & Nguyen, A. (2021). Graph-based person signature for person re-identifications. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(pp. 3492–3501)
2021
-
[45]
Ni, X., Fang, L., & Huttunen, H. (2021). Adaptive l2 regularization in person re-identification. In Int. Conf. Pattern Recognit. (pp. 9601–9607)
2021
-
[46]
& Pedrini, H
Quispe, R. & Pedrini, H. (2021). Top-db-net: Top dropblock for activation enhancement in person re-identification. In Int. Conf. Pattern Recognit. (pp. 2980–2987)
2021
-
[47]
Ren, M., He, L., Liao, X., Liu, W., Wang, Y ., & Tan, T. (2021). Learning instance-level spatial-temporal patterns for person re-identification. In Proc. IEEE Int. Conf. Comput. Vis.(pp. 14930–14939)
2021
-
[48]
Ren, S., He, K., Girshick, R., & Sun, J. (2017). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell., 39(6), 1137–1149
2017
-
[49]
& Stiefelhagen, R
Schumann, A. & Stiefelhagen, R. (2017). Person re-identification by deep learning attribute-complementary information. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshops(pp. 20–28)
2017
-
[50]
Somers, V ., De Vleeschouwer, C., & Alahi, A. (2023). Body part-based representation learning for occluded person re-identification. In Proc. IEEE Winter Conf. Comput. Vis.(pp. 1613–1623)
2023
-
[51]
Specker, A., Cormier, M., & Beyerer, J. (2023). UPAR: Unified pedestrian attribute recognition and person retrieval. In Proc. IEEE Winter Conf. Comput. Vis. (pp. 981–990)
2023
-
[52]
Su, C., Li, J., Zhang, S., Xing, J., Gao, W., & Tian, Q. (2017). Pose-driven deep convolutional model for person re-identification. In Proc. IEEE Int. Conf. Comput. Vis. (pp. 3960–3969)
2017
-
[53]
Sudowe, P., Spitzer, H., & Leibe, B. (2015). Person attribute recognition with a jointly-trained holistic cnn model. In Proc. IEEE Int. Conf. Comput. Vis. Workshop(pp. 87–95)
2015
-
[54]
Suh, Y ., Wang, J., Tang, S., Mei, T., & Mu Lee, K. (2018). Part-aligned bilinear representations for person re-identification. In Proc. Eur. Conf. Comput. Vis. (pp. 402–419)
2018
-
[55]
Tan, Z., Yang, Y ., Wan, J., Guo, G., & Li, S. Z. (2020). Relation-aware pedestrian attribute recognition with graph convolutional networks. In Proc. AAAI. Conf. Artif. Intell., volume 34 (pp. 12055–12062)
2020
-
[56]
Tang, C., Sheng, L., Zhang, Z., & Hu, X. (2019). Improving pedestrian at- tribute recognition with weakly-supervised multi-scale attribute-specific localization. In Proc. IEEE Int. Conf. Comput. Vis.(pp. 4997–5006)
2019
-
[57]
Tay, C.-P., Roy, S., & Yap, K.-H. (2019). AANet: Attribute attention network for person re-identifications. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (pp. 7134–7143)
2019
-
[58]
Wang, G., Lai, J., Huang, P., & Xie, X. (2019). Spatial-temporal person re-identification. In Proc. AAAI. Conf. Artif. Intell. , volume 33 (pp. 8933–8940)
2019
-
[59]
Wang, J., Zhu, X., Gong, S., & Li, W. (2017). Attribute recognition by joint recurrent learning of context and correlation. In Proc. IEEE Int. Conf. Comput. Vis. (pp. 531–540)
2017
-
[60]
Wang, P., Zhao, Z., Su, F., & Meng, H. (2022). LTReID: Factorizable feature generation with independent components for long-tailed person re-identification. IEEE Trans. on Multimedia, 25, 4610–4622
2022
-
[61]
Wang, Z., Fang, Z., Wang, J., & Yang, Y . (2020). Vitaa: Visual-textual attributes alignment in person search by natural language. In Proc. Eur. Conf. Comput. Vis. (pp. 402–420)
2020
-
[62]
& Kipf, T
Welling, M. & Kipf, T. N. (2017). Semi-supervised classification with graph convolutional networks. In Proc. Int. Conf. Learn. Representations
2017
-
[63]
Yan, Y ., Yu, H., Li, S., Lu, Z., He, J., Zhang, H., & Wang, R. (2022). Weaken- ing the influence of clothing: universal clothing attribute disentanglement for person re-identification. In Proc. Int. Joint Conf. Artif. Intell. (pp. 1523–1529)
2022
-
[64]
Yang, J., Fan, J., Wang, Y ., Wang, Y ., Gan, W., Liu, L., & Wu, W. (2020). Hierarchical feature embedding for attribute recognition. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(pp. 13055–13064)
2020
-
[65]
Lai, J. (2018). Adversarial attribute-image person re-identification. In Proc. Int. Joint Conf. on Artificial Intelligence (pp. 1100–1106)
2018
-
[66]
Zhang, Z., Lan, C., Zeng, W., Jin, X., & Chen, Z. (2020). Relation-aware global attention for person re-identification. In Proc. IEEE Conf. Comput. Vis. Pattern Recog.(pp. 3186–3195)
2020
-
[67]
Zhao, H., Tian, M., Sun, S., Shao, J., Yan, J., Yi, S., Wang, X., & Tang, X. (2017). Spindle Net: Person re-identification with human body region guided feature decomposition and fusion. In Proc. IEEE Conf. Comput. Vis. Pattern Recog.(pp. 1077–1085)
2017
-
[68]
Zhao, S., Gao, C., Shao, Y ., Zheng, W.-S., & Sang, N. (2021). Weakly supervised text-based person re-identification. In Proc. IEEE Int. Conf. Comput. Vis. (pp. 11395–11404)
2021
-
[69]
Zhao, S., Gao, C., Zhang, J., Cheng, H., Han, C., Jiang, X., Guo, X., Zheng, W.-S., Sang, N., & Sun, X. (2020). Do Not Disturb Me: Person re- identification under the interference of other pedestrians. In Proc. Eur. Conf. Comput. Vis. (pp. 647–663)
2020
-
[70]
Zhao, X., Sang, L., Ding, G., Guo, Y ., & Jin, X. (2018). Grouping attribute recognition for pedestrian with joint recurrent learning. In Proc. Int. Joint Conf. on Artificial Intelligence (pp. 3177–3183)
2018
-
[71]
Zheng, L., Shen, L., Tian, L., Wang, S., Wang, J., & Tian, Q. (2015). Scalable person re-identification: A benchmark. In Proc. IEEE Int. Conf. Comput. Vis. (pp. 1116–1124)
2015
-
[72]
Zheng, Z., Zheng, L., & Yang, Y . (2017). Unlabeled samples generated by GAN improve the person re-identification baseline in vitro. In Proc. IEEE Int. Conf. Comput. Vis. (pp. 3754–3762)
2017
-
[73]
Zhong, Z., Zheng, L., Cao, D., & Li, S. (2017). Re-ranking person re- identification with k-reciprocal encoding. In Proc. IEEE Conf. Comput. Vis. Pattern Recog.(pp. 1318–1327)
2017
-
[74]
Zhong, Z., Zheng, L., Kang, G., Li, S., & Yang, Y . (2020). Random erasing data augmentation. In Proc. AAAI. Conf. Artif. Intell. (pp. 13001–13008). Eom et al.: Preprint submitted to Elsevier Page 16 of 17
2020
-
[75]
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2016). Learning deep features for discriminative localization. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(pp. 2921–2929)
2016
-
[76]
Zhu, H., Ke, W., Li, D., Liu, J., Tian, L., & Shan, Y . (2022). Dual cross- attention learning for fine-grained visual categorization and object re- identification. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit.(pp. 4692–4702)
2022
-
[77]
Zhu, K., Guo, H., Liu, Z., Tang, M., & Wang, J. (2020). Identity-guided human semantic parsing for person re-identification. In Proc. Eur. Conf. Comput. Vis. (pp. 346–363). Eom et al.: Preprint submitted to Elsevier Page 17 of 17
2020
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.