REVIEW 5 major objections 6 minor 54 references
Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation
T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper argues that multimodal entity linking fails for two reasons: random negative sampling lets the model lean on easy attributes, and the visual content of a mention often does not match the knowledge base image.
desk verdict A plausible, well-intended MEL paper whose central hard-negative mechanism rests on entity attributes it never describes; worth refereeing, but the reported numbers are provisional until code and attribute details are released. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is JD-CCL plus CVaCPT. JD-CCL computes a Jaccard distance $J(i,j) = |e a_i \cap e a_j| / |e a_i \cup e a_j|$ between the attribute sets of entities, sorts them, and draws the top-$k$ most similar entities as negative samples in a conditional contrastive loss. CVaCPT feeds the mention sentence into a text-to-image diffusion model to produce several synthetic images, pools their global features via max-pooling, mixes the pooled feature with the mention's text representation in a contextual encoder, and predicts per-patch affine parameters $\alpha$ and $\beta$ that scale and shift the mention image's patch features before matching. Together they force the model to compare mentions against near-duplicate entities and to focus image encoding on patches that are relevant to the queried entity.
What would settle it
Train the same model on WikiDiverse but replace every entity's attribute set with a random token list before computing Jaccard distances; if H@1 stays near 70.25, the hard-negative selection is not what drives the gain, and if it falls back toward the MIMIC baseline of 63.51, the attributes are load-bearing.
Extended reading notes
Core claim
The central claim is that replacing random negatives with Jaccard-distance-selected hard negatives, and conditioning visual features on multi-view synthetic images, makes a multimodal entity linking model learn more discriminative mention-entity matches. The authors build on the MIMIC architecture, swap its contrastive objective for JD-CCL, and wrap the mention visual encoder with CVaCPT. They report H@1 of 70.25 on WikiDiverse versus 63.51 for MIMIC, 83.38 versus 81.02 on RichpediaMEL, and 89.28 versus 87.98 on WikiMEL, with the improvements shown to be significant by paired t-tests. Their conclusion is that the combination of attribute-conditioned hard negatives and controllable visual augmentation is the reason for the gains.
Load-bearing premise
The whole approach assumes that every entity in the knowledge base comes with attribute sets that are complete enough and discriminative enough that Jaccard distance ranks the truly hardest negatives; the paper's own limitations section concedes that poorly annotated datasets would hurt training.
Editorial extensions
If this is right
- Any MEL system that uses contrastive training can swap in JD-CCL by precomputing Jaccard similarities over entity attributes, with no change to the matcher.
- CVaCPT can be attached to existing visual encoders as a trainable preprocessing head, meaning the improvement is architecture-agnostic.
- The reported gains are largest when the knowledge base contains many same-name or near-duplicate entities, as in WikiDiverse.
- Merging small datasets into larger knowledge bases (WikiDM, WikiDR) still shows the method ahead of MIMIC, suggesting it scales to bigger candidate sets.
Reading between the lines
- The method's success hinges on attribute availability: if attributes are missing or noisy, the top-k negatives by Jaccard distance may not be genuinely similar, and the H@1 gain would shrink.
- A natural testable extension is to apply JD-CCL to other retrieval problems that have structured entity attributes, such as product or legal entity search.
- The paper leaves implicit that the diffusion-generated images are a form of data augmentation; using retrieved or web-mined images as a cheaper alternative would test whether generation is essential or merely sufficient.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two augmentations to the MIMIC-style multimodal entity linking framework. JD-CCL replaces random negative sampling in the contrastive objective with hard negatives selected by Jaccard similarity over entity attribute sets. CVaCPT generates multiple diffusion-based synthetic images from the mention sentence and uses their pooled global features, together with contextual text features, to compute affine scale-and-shift parameters that modulate the mention image patch representations. Experiments on WikiDiverse, RichpediaMEL, WikiMEL, and two merged larger knowledge bases report consistent improvements over the MIMIC baseline, most notably a 6.74% H@1 gain on WikiDiverse. The paper also includes ablations showing that removing either module hurts performance and that varying k (hard-negative count) and n_s (synthetic image count) changes results in the expected direction.
Significance. If the reported results are reproducible, the work makes a useful practical contribution: it demonstrates that knowledge-base attribute information can guide harder contrastive negatives for multimodal entity linking, and that diffusion-generated multi-view images can help control noisy visual patches. The paper explicitly reports ablations for both components and acknowledges computational and annotation-quality limitations. However, the absence of code, the unverified source and coverage of the attribute sets on which JD-CCL rests, and the failure to fix the main-run hyperparameters prevent the central claim from being fully assessed. The contribution is plausible and the direction is worth pursuing, but the manuscript currently does not provide enough evidence to certify the mechanism behind the gains.
major comments (5)
- [§4.2, Eq. (3)] The Jaccard formula contains a typo in the denominator: it reads |eai ∪ ebi|, but ebi is not defined anywhere, and the intended expression is |eai ∪ eaj|. In addition, Eq. (3) is called a Jaccard distance but it is actually the Jaccard similarity index; the text then says 'top k scores as the k most similar or hard negative samples', which is consistent with similarity but contradicts the earlier 'Jaccard distance' wording. Since this preprocessing step is the entire basis of JD-CCL, the manuscript must correct the formula and use consistent terminology.
- [§3, §4.2, §8 (Limitations)] The paper never states where the entity attribute sets eai come from, how they are obtained, or how they are cleaned. JD-CCL's hard-negative selection relies entirely on the assumption that the top-k entities by Jaccard similarity are genuinely hard negatives and not arbitrary due to sparse, noisy, or coarse attribute sets. Section 8 concedes that 'If the dataset is not well-annotated, this could negatively impact the training stage', but no statistics are provided (e.g., distribution of attribute-set sizes, number of zero-similarity pairs, fraction of ties in the top-k lists). Without such diagnostics, the 6.74% H@1 gain on WikiDiverse cannot be unambiguously attributed to the proposed Jaccard-based negative sampling rather than to the general use of attribute text in the entity encoder.
- [§5.4, §6, Table 1 and Table 2] The main-run values of k (number of JD-CCL negatives) and n_s (number of synthetic images in CVaCPT) are never fixed in the paper. Table 2 reports sensitivity experiments over k ∈ {2,4,6} and n_s ∈ {1,2,3}, but the main results in Table 1 do not state which of these values were used, and the coefficient w in Eqs. (5)–(6) is also left unspecified. This makes the main results irreproducible as written and also prevents the reader from knowing whether the reported gains correspond to a carefully tuned configuration.
- [§4.1, §4.2, Table 2 (w/o LD-CCL)] The ablation 'w/o LD-CCL' does not isolate the effect of Jaccard-based hard negative sampling. Section 4.1 states that entity inputs are formed by concatenating the entity name with its attributes eai, so removing JD-CCL likely also removes the attribute text from the entity encoder. As a result, the ablation confounds two changes: the loss-relevant negative sampling strategy and the input representation of entities. To support the claim that 'hard negative samples play an important role' (§7.1), the authors should compare against a variant that keeps attribute text in the entity input but uses random negative sampling, in addition to the variant that removes JD-CCL entirely.
- [§5.4, §6, Table 1] The paper reports that models were run five times and that paired t-tests with p < 0.05 show significant improvements, but it does not report standard deviations, confidence intervals, or the number of paired evaluations. This is especially important because the gains on RichpediaMEL and WikiMEL are small (2.36% and 1.30% H@1), while the paper claims significance. Without variance information, the reader cannot judge whether the differences on the smaller-gain datasets are meaningful.
minor comments (6)
- [Table 2] The table header uses 'LD-CCL' where the text uses 'JD-CCL'; this typo should be corrected.
- [§4.3, Eqs. (5)–(8)] The notation is inconsistent: the text introduces PθX for the transformation, but the equations use AG and AL, and AG and AL are never formally defined. Please define all components of the patch transform explicitly.
- [§5.1, §7.2] The symbol for the number of synthetic images alternates between s and n_s; Section 7.2 writes 'the number of synthetic images, s' while Section 4.3 uses ns. Please use one notation consistently.
- [§6] The dataset name is misspelled as 'RipediaMEL' in the sentence reporting the 2.36% improvement; the correct spelling is RichpediaMEL.
- [§7.4] The caption of Figure 6 says 'In this appendix', but the figure appears in the main text, not in an appendix. This should be corrected.
- [§5.4] The sentence 'For the mention or entity input that do not have image, will initialize with blank image (while) with zero padding' contains a grammatical error and an unclear parenthetical '(while)'. Please clarify the placeholder-image initialization procedure.
Circularity Check
No circularity: the reported benchmark gains come from external KB attributes and diffusion-generated images, and the evaluation is empirical against independent baselines.
full rationale
The derivation chain is self-contained and the reported improvements are empirical benchmark results, not consequences of a definitional identity. JD-CCL selects hard negatives by Jaccard similarity over entity attribute sets in a preprocessing step (Eq. 3), then trains the contrastive objective (Eq. 4) on mention-entity pairs from the benchmark datasets. The evaluation metrics (H@1, H@3, H@5, MRR) measure agreement with held-out ground-truth entity links; nothing in the Jaccard preprocessing or the contrastive loss encodes those test labels. Similarly, CVaCPT uses Stable Diffusion synthetic images and CLIP features to modulate patch representations, again without using task labels to construct the module. The comparison against external baselines such as CLIP, MIMIC, and ALBEF is therefore independent of the paper's own method. Self-citations appear only in the related-work enumeration and in lists of the authors' prior contrastive-learning papers; none of these citations is load-bearing for the central claim, and no uniqueness theorem or ansatz is imported from the authors' prior work. The ablation 'w/o JD-CCL' may not fully separate the effect of attribute text in the entity encoder from the effect of Jaccard-based negative sampling, and Eq. (3) contains an undefined symbol 'ebi', and the paper does not describe how attributes are obtained or cleaned. These are reproducibility and robustness limitations, which the paper itself acknowledges in Section 8 ('If the dataset is not well-annotated, this could negatively impact the training stage'), but they are not circularity. No step in the paper reduces a prediction to its own input by construction.
Assumptions & free parameters
free parameters (3)
- w in Eqs. 5-6 (CPT mixing coefficient) =
not reported
- k (number of JD-CCL hard negatives) =
not stated for main results
- n_s (number of synthetic images in CVaCPT) =
not stated for main results
assumptions (5)
- domain assumption Entity attribute sets eai are available, complete, and discriminative for every entity used in training and evaluation.
- ad hoc to paper Entities with high Jaccard attribute overlap are the most useful hard negatives for contrastive learning.
- domain assumption Stable Diffusion synthetic images generated from the mention sentence are semantically aligned with the mention and improve visual matching.
- domain assumption Pre-trained CLIP features and the MIMIC matcher transfer well when fine-tuned with the new objectives.
- domain assumption Zero-padded blank images are acceptable placeholders for missing images.
Cite this review
Pith. "Pith review of Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation." pith.science (2026). https://pith.science/paper/L3ZHIK2N
@misc{pith2026250114166,
author = {Pith},
title = {Pith review of: Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/L3ZHIK2N}},
note = {Machine review of arXiv:2501.14166}
}
read the original abstract
Previous research on multimodal entity linking (MEL) has primarily employed contrastive learning as the primary objective. However, using the rest of the batch as negative samples without careful consideration, these studies risk leveraging easy features and potentially overlook essential details that make entities unique. In this work, we propose JD-CCL (Jaccard Distance-based Conditional Contrastive Learning), a novel approach designed to enhance the ability to match multimodal entity linking models. JD-CCL leverages meta-information to select negative samples with similar attributes, making the linking task more challenging and robust. Additionally, to address the limitations caused by the variations within the visual modality among mentions and entities, we introduce a novel method, CVaCPT (Contextual Visual-aid Controllable Patch Transform). It enhances visual representations by incorporating multi-view synthetic images and contextual textual representations to scale and shift patch representations. Experimental results on benchmark MEL datasets demonstrate the strong effectiveness of our approach.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Vassileios Balntas, Edgar Riba, Daniel Ponsa, and Krystian Mikolajczyk. 2016. https://api.semanticscholar.org/CorpusID:27938870 Learning local feature descriptors with triplets and shallow convolutional neural networks . In British Machine Vision Conference
work page 2016
-
[2]
S. Chopra, R. Hadsell, and Y. LeCun. 2005. https://doi.org/10.1109/CVPR.2005.202 Learning a similarity metric discriminatively, with application to face verification . In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05), volume 1, pages 539--546 vol. 1
-
[3]
Jiankang Deng, Jia Guo, and Stefanos Zafeiriou. 2018. https://arxiv.org/abs/1801.07698 Arcface: Additive angular margin loss for deep face recognition . CoRR, abs/1801.07698
arXiv 2018
-
[4]
Jacob Devlin, Ming - Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1) , pages 4171--4186. Association for Computational Linguistics
2019
-
[5]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR . OpenReview.net
work page 2021
-
[6]
Zi - Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, Zicheng Liu, and Michael Zeng. 2022. An empirical study of training end-to-end vision-and-language transformers. In CVPR , pages 18145--18155. IEEE
work page 2022
-
[7]
Jingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang, Wei He, and Qingming Huang. 2021. Multimodal entity linking: A new dataset and A baseline. In ACM Multimedia , pages 993--1001. ACM
work page 2021
-
[8]
Gerritse, Faegheh Hasibi, and Arjen P
Emma J. Gerritse, Faegheh Hasibi, and Arjen P. de Vries. 2022. Entity-aware transformers for entity search. In SIGIR , pages 1455--1465. ACM
work page 2022
Show all 54 references
-
[9]
Minguk Kang and Jaesik Park. 2020. Contragan: Contrastive learning for conditional image generation. Advances in Neural Information Processing Systems, 33:21357--21369
2020
-
[10]
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. arXiv preprint arXiv:2004.11362
2020 arXiv
-
[11]
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In ICML , volume 139 of Proceedings of Machine Learning Research, pages 5583--5594. PMLR
2021
-
[12]
Selvaraju, Akhilesh Gotmare, Shafiq R
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu - Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In NeurIPS, pages 9694--9705
2021
-
[13]
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song. 2017. https://doi.org/10.1109/CVPR.2017.713 Sphereface: Deep hypersphere embedding for face recognition . In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6738--6746
2017 doi
-
[14]
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In EMNLP (1) , pages 7052--7063. Association for Computational Linguistics
2021
-
[15]
Pengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu, Linli Xu, and Enhong Chen. 2023. https://arxiv.org/abs/2307.09721 Multi-grained multimodal interaction network for entity linking . Preprint, arXiv:2307.09721
2023 arXiv
-
[16]
Ma, Yao-Hung Hubert Tsai, Paul Pu Liang, Han Zhao, Kun Zhang, Ruslan Salakhutdinov, and Louis-Philippe Morency
Martin Q. Ma, Yao-Hung Hubert Tsai, Paul Pu Liang, Han Zhao, Kun Zhang, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022. https://arxiv.org/abs/2106.02866 Conditional contrastive learning for improving fairness in self-supervised learning . Preprint, arXiv:2106.02866
2022 arXiv
-
[17]
Seungwhan Moon, Leonardo Neves, and Vitor Carvalho. 2018. Multimodal named entity disambiguation for noisy social media posts. In ACL (1) , pages 2000--2008. Association for Computational Linguistics
2018
-
[18]
Cong-Duy Nguyen, Thong Nguyen, Duc Anh Vu, and Luu Anh Tuan. 2023 a . https://arxiv.org/abs/2312.02227 Improving multimodal sentiment analysis: Supervised angular margin-based contrastive learning for enhanced fusion representation . Preprint, arXiv:2312.02227
2023 arXiv
-
[19]
Thong Nguyen, Yi Bin, Xiaobao Wu, Xinshuai Dong, Zhiyuan Hu, Khoi Le, Cong-Duy Nguyen, See-Kiong Ng, and Luu Anh Tuan. 2025. Meta-optimized angular margin contrastive framework for video-language representation learning. In European Conference on Computer Vision, pages 77--98....
2025
-
[20]
Thong Nguyen and Anh Tuan Luu. 2021. https://arxiv.org/pdf/2110.12764 Contrastive learning for neural topic model . Advances in Neural Information Processing Systems, 34
2021 arXiv
-
[21]
Thong Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy Nguyen, See-Kiong Ng, and Luu Anh Tuan. 2023 b . Demaformer: Damped exponential moving average transformer with energy-based modeling for temporal language grounding. arXiv preprint arXiv:2312.02549
2023 arXiv
-
[22]
Thong Nguyen, Xiaobao Wu, Anh Tuan Luu, Zhen Hai, and Lidong Bing. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.686 Adaptive contrastive learning on multimodal transformer for review helpfulness prediction . In Proceedings of the 2022 Conference on Empirical Methods in Na...
2022 doi
-
[23]
Thong Thanh Nguyen, Yi Bin, Xiaobao Wu, Zhiyuan Hu, Cong-Duy T Nguyen, See-Kiong Ng, and Anh Tuan Luu. 2024 a . Multi-scale contrastive learning for video temporal grounding. arXiv preprint arXiv:2412.07157
2024 arXiv
-
[24]
Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T Nguyen, See-Kiong Ng, and Anh Tuan Luu. 2024 b . Motion-aware contrastive learning for temporal panoptic scene graph generation. arXiv preprint arXiv:2412.07160
2024 arXiv
-
[25]
Thong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy T Nguyen, See-Kiong Ng, and Anh Tuan Luu. 2024 c . https://openreview.net/forum?id=HdAoLSBYXj Topic modeling as multi-objective optimization with setwise contrastive learning . In The Twelfth International Conference on L...
2024
-
[26]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In ICML , volu...
2021
-
[27]
Senbao Shi, Zhenran Xu, Baotian Hu, and Min Zhang. 2024. https://arxiv.org/abs/2306.12725 Generative multimodal entity linking . Preprint, arXiv:2306.12725
2024 arXiv
-
[28]
Kihyuk Sohn. 2016. https://proceedings.neurips.cc/paper_files/paper/2016/file/6b180037abbebea991d8b1232f8a8ca9-Paper.pdf Improved deep metric learning with multi-class n-pair loss objective . In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc
2016
-
[29]
Shezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan, Shasha Li, Xiaoguang Mao, and Meng Wang. 2023. https://arxiv.org/abs/2312.11816 A dual-way enhanced framework from text matching point of view for multimodal entity linking . Preprint, arXiv:2312.11816
2023 arXiv
-
[30]
Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. 2020. What makes for good views for contrastive learning? arXiv preprint arXiv:2005.10243
2020 arXiv
-
[31]
Yao-Hung Hubert Tsai, Tianqin Li, Weixin Liu, Peiyuan Liao, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2021. Integrating auxiliary information in self-supervised learning. arXiv preprint arXiv:2106.02869
2021 arXiv
-
[32]
Yao-Hung Hubert Tsai, Tianqin Li, Martin Q Ma, Han Zhao, Kun Zhang, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2022. Conditional contrastive learning with kernel. arXiv preprint arXiv:2202.05458
2022 arXiv
-
[33]
Denny Vrandecic and Markus Kr \" o tzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM , 57(10):78--85
2014
-
[34]
Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Zhifeng Li, Dihong Gong, Jingchao Zhou, and Wei Liu. 2018. https://arxiv.org/abs/1801.09414 Cosface: Large margin cosine loss for deep face recognition . CoRR, abs/1801.09414
2018 arXiv
-
[35]
Meng Wang, Haofen Wang, Guilin Qi, and Qiushuo Zheng. 2020. Richpedia: A large-scale, comprehensive multi-modal knowledge graph. Big Data Res., 22:100159
2020
-
[36]
Peng Wang, Jiangheng Wu, and Xiaohang Chen. 2022 a . https://doi.org/10.1145/3477495.3531867 Multimodal entity linking with gated hierarchical fusion and contrastive training . In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Informa...
2022
-
[37]
Xuwu Wang, Junfeng Tian, Min Gui, Zhixu Li, Rui Wang, Ming Yan, Lihan Chen, and Yanghua Xiao. 2022 b . Wikidiverse: A multimodal entity linking dataset with diversified contextual topics and entity types. In ACL (1) , pages 4785--4797. Association for Computational Linguistics
2022
-
[38]
Jie Wei, Guanyu Hu, Xinyu Yang, Anh Tuan Luu, and Yizhuo Dong. 2024. Learning facial expression and body gesture visual information for video emotion recognition. Expert Systems with Applications, 237:121419
2024
-
[39]
Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao. 2016. A discriminative feature learning approach for deep face recognition. In Computer Vision -- ECCV 2016, pages 499--515, Cham. Springer International Publishing
2016
-
[40]
Mike Wu, Milan Mosse, Chengxu Zhuang, Daniel Yamins, and Noah Goodman. 2020 a . Conditional negative sampling for contrastive learning of visual representations. arXiv preprint arXiv:2010.02037
2020 arXiv
-
[41]
Xiaobao Wu, Xinshuai Dong, Thong Nguyen, Chaoqun Liu, Liang-Ming Pan, and Anh Tuan Luu. 2023 a . https://arxiv.org/abs/2304.03544 Infoctm: A mutual information maximization perspective of cross-lingual topic modeling . In Proceedings of the AAAI Conference on Artificial Intell...
2023 arXiv
-
[42]
Xiaobao Wu, Xinshuai Dong, Thong Nguyen, and Anh Tuan Luu. 2023 b . https://arxiv.org/pdf/2306.04217 Effective neural topic modeling with embedding clustering regularization . In International Conference on Machine Learning. PMLR
2023 arXiv
-
[43]
Xiaobao Wu, Xinshuai Dong, Liangming Pan, Thong Nguyen, and Anh Tuan Luu. 2024 a . https://arxiv.org/pdf/2405.17957 Modeling dynamic topics in chain-free fashion by evolution-tracking contrastive learning and unassociated word exclusion . In Findings of the Association for Com...
2024 arXiv
-
[44]
Xiaobao Wu, Chunping Li, Yan Zhu, and Yishu Miao. 2020 b . https://aclanthology.org/2020.emnlp-main.138.pdf Short text topic modeling with topic distribution quantization and negative sampling decoder . In Proceedings of the 2020 Conference on Empirical Methods in Natural Lang...
2020
-
[45]
Xiaobao Wu, Anh Tuan Luu, and Xinshuai Dong. 2022. https://aclanthology.org/2022.emnlp-main.176 Mitigating data sparsity for short text topic modeling by topic-semantic contrastive learning . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Proces...
2022
-
[46]
Xiaobao Wu, Fengjun Pan, and Anh Tuan Luu. 2024 b . https://aclanthology.org/2024.acl-demos.4 Towards the T op M ost: A topic modeling system toolkit . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations),...
2024
-
[47]
Xiaobao Wu, Fengjun Pan, Thong Nguyen, Yichao Feng, Chaoqun Liu, Cong-Duy Nguyen, and Anh Tuan Luu. 2024 c . https://arxiv.org/pdf/2401.14113.pdf On the affinity, rationality, and diversity of hierarchical topic modeling . In Proceedings of the AAAI Conference on Artificial In...
2024 arXiv
-
[48]
Yadollah Yaghoobzadeh, Heike Adel, and Hinrich Sch \"u tze. 2017. https://aclanthology.org/E17-1111 Noise mitigation for neural entity typing and relation extraction . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguisti...
2017
-
[49]
Chengmei Yang, Bowei He, Yimeng Wu, Chao Xing, Lianghua He, and Chen Ma. 2023. https://proceedings.mlr.press/v216/yang23d.html MMEL : A joint learning framework for multi-mention entity linking . In Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intell...
2023
-
[50]
Li Zhang, Zhixu Li, and Qiang Yang. 2021. Attention-based multimodal entity linking with high-quality images. In DASFAA (2) , volume 12682 of Lecture Notes in Computer Science, pages 533--548. Springer
2021
-
[51]
Yuhao Zhang, Hongji Zhu, Yongliang Wang, Nan Xu, Xiaobo Li, and Binqiang Zhao. 2022. https://doi.org/10.18653/v1/2022.acl-long.336 A contrastive framework for learning sentence representations from pairwise and triple-wise perspective in angular space . In Proceedings of the 6...
2022 doi
-
[52]
Qiushuo Zheng, Hao Wen, Meng Wang, and Guilin Qi. 2022. Visual entity linking via multi-modal learning. Data Intell., 4(1):1--19
2022
-
[53]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[54]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.