Pith. sign in

REVIEW 4 major objections 5 minor 67 references

AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper shows a retrieval-augmented LLM can rewrite abnormal Chinese addresses, correcting 43.9% offline and cutting live parcel re-routing by over 40%.

desk verdict AddrLLM is a real industrial LLM+RAG system with a genuine deployment, but the headline 43% re-routing claim is not supported by the evidence in the paper. read the letter →

arxiv 2411.13584 v1 pith:H5QDTVEN submitted 2024-11-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords addressrewritinglargelanguagemodelsretrieval-augmentedgenerationgeocodinglogisticsreinforcementlearningqueryreformulationabnormaldetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Abnormal addresses—those with missing regions, nested cites, aliases, or misspellings—routinely send parcels to the wrong delivery station and force costly re-routing. The paper argues that a retrieval-augmented large language model, fine-tuned on address parsing, entity prediction, and rewriting tasks, can fix these errors in a unified model without retraining when new addresses appear. The model is aligned by a reward computed from the logistics company's own geocoding service and courier-reported delivery coordinates, which the authors call bias-free feedback. Offline, the system corrects 43.9% of abnormal addresses; in a live deployment covering roughly two million daily parcels, it reduced abnormal-address re-routing by over 40%.

What carries the argument

The central machinery is a retrieval-augmented LLM whose retriever is fine-tuned on geocoding so that relevance means spatial proximity, and whose generator is trained by supervised fine-tuning plus PPO with a reward computed directly from the logistics geocoding system. The reward is the sum of a semantic cosine score, a reverse-geocoding score, and a geocoding-distance score, which lets the model learn to correct addresses that the existing system cannot handle without a trained reward model or manual labels.

What would settle it

Take a random sample of addresses AddrLLM rewrote during deployment, verify the intended destination by direct courier follow-up or independent ground-truth inspection, and compare delivery success for rewritten versus original addresses; if rewritten addresses do not reach the correct station more often than the originals, the claimed 40% re-routing reduction collapses.

Watch

Extended reading notes

Core claim

On its own terms, the paper claims that LLM-based address rewriting works in production. AddrLLM combines multi-instruction supervised fine-tuning on 20 million parsing, 20 million entity-prediction, and 20 million rewriting samples with an address-centric retrieval module whose retriever is fine-tuned on 200 million geocoding pairs so that retrieved addresses are geographically close rather than merely semantically similar. A PPO objective-alignment stage then optimizes a three-term reward: semantic similarity to the original address, similarity to the reverse-geocoded delivery coordinates, and geocoding distance to the courier-reported delivery point. The reported offline results include 91.8% trigger prediction and 90.3% accuracy on entity prediction, 89.7% hit rate on direct rewriting, 94.3% station-level geocoding accuracy, and 43.9% correction of abnormal addresses, with 99.9% robustness on standard addresses; the authors attribute these gains to the synergy of all three modules.

Load-bearing premise

The load-bearing premise is that the logistics system's geocoding, reverse geocoding, and courier-reported delivery coordinates provide an unbiased definition of a correct rewritten address; if that feedback is noisy or systematically biased, the measured correction and re-routing reductions would be overstated.

Editorial extensions

If this is right

  • Offline, AddrLLM corrects 43.9% of abnormal addresses, a correction rate 24.2 percentage points higher than the best baseline.
  • The model keeps 99.9% of standard addresses unchanged and correct, so applying it to normal traffic poses little risk of breaking good addresses.
  • Station-level geocoding accuracy rises to 94.3%, corresponding to roughly a 43% reduction in station-level misrouting compared with the existing geocoding service.
  • Live deployment in Zhejiang corrected over 40% of abnormal addresses per day and remained stable over 60 days including a peak sales event.
  • Because address knowledge is stored in an external retrieval database, adding new addresses to the database does not require retraining the LLM.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same recipe—use a downstream operational outcome as the RL reward instead of a learned reward model—could transfer to other rewriting or normalization tasks where a measurable system outcome exists, such as product-query rewriting or clinical-text normalization.
  • The spatially fine-tuned retriever may be reusable as a geocoding-aware encoder for other geo-NLP tasks, since its embedding distances were shown to correlate linearly with real geographical distances.
  • A natural next experiment, acknowledged as missing by the paper itself, is to report correction rates by error type (missing region, nested address, alias, irrelevant words, misspelling); that breakdown would show where the LLM's gain over older systems actually comes from.
  • If the geocoding reward is noisy near delivery-station boundaries, the station-level gains may overstate real delivery improvements; a courier-confirmed outcome study would settle the practical value.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces AddrLLM, an LLM-based address rewriting framework for logistics, combining multi-instruction supervised fine-tuning (SFT), an address-centric retrieval-augmented generation (RAG) module, and an objective alignment stage trained with rewards derived from JD Logistics' LBS system (geocoding, reverse geocoding, and courier-reported delivery coordinates). Offline experiments on nationwide data show AddrLLM outperforming existing methods on address entity prediction, direct rewriting hit rate, and geocoding accuracy (e.g., Acc@Station 94.3 vs. 90.0 for SoP). The model was deployed in Zhejiang province for about two months; the paper reports an average of 754 rewrites per day and that over 40% of detected abnormal addresses were corrected. The abstract and introduction further claim that deployment significantly decreased parcel re-routing by approximately 43% and reduced over 40% of re-routing among 2 million daily parcels.

Significance. If the deployment claim were substantiated, this would be a notable industrial application of LLM-based address rewriting with clear practical impact. The offline comparisons are extensive, use real-world data at scale, and include ablation studies that help isolate the contributions of SFT, RAG, and objective alignment. The paper also explicitly acknowledges several limitations, including the lack of a comprehensive error-type analysis and the conservative deployment design. However, the headline 43% re-routing reduction is not supported by the presented evidence; it conflates an offline station-level geocoding error reduction with a system-level re-routing outcome. The potential circularity between the RL reward and the evaluation ground truth also needs to be examined before the offline gains can be taken at face value.

major comments (4)
  1. [Abstract; §1; §3.4.1; §3.5] The claim that deployment 'significantly decreased the rate of parcel re-routing by approximately 43%' and 'reduced over 40% parcel re-routing caused by abnormal addresses among around 2 million daily parcels' is not supported by the data presented. Section 3.5 reports only daily rewriting volume (average 754) and an accuracy defined as the percentage of abnormal addresses corrected by AddrLLM; no before/after re-routing counts are given. With roughly 2 million parcels per day, 754 rewrites is about 0.04% of parcels, so even a 40% correction rate affects only ~300 parcels daily. The 43% figure in §3.4.1 is a relative reduction in station-level geocoding inaccuracies (SoP Acc@Station 90.0 vs. AddrLLM 94.3, i.e., inaccuracy 10.0 vs. 5.7), not a measured re-routing reduction. Please either supply deployment-level re-routing statistics or revise the abstract and contributions to state the actual measured quantity.
  2. [§2.2, Eq. (3); §3.1 Hit Rate] The RL reward includes a reverse-geocoding score, r_revgeo(y,c)=cos(f(y), f(reverse(c))), and the direct evaluation Hit Rate derives its ground truth by reverse geocoding the delivery coordinates using the same JD service. Because the same service supplies both the training reward and the test ground truth, the evaluation cannot detect systematic biases or errors in that service, and calling the feedback 'bias-free' in §2.2 is not justified. Please provide an independent validation set (e.g., manually inspected samples) or an analysis of reverse-geocoding error rates to show that the reward and the Hit Rate metric are not circularly reinforcing the model's outputs. If such validation is not feasible, the paper should explicitly state this limitation and temper the corresponding claims.
  3. [§3.5] The deployed-system 'accuracy' is defined as the percentage of abnormal addresses corrected by AddrLLM, but the paper does not specify how a correction is verified after deployment. If the verification uses the same geocoding or abnormal-address detection pipeline that triggered the rewrite, the reported correction rate may partly reflect the detector's own errors. Please describe the verification procedure used in the online system and, if possible, report a sample of rewrites that were manually checked.
  4. [§3.5; Appendix E] The deployment results for Zhejiang, Yulin, and Yangjiang are reported without a comparison period or a control group. To substantiate a re-routing reduction claim, the paper needs at least a before/after comparison of re-routing rates in the deployment region, or a controlled A/B experiment. As presented, the online section demonstrates stable operation and that some detected abnormal addresses are changed, but it does not demonstrate a reduction in re-routing events.
minor comments (5)
  1. [Abstract; §1; §3.4.1] The number 43% is used for two different quantities (station-level geocoding inaccuracy reduction and parcel re-routing reduction); please standardize the terminology so readers can distinguish the offline metric from the deployment outcome.
  2. [§3.5; Figure 5] The text says monitoring covered 60 days, but the date range in Figure 5 (2024/05/19 to 2024/07/14) spans 56 days; please reconcile this discrepancy.
  3. [References] References [61] and [62] are the same paper (Zhu et al., 2024); please deduplicate.
  4. [§3.4.5; Figure 3] The t-SNE visualization of station-level embeddings is qualitative; reporting a quantitative station-level classification accuracy (e.g., k-NN accuracy on held-out addresses) would strengthen RQ5.
  5. [§2.3; §3.2; Appendix C Table 7] The prompt in Appendix C says 'Related Address:{addresses returned by retriever}' without specifying the number of retrieved addresses, whereas §3.2 states top-10 is used; please make the prompt template consistent with the experimental setting.

Circularity Check

3 steps flagged · score 6.0 of 10

Offline Hit Rate and geocoding metrics reuse the same JD LBS reverse-geocoding/geocoding oracle that supplies the RL reward; the headline 43% re-routing reduction renames the offline station-error reduction.

  1. fitted input called prediction [Section 2.2, Eq. (3) and Section 3.1, Direct evaluation]
    "r_revgeo(y,c) = cos(f(y), f(reverse(c))) (3) where y is rewritten address, c is coordinates of successful delivery, reverse is reverse-geocoding service. ... Hit Rate: We derived the groundtruth for levels 1 to 4 (Appendix A) by reverse geocoding the delivery coordinates, complementing them with level 5 and 6 from the input addresses."

    The reward in Eq. (3) trains the model so that the rewritten address is semantically close to the output of JD's reverse-geocoding service applied to delivery coordinates. Section 3.1 then builds the direct-evaluation Hit Rate ground truth from the same reverse-geocoding service applied to the same delivery coordinates. The training target and the test label are therefore produced by the same oracle, so the metric cannot independently validate the rewrite. Any bias or noise in JD's reverse geocoder is baked into both the reward and the ground truth, making the reported Hit Rate a measure of fit to the training signal rather than an external correctness check.

  2. fitted input called prediction [Section 2.2, Eq. (4) and Section 3.1, Geocoding metrics]
    "geo(y,c) = ... k = dis(geocoding(y), c) (4) ... where dis is Euclidean distance function, geocoding is JD geocoding system that maps from address to coordinates. ... Acc@300m: The percentage of coordinates that locates within 300-meter radius of the groundtruth coordinates, with the groundtruth being the delivery coordinates as reported by couriers."

    The geocoding component of the reward, weighted by lambda_3 = 0.6, directly maximizes the closeness between JD geocoding of the rewrite and courier-reported delivery coordinates. The offline geocoding metrics Acc@300m, Acc@500m, and Acc@Station evaluate exactly that same distance to the same courier-reported coordinates. Thus the paper's reported geocoding accuracy and its 43% station-level inaccuracy reduction relative to SoP are measured on the same objective the model was optimized against. Held-out test addresses reduce data leakage but do not break the shared-oracle circularity: the evaluation benchmark and the training reward are the same quantity.

1 more flagged steps
  1. renaming known result [Abstract vs Section 3.4.1 and Section 3.5]
    "It has significantly decreased the rate of parcel re-routing by approximately 43%. ... compared with SoP method, AddrLLM reduces station-level inaccuracies by a substantial 43%. ... On average, there are 754 instances of address rewriting each day. Our specialized address rewriting model, AddrLLM, has effectively corrected over 40% of these erroneous addresses."

    The headline 'parcel re-routing' reduction is not a measured re-routing rate. The only 43% figure in the offline experiments is the relative reduction in station-level geocoding inaccuracy (Acc@Station 90.0 for SoP vs 94.3 for AddrLLM), while the online section reports 754 rewrites per day, roughly 0.04% of the 'around 2 million daily parcels,' with an unverified 'accuracy' of over 40% for those rewrites. Presenting the offline error reduction as a 43% drop in system-level parcel re-routing renames a different quantity and is not derivable from the monitoring data actually reported.

full rationale

The paper's core engineering contributions, including the SFT dataset construction, the address-centric RAG retriever, and the ablation study, are substantive and not circular by themselves. I found no load-bearing self-citation chain: the citations to the authors' prior FastAddr and CoMiner work are used for abnormal-address detection and coordinate mining, but the central rewriting claims do not reduce to those papers. However, the evaluation design is partially self-referential. The RL reward in Eq. (3) uses JD's reverse-geocoding service on delivery coordinates, and Section 3.1 constructs the Hit Rate ground truth from that same service; Eq. (4) trains the model to make JD geocoding land near delivery coordinates, and the Acc@300m/500m/Station metrics measure exactly that same distance. Consequently, the offline metrics are not independent of the training signal: they share the same LBS oracle, so a model optimized for that oracle will score well by construction. This is a partial circularity rather than a complete one, because held-out addresses, token-level Hit Rate, and comparisons against untrained baselines still provide some independent signal. Separately, the abstract's 43% parcel re-routing claim is unsupported: the reported 43% is an offline station-level inaccuracy reduction, and the online data show only 754 rewrites per day, which cannot substantiate a 43% reduction among roughly two million daily parcels. That overclaim is better classified as a correctness issue, but it also functions as a renaming of the offline metric as an online outcome. Overall, the paper is not wholly circular, but the main offline predictions are measurably entangled with their own training oracle, warranting a score of 6.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central results rest on JD's proprietary address services serving as both supervision and evaluation, on manually chosen reward weights and thresholds, and on a largely unspecified KL coefficient. These are load-bearing for the claimed improvements and none are independently verified.

free parameters (4)
  • reward weights lambda_1, lambda_2, lambda_3 = 0.2, 0.2, 0.6
    Manually set in Eq. 5 to combine semantics, reverse geocoding, and geocoding scores. No sensitivity analysis is reported, and these weights directly shape the RL-optimized policy.
  • geocoding score thresholds theta_1, theta_2 = 100 m, 1000 m
    Manually set in Eq. 4 to normalize the geocoding distance score. No ablation or sensitivity test is provided.
  • KL regularization coefficient beta = not reported
    Introduced in Eq. 8 to limit deviation from the SFT policy, but the paper never states its value, leaving the RL objective underspecified.
  • number of retrieved addresses (top-k) = 10
    Chosen in Section 3.2 for the RAG module; no comparison of different k values is reported.
assumptions (5)
  • standard math Autoregressive language modeling objective (Eq. 1) is a valid surrogate for address rewriting quality
    Standard maximum-likelihood training for LLMs; unproved background.
  • domain assumption The Chinese address hierarchy in Appendix A is the canonical standard and the target of rewriting
    The definition of 'standard' address is taken from JD's LBS system; no external or government standard is cited.
  • domain assumption The testing set ratio of 90% standard to 10% abnormal addresses matches real-world conditions
    Stated in Section 3.1 without empirical justification; this ratio drives the Robustness and Correction metrics.
  • ad hoc to paper JD LBS geocoding and reverse geocoding services provide unbiased and accurate coordinates and addresses for reward and evaluation
    The 'bias-free' objective alignment (Section 2.2) relies on these services, yet the same services define the evaluation ground truths (Section 3.1), so they cannot provide independent evidence.
  • ad hoc to paper Spatial proximity in a fine-tuned BERT embedding space is the correct relevance criterion for address retrieval
    Section 2.3 asserts this without end-to-end comparison of retrieval strategies; the t-SNE and R^2 analysis in Section 3.4.5 shows correlation but not downstream benefit.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data." pith.science (2026). https://pith.science/paper/H5QDTVEN

@misc{pith2026241113584,
  author       = {Pith},
  title        = {Pith review of: AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H5QDTVEN}},
  note         = {Machine review of arXiv:2411.13584}
}
read the original abstract

Textual description of a physical location, commonly known as an address, plays an important role in location-based services(LBS) such as on-demand delivery and navigation. However, the prevalence of abnormal addresses, those containing inaccuracies that fail to pinpoint a location, have led to significant costs. Address rewriting has emerged as a solution to rectify these abnormal addresses. Despite the critical need, existing address rewriting methods are limited, typically tailored to correct specific error types, or frequently require retraining to process new address data effectively. In this study, we introduce AddrLLM, an innovative framework for address rewriting that is built upon a retrieval augmented large language model. AddrLLM overcomes aforementioned limitations through a meticulously designed Supervised Fine-Tuning module, an Address-centric Retrieval Augmented Generation module and a Bias-free Objective Alignment module. To the best of our knowledge, this study pioneers the application of LLM-based address rewriting approach to solve the issue of abnormal addresses. Through comprehensive offline testing with real-world data on a national scale and subsequent online deployment, AddrLLM has demonstrated superior performance in integration with existing logistics system. It has significantly decreased the rate of parcel re-routing by approximately 43\%, underscoring its exceptional efficacy in real-world applications.

Figures

Figures reproduced from arXiv: 2411.13584 by the authors.

Figure 1
Figure 1. Parcel dispatching, re-routing and address rewriting. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Framework of AddrLLM [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 5
Figure 5. Result of Deployment that can be corrected by AddrLLM) is depicted in [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3: T-SNE visualization of address embedding generated [PITH_FULL_IMAGE:figures/full_fig_p008_3.png]
Figure 4
Figure 4. Figure 4: Embedding visualization of 50000 address pairs in [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 7
Figure 7. Figure 7: Result of deployment in Yangjiang [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

67 extracted references · 29 canonical work pages

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Flo- rencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...

  3. [3]

    Baichuan Inc. [n. d.]. Baichuan-7B Model. https://huggingface.co/baichuan- inc/Baichuan-7B. Accessed: 2024-07-11

  4. [4]

    Do, Yan Xu, and Pascale Fung

    Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V . Do, Yan Xu, and Pascale Fung. 2023. A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity. In Proceedings of the 13th International Joint Conference on Natural Langu...

  5. [5]

    Nils Boysen, Stefan Fedtke, and Stefan Schwerdfeger. 2021. Last-mile delivery concepts: a survey from an operational research perspective. OR Spectr. 43, 1 (2021), 1–58. https://doi.org/10.1007/S00291-020-00607-8

  6. [6]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  7. [7]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al . 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  8. [8]

    Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014)

Show all 67 references
  1. [9]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Under- standing. In Proceedings of the 2019 Conference of the North American Chap- ter of the Association for Computational Linguistics: H...

  2. [10]

    Ruixue Ding, Boli Chen, Pengjun Xie, Fei Huang, Xin Li, Qiang Zhang, and Yao Xu. 2023. MGeo: Multi-Modal Geographic Language Model Pre-Training. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, Taip...

  3. [11]

    Jianfeng Gao, Xiaodong He, and Jian-Yun Nie. 2010. Clickthrough-based transla- tion models for web search: from word models to phrase models. In Proceedings of the 19th ACM international conference on Information and knowledge manage- ment. 1139–1148

  4. [12]

    Jianfeng Gao and Jian-Yun Nie. 2012. Towards concept-based translation models using search logs for query expansion. In Proceedings of the 21st ACM interna- tional conference on Information and knowledge management. 1–10

  5. [13]

    Jianfeng Gao, Shasha Xie, Xiaodong He, and Alnur Ali. 2012. Learning lexicon models from search logs for query expansion. In Proceedings of EMNLP

  6. [14]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 (2023)

  7. [15]

    Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang

  8. [16]

    Muhammad Usman Hadi, Rizwan Qureshi, Abbas Shah, Muhammad Irfan, Anas Zafar, Muhammad Bilal Shaikh, Naveed Akhtar, Jia Wu, Seyedali Mirjalili, et al

  9. [17]

    Zhiqing Hong, Guang Wang, Wenjun Lyu, Baoshen Guo, Yi Ding, Haotian Wang, Shuai Wang, Yunhuai Liu, and Desheng Zhang. 2022. CoMiner: nationwide behavior-driven unsupervised spatial coordinate mining from uncertain delivery events. In Proceedings of the 30th International Confe...

  10. [18]

    Zhiqing Hong, Heng Yang, Haotian Wang, Wenjun Lyu, Yu Yang, Guang Wang, Yunhuai Liu, Yang Wang, and Desheng Zhang. 2022. FastAddr: real-time abnor- mal address detection via contrastive augmentation for location-based services. In Proceedings of the 30th International Conferen...

  11. [19]

    Mengting Hu, Xiaoqun Zhao, Jiaqi Wei, Jianfeng Wu, Xiaosu Sun, Zhengdan Li, Yike Wu, Yufei Sun, and Yuzhi Zhang. 2023. rT5: A Retrieval-Augmented Pre-trained Model for Ancient Chinese Entity Description Generation. In CCF In- ternational Conference on Natural Language Processi...

  12. [20]

    Jizhou Huang, Haifeng Wang, Yibo Sun, Yunsheng Shi, Zhengjie Huang, An Zhuo, and Shikun Feng. 2022. Ernie-geol: A geography-and-language pre-trained model and its applications in baidu maps. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Minin...

  13. [21]

    Yizheng Huang and Jimmy Huang. 2024. A Survey on Retrieval-Augmented Text Generation for Large Language Models. arXiv preprint arXiv:2404.10981 (2024)

  14. [22]

    Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research 24, 251 ...

  15. [23]

    Nicolas Jonason, Luca Casini, Carl Thomé, and Bob LT Sturm. 2023. Re- trieval Augmented Generation of Symbolic Music with LLMs. arXiv preprint arXiv:2311.10384 (2023)

  16. [24]

    Vishal Kakkar and T Ravindra Babu. 2018. Address Clustering for e-Commerce Applications.. In eCOM@ SIGIR

  17. [25]

    Ravindra Babu

    Vishal Kakkar and T. Ravindra Babu. 2018. Address Clustering for e-Commerce Applications. In The SIGIR 2018 Workshop On eCommerce co-located with the 41st International ACM SIGIR Conference on Research and Development in Infor- mation Retrieval (SIGIR 2018), Ann Arbor, Michiga...

  18. [26]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http://arxiv.org/abs...

  19. [27]

    Simone Kresevic, Mauro Giuffrè, Milos Ajcevic, Agostino Accardo, Lory S Crocè, and Dennis L Shung. 2024. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. NPJ Digital Medicine 7, 1 (2024), 102

  20. [28]

    Rossi, and Thien Huu Nguyen

    Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Der- noncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2023. Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback. In Proceedings of the 2023 Conf...

  21. [29]

    Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. nature 521, 7553 (2015), 436–444

  22. [30]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58t...

  23. [31]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing S...

  24. [32]

    Dongyang Li, Ruixue Ding, Qiang Zhang, Zheng Li, Boli Chen, Pengjun Xie, Yao Xu, Xin Li, Ning Guo, Fei Huang, et al. 2023. Geoglue: A geographic language understanding evaluation benchmark. arXiv preprint arXiv:2305.06545 (2023)

  25. [33]

    Haochen Li, Xin Zhou, and Zhiqi Shen. 2024. Rewriting the Code: A Simple Method for Large Language Model Augmented Code Search. arXiv preprint arXiv:2401.04514 (2024)

  26. [34]

    Jie Li, Haifeng Liu, Chuanghua Gui, Jianyu Chen, Zhenyun Ni, and Ning Wang

  27. [35]

    Jie Liu and Barzan Mozafari. 2024. Query Rewriting via Large Language Models. arXiv preprint arXiv:2403.09060 (2024)

  28. [36]

    Xiao Liu, Juan Hu, Qi Shen, and Huan Chen. 2021. Geo-bert pre-training model for query rewriting in poi search. InFindings of the Association for Computational Linguistics: EMNLP 2021. 2209–2214. AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data ...

  29. [37]

    Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net. https://openreview.net/ forum?id=Bkg6RiCqY7

  30. [38]

    Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A Smith. 2021. Time waits for no one! analysis and challenges of temporal misalignment. arXiv preprint arXiv:2111.07408 (2021)

  31. [39]

    Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query Rewriting in Retrieval-Augmented Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamo...

  32. [40]

    Wenjun Peng, Guiyang Li, Yue Jiang, Zilong Wang, Dan Ou, Xiaoyi Zeng, Derong Xu, Tong Xu, and Enhong Chen. 2024. Large Language Model based Long-tail Query Rewriting in Taobao Search. In Companion Proceedings of the ACM on Web Conference 2024, WWW 2024, Singapore, Singapore, M...

  33. [41]

    Yiming Qiu, Kang Zhang, Han Zhang, Songlin Wang, Sulong Xu, Yun Xiao, Bo Long, and Wen-Yun Yang. 2021. Query Rewriting via Cycle-Consistent Translation for E-Commerce Search. In 37th IEEE International Conference on Data Engineering, ICDE 2021, Chania, Greece, April 19-22, 202...

  34. [42]

    Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi. 2022. Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural langua...

  35. [43]

    Stefan Riezler and Yi Liu. 2010. Query rewriting using monolingual statistical machine translation. Computational Linguistics 36, 3 (2010), 569–582

  36. [44]

    Stefan Riezler, Yi Liu, and Alexander Vasserman. 2008. Translating queries into snippets for improved query expansion. In Proceedings of the 22nd international conference on computational linguistics (Coling 2008). 737–744

  37. [45]

    Pravakar Roy, Chirag Sharma, Chao Gao, and Kumarswamy Valegerepura. 2023. Deep Query Rewriting For Geocoding. In Proceedings of the 32nd ACM Interna- tional Conference on Information and Knowledge Management. 4801–4807

  38. [46]

    John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel

  39. [47]

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov

  40. [48]

    Lei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu, Yinxiao Liu, Simon Tong, Jindong Chen, and Lei Meng. 2024. RewriteLM: An Instruction-Tuned Large Language Model for Text Rewriting. In Thirty-Eighth AAAI Conference on Artifi- cial Intelligence, AAAI 2024, Thirty-Sixth Confer...

  41. [49]

    Lei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu, Yinxiao Liu, Simon Tong, Jindong Chen, and Lei Meng. 2024. Rewritelm: An instruction-tuned large language model for text rewriting. In Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 38. 18970–18980

  42. [50]

    Ravindra Babu Tallamraju. 2022. Geographical Address Models in the Indian e-Commerce. In Proceedings of the 31st ACM International Conference on Infor- mation & Knowledge Management. 5096–5097

  43. [51]

    Qin Tian, Fu Ren, Tao Hu, Jiangtao Liu, Ruichang Li, and Qingyun Du. 2016. Using an optimized Chinese address matching method to develop a geocoding service: a case study of Shenzhen, China. ISPRS International Journal of Geo- Information 5, 5 (2016), 65

  44. [52]

    Hanwen Tong, Chenhao Xie, Jiaqing Liang, Qianyu He, Zhiang Yue, Jingping Liu, Yanghua Xiao, and Wenguang Wang. 2022. A context-enhanced generate-then- evaluate framework for chinese abbreviation prediction. In Proceedings of the 31st ACM International Conference on Information...

  45. [53]

    Christopher Toukmaji and Allison Tee. 2024. Retrieval-Augmented Generation and LLM Agents for Biomimicry Design Solutions. In Proceedings of the AAAI Symposium Series, V ol. 3. 273–278

  46. [54]

    Binghai Wang, Rui Zheng, Lu Chen, Yan Liu, Shihan Dou, Caishuang Huang, Wei Shen, Senjie Jin, Enyu Zhou, Chenyu Shi, Songyang Gao, Nuo Xu, Yuhao Zhou, Xiaoran Fan, Zhiheng Xi, Jun Zhao, Xiao Wang, Tao Ji, Hang Yan, Lixing Shen, Zhan Chen, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjin...

  47. [55]

    Lixia Wu, Jianlin Liu, Junhong Lou, Haoyuan Hu, Jianbin Zheng, Haomin Wen, Chao Song, and Shu He. 2023. G2ptl: A pre-trained model for delivery address and its applications in logistics system. arXiv preprint arXiv:2304.01559 (2023)

  48. [56]

    Mengyuan Yang, Mengying Zhu, Yan Wang, Linxun Chen, Yilei Zhao, Xiuyuan Wang, Bing Han, Xiaolin Zheng, and Jianwei Yin. 2024. Fine-Tuning Large Language Model Based Explainable Recommendation with Explainable Quality Reward. In Thirty-Eighth AAAI Conference on Artificial Intel...

  49. [57]

    Narasimhan, and Yuan Cao

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. Open...

  50. [58]

    Fanghua Ye, Meng Fang, Shenghui Li, and Emine Yilmaz. 2023. Enhanc- ing Conversational Search: Large Language Model-Aided Informative Query Rewriting. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan...

  51. [59]

    Han Yu, Peikun Guo, and Akane Sano. 2023. Zero-Shot ECG Diagnosis with Large Language Models and Retrieval-Augmented Generation. In Machine Learning for Health (ML4H). PMLR, 650–663

  52. [60]

    Mengxi Yu, Ziyu Liu, Yuhang Tang, and Jianfeng Jiang. 2021. Recognition algorithm of e-commerce click farming based on K-means technology. In 2021 6th International Conference on Intelligent Computing and Signal Processing (ICSP). IEEE, 103–106

  53. [61]

    Hongyi Zhu, Jia-Hong Huang, Stevan Rudinac, and Evangelos Kanoulas. 2024. En- hancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models. In Proceedings of the 2024 International Conference on Multimedia Retrieval. 978–987

  54. [62]

    Hongyi Zhu, Jia-Hong Huang, Stevan Rudinac, and Evangelos Kanoulas. 2024. En- hancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models. In Proceedings of the 2024 International Conference on Multimedia Retrieval. 978–987. ...

  55. [2015]

    arXiv preprint arXiv:1506.02438 (2015)

    High-dimensional continuous control using generalized advantage estima- tion. arXiv preprint arXiv:1506.02438 (2015)

  56. [2017]

    arXiv preprint arXiv:1707.06347 (2017)

    Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)

  57. [2019]

    arXiv:1908.07389 [cs.IR]

    The Design and Implementation of a Real Time Visual Search System on JD E-commerce Platform. arXiv:1908.07389 [cs.IR]

  58. [2020]

    In International confer- ence on machine learning

    Retrieval augmented language model pre-training. In International confer- ence on machine learning. PMLR, 3929–3938

  59. [2023]

    Authorea Preprints (2023)

    A survey on large language models: Applications, challenges, limitations, and practical usage. Authorea Preprints (2023)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.