REVIEW 4 major objections 5 minor 2 cited by
Time Series Language Model for Descriptive Caption Generation
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that a 1-billion-parameter encoder-decoder fed a phase-tagged text plus reprogrammed numeric embedding of a time series captions stock and synthetic series better than far larger vision-language and chat models.
desk verdict A solid, honest engineering paper for time series captioning whose headline ROUGE/BERTScore margins hinge on an underspecified best-of-K evaluation that must be clarified before the numbers can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the joint representation $JR(T)$ of a time series, which concatenates a phase-tagged textual series with a reprogrammed time series embedding. Three supporting mechanisms carry the argument: the in-context prompting generator, which bootstraps LLaMA2-13B-Chat with grouped demonstrations to produce diverse synthetic pairs; the cross-modal dense retrieval scorer, trained with in-batch negatives to maximize the dot-product similarity between the [CLS] vectors of series and caption, which filters the synthetic data and then doubles as the TSLMScore evaluation metric; and the reprogramming layer, which uses cross-attention between time series embeddings and text prototypes to align the two modalities before the transformer blocks. The ablation variants TSLM (Text), TSLM (TimeSeries), and TSLM (w/o denoising) define what each mechanism contributes to the reported results.
What would settle it
Have human annotators label a random sample of the 203,554 generated pairs as matching or non-matching without seeing the scorer's outputs, then compare the scorer's scores on the kept pairs ($T_h \geq 0$) with those on the 15,473 removed pairs; if the removed pairs are not systematically the ones humans reject, the reported 6.71% R-L gain from denoising cannot be attributed to the filter as described.
Extended reading notes
Core claim
The central claim is that the captioning task is best served by jointly representing the time series in two modalities: a position-aware textual form obtained by writing the values as tokens and wrapping the sequence in <start>, <middle>, and <end> tags, and a fine-grained embedding produced by a frozen 1D-CNN autoencoder. The embedding is reprogrammed into the text-embedding space through a cross-attention layer over learned text prototypes, and the concatenated sequence is fed to a transformer encoder-decoder initialized from T5-large, trained with next-token prediction. To obtain enough training data, the authors generate synthetic pairs by prompting LLaMA2-13B-Chat in-context with grouped demonstrations, then train a cross-modal dense retrieval scorer on the original pairs and discard every synthetic pair whose series-caption dot-product similarity falls below zero, removing 15,473 of 203,554 pairs. The paper reports that the resulting model outperforms all baselines on both datasets — for instance reaching R-L 66.45 and BERTScore 0.80 on STOCK against 63.25 and 0.78 for LLaMA2-70B-Chat — and that the same dense retrieval scorer, reused as the TSLMScore evaluation metric, ranks the joint model first.
Load-bearing premise
The load-bearing premise is that a similarity scorer trained on a few thousand original pairs reliably distinguishes, across more than two hundred thousand generated pairs, which captions actually match their series — even though the paper itself notes in Appendix F that this denoising step was validated only by qualitative manual inspection.
Editorial extensions
If this is right
- If the reported numbers hold, a 1-billion-parameter model with a purpose-built time series encoder beats conversational models with 70 billion parameters on this task, making captioning at scale far cheaper in time and memory.
- The 6.71% R-L gain of TSLM over TSLM (w/o denoising) implies that LLM-generated synthetic captioning data contains enough plausible-but-wrong pairs that filtering them is worth more than adding more raw data.
- Because TSLM (Text) beats all text-only baselines except the 70B model, the three-phase tagging itself is a transferable, parameter-free way to give an LLM positional information about a numeric sequence.
- The generation pipeline of multiple sampled captions plus an LLM summarizer means a general-purpose LLM can describe time series without ever seeing raw numbers, reducing its direct hallucination burden.
- The threshold analysis, showing best results at $T_h = 0$, implies that trimming only the left tail of the score distribution suffices and that aggressive filtering removes useful data.
Reading between the lines
- If the recipe transfers, the same combination of phase tagging, a frozen convolutional encoder, a reprogramming cross-attention layer, and denoised synthetic data from an open-source LLM could apply to other data-scarce series-to-text tasks such as sensor logs, patient vitals, or network alarms, since none of the components is stock-specific.
- Because the TSLMScore metric is computed by the same dense retrieval scorer that filters the training data, its rankings could reflect the scorer's own preferences; an independent evaluation with human judgments or a separately trained scorer would test whether the filter genuinely improves caption quality rather than merely matching its own signal.
- A testable extension would be to run the same threshold-based filtering with a rule-based or independent LLM judge and compare the kept sets; if the two filters agree on which pairs are noisy, the dense-retrieval scorer's specific architecture matters less than the act of filtering itself.
- The temperature sweep suggests that generating diverse candidate captions and then summarizing them drives accuracy, so weighting the summarizer's input by each candidate's score or self-consistency is a natural untested follow-up.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes TSLM, an encoder-decoder model for time series captioning. Time series are represented jointly by a phase-tagged textual sequence and by embeddings from a frozen 1D-CNN autoencoder; a reprogramming layer aligns the embeddings with text prototypes, and the fused representation is fed to a T5-initialized transformer. To address data scarcity, the authors generate 203,554 synthetic time series-caption pairs via in-context prompting with LLaMA2-13B-Chat, then filter noisy pairs with a cross-modal dense retrieval scorer trained on the groundtruth data. The model is trained on the original plus denoised generated data and evaluated on the STOCK and SYNTH datasets against time-series, image, text-decoder, and text-encoder-decoder baselines. The paper reports that TSLM outperforms all baselines on ROUGE-1/2/L, BERTScore, and a proposed TSLMScore, with ablations showing the joint representation and denoising each contribute gains.
Significance. If the reported comparisons hold, the paper makes a useful contribution to an under-explored task, time series captioning, and demonstrates a practical recipe: joint phase-tagged text plus reprogrammed CNN embeddings, trained on denoised LLM-generated synthetic data. The paper is clearly structured, includes ablations over modality, model size, denoising, and data percentage, and covers baselines from several modalities. The central claim is nevertheless not yet verifiable because the protocol for scoring the K=3 generated captions is unspecified. In addition, the proposed TSLMScore is computed with the paper's own retrieval model and is therefore not an independent evaluation metric. The ROUGE/BERTScore ranking is externally grounded and does not depend on TSLMScore, so the main comparison could be salvaged by a precise and equal evaluation protocol. The paper does not provide code, seeds, or error bars, which further limits reproducibility.
major comments (4)
- [Section 6; Section 7.4, Table 1] The evaluation protocol for the K=3 generated captions is unspecified. Section 6 says TSLM generates K captions and then summarizes them with LLaMA2-13B-Chat into a "descriptive caption," and Section 7.3 sets K=3. Section 7.4 says "The generated captions from TSLM and baselines are compared to the groundtruth captions" but does not state whether the quantitative scores are computed on each of the K captions and then averaged or maximized, on a single sampled caption, or on the LLaMA-summarized descriptive caption, nor whether baselines also produce K captions. If the reported TSLM numbers are the best of K while baselines are scored on one caption, or if the summarizer is applied only to TSLM, Table 1 is not a comparison of equal prediction tasks. This is the load-bearing step for the headline claim and must be documented precisely.
- [Section 7.4, Equations (13)-(14)] TSLMScore is computed with the same cross-modal dense retrieval model that is trained on the groundtruth caption pairs and shares the embedding and transformer layers with the multi-modal encoder, and it uses the same joint representation construction as TSLM. As a result, TSLMScore measures proximity in TSLM's own embedding space; it will tend to reward captions that lie close to that model's representations and is not an independent measure of caption quality. The ROUGE/BERTScore columns do not depend on this metric, but the paper should either remove TSLMScore from the headline comparisons or validate it against human judgments across model families.
- [Appendix C, Table 3; Section 7.3] The denoising threshold Th appears to be selected using test-set results. Section 7.3 fixes Th=0, and Appendix C compares multiple thresholds by reporting "evaluation metrics on the testing sets" and concludes the optimal range from those test metrics. If the test set was used to choose Th, the reported numbers for the full model and the claimed 6.71% R-L gain from denoising are optimistically selected. Please describe a validation-based selection of Th or otherwise justify that Th=0 was fixed before test-set evaluation.
- [Section 7.3; Table 1] No variance information is reported for any method: there are no multiple seeds, standard deviations, or significance tests. Without this, the abstract's claim of outperforming state-of-the-art approaches "by a significant margin" is not statistically supported, and the magnitude of the margins in Table 1 cannot be assessed for seed sensitivity. Please report runs over at least three seeds and include variance or a paired test for the main comparisons.
minor comments (5)
- [Section 4.1.1, Eq. (1)] The phase split uses T_{1:l/3}, T_{l/3+1:2l/3}, and T_{2l/3+1:l}; for lengths not divisible by 3, this division is undefined. Please clarify the rounding or use explicit index sets.
- [Table 1] The TRUCE row shows "–" for R-1, R-2, and TSLMScore; please clarify whether these values were not reported in the original TRUCE paper or are not applicable.
- [Appendix D, Figure 6] The temperature analysis does not state which sampling parameters, such as top-k and top-p, are held fixed. Since Section 7.3 fixes top-p=0.95 and top-k=50, please clarify whether these values are also used in Figure 6.
- [Section 5.3] The passage "Once the joint embedding space is learned, ... Therefore, Once the joint embedding space is learned" contains a capitalization/line-break error; please revise for readability.
- [Section 7.3] The paper states that "We use the same training, validation, and testing splits of TRUCE" but does not report the size of the validation split or how hyperparameters other than Th were selected; please state the split sizes and selection procedure.
Circularity Check
Partial circularity: TSLMScore reuses the paper's own denoising model, but the main ROUGE/BERTScore comparisons remain externally grounded.
-
fitted input called prediction
[Section 7.4 (TSLMScore), reusing Eq. (14) from Section 5.3]
"In addition, we use our trained denoising cross-modal dense retrieval model to report a new score denoted by TSLMScore. Given an unseen time series 𝑇 and a predicted caption𝑐, we compute their similarity using the dot product of their embeddings that are extracted using the trained cross-modal dense retrieval model as shown by Equation (14)."
The TSLMScore is computed with the same model that, in Section 5.3, was trained on the groundtruth pairs and then used to filter the synthetic training data: 'we score each pair in the generated data using sim(𝑇,𝑐)' and remove pairs below threshold 𝑇ℎ (𝑇ℎ=0 in Section 7.3). TSLM is trained on the resulting denoised data. Reusing this same fitted similarity function as the evaluation metric means TSLM is scored by the very function that selected its training set, so a high TSLMScore for TSLM is partly guaranteed by construction rather than by independent verification. This self-referentiality is compounded by Appendix F's admission that the denoising step 'is only evaluated qualitatively with manual inspection.' The ROUGE/BERTScore columns are external and do not inherit this circularity.
full rationale
The main quantitative claim—that TSLM outperforms baselines on ROUGE-1/2/L and BERTScore—is evaluated against external text-matching metrics and is not circular; those columns in Table 1 are self-contained and would stand even if TSLMScore were removed. No load-bearing self-citation or uniqueness import is present: the reprogramming reference [27] is an external method, and the authors' own prior work appears only in the related-work list. The one genuine circular element is TSLMScore: Eq. (14) defines a dot-product similarity using a cross-modal retrieval model that was first fit to the groundtruth pairs and then used to filter the generated training data; Section 7.4 reuses that same trained model as a headline evaluation score. Because TSLM was trained on data filtered by this very scorer, its TSLMScore advantage is partly by construction and is further weakened by Appendix F's admission that the denoising was only qualitatively validated. The K-caption aggregation ambiguity noted by the skeptic is a reproducibility issue, not a circularity, and is not scored here. Overall, the paper's central derivation is independent, with one self-referential auxiliary metric; score 4.
Assumptions & free parameters
free parameters (5)
- Denoising threshold T_h =
0 in main text; T_h=1 (about mu-sigma) is better on several SYNTH rows and on STOCK TSLMScore in Appendix C
- Temperature of LLaMA2-13B-Chat =
0.7 for data generation; 0.95 for caption sampling
- Number of phase tags =
3 (start, middle, end)
- Number of generated captions K =
3
- Compressed series length f and text prototype count p =
f=6, p=1000
assumptions (5)
- domain assumption LLaMA2-13B-Chat, prompted with training-set demonstrations, generates time series-caption pairs whose joint distribution is close enough to real data to improve a downstream captioning model after filtering.
- domain assumption A cross-modal dense retrieval scorer trained on the small ground-truth sets reliably separates correct from hallucinated generated pairs across 203,554 synthetic samples.
- domain assumption Three-phase segmentation of the numeric series, with values written as text tokens, preserves the information needed to generate the reference captions.
- domain assumption A single model trained on the merged STOCK+SYNTH data plus synthetic data is a fair comparison against baselines reported per dataset.
- standard math In-batch negatives provide a valid contrastive signal for the retrieval scorer at batch size 8.
invented entities (1)
-
TSLMScore
Cite this review
Pith. "Pith review of Time Series Language Model for Descriptive Caption Generation." pith.science (2026). https://pith.science/paper/VXVIXEJB
@misc{pith2026250101832,
author = {Pith},
title = {Pith review of: Time Series Language Model for Descriptive Caption Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VXVIXEJB}},
note = {Machine review of arXiv:2501.01832}
}
read the original abstract
The automatic generation of representative natural language descriptions for observable patterns in time series data enhances interpretability, simplifies analysis and increases cross-domain utility of temporal data. While pre-trained foundation models have made considerable progress in natural language processing (NLP) and computer vision (CV), their application to time series analysis has been hindered by data scarcity. Although several large language model (LLM)-based methods have been proposed for time series forecasting, time series captioning is under-explored in the context of LLMs. In this paper, we introduce TSLM, a novel time series language model designed specifically for time series captioning. TSLM operates as an encoder-decoder model, leveraging both text prompts and time series data representations to capture subtle temporal patterns across multiple phases and generate precise textual descriptions of time series inputs. TSLM addresses the data scarcity problem in time series captioning by first leveraging an in-context prompting synthetic data generation, and second denoising the generated data via a novel cross-modal dense retrieval scoring applied to time series-caption pairs. Experimental findings on various time series captioning datasets demonstrate that TSLM outperforms existing state-of-the-art approaches from multiple data modalities by a significant margin.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 2 Pith papers
-
Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language
CGTime trains a time-series-language model on deterministically computed statistics, using them as both supervision and reward, and reports strong performance on its own multivariate benchmark compared with larger gen...
-
A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
The authors organize LLM-based time series reasoning into three exclusive topologies (direct, chain, branch) crossed with four objectives, and use them to label 125 papers, benchmarks, and resources.
Reference graph
Works this paper leans on
-
[1]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binko...
work page 2022
-
[2]
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016. Neural Module Networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016. IEEE Computer Society, 39–48
work page 2016
-
[3]
Kamal Berahmand, Fatemeh Daneshfar, Elaheh Sadat Salehi, Yuefeng Li, and Yue Xu. 2024. Autoencoders and their applications in machine learning: a survey. Artif. Intell. Rev. 57, 2 (2024), 28
work page 2024
-
[4]
Ching Chang, Wen-Chih Peng, and Tien-Fu Chen. 2023. LLM4TS: Two-Stage Fine- Tuning for Time-Series Forecasting with Pre-Trained LLMs.CoRR abs/2308.08469 Mohamed Trabelsi, Aidan Boyd, Jin Cao, and Huseyin Uzunalioglu (2023)
arXiv 2023
-
[5]
Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, and Brian D. Davison
-
[6]
Hao Cheng, Qingsong Wen, Yang Liu, and Liang Sun. 2024. RobustTSF: Towards Theory and Design of Robust Time Series Forecasting with Anomalies. CoRR abs/2402.02032 (2024)
arXiv 2024
-
[7]
Yew Ken Chia, Lidong Bing, Soujanya Poria, and Luo Si. 2022. RelationPrompt: Leveraging Prompts to Generate Synthetic Data for Zero-Shot Relation Triplet Extraction. In Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022 . Association for Computational Linguistics, 45–57
work page 2022
-
[8]
Zhuyun Dai and Jamie Callan. 2019. Deeper Text Understanding for IR with Contextual Neural Language Modeling. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval . 4 pages
work page 2019
Show all 78 references
-
[9]
Fatoumata Dama and Christine Sinoquet. 2021. Time Series Analysis and Model- ing to Forecast: a Survey. https://api.semanticscholar.org/CorpusID:237940640
2021
-
[10]
Webb, Shirui Pan, Charu C
Zahra Zamanzadeh Darban, Geoffrey I. Webb, Shirui Pan, Charu C. Aggarwal, and Mahsa Salehi. 2022. Deep Learning for Time Series Anomaly Detection: A Survey. CoRR abs/2211.05244 (2022)
2022 arXiv
-
[11]
Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. 2023. A decoder-only foundation model for time-series forecasting. CoRR abs/2310.10688 (2023)
2023 arXiv
-
[12]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs. In Advances in Neural Informa- tion Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023
2023
-
[13]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT
2019
-
[14]
Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia, Boyang Li, Shafiq Joty, and Lidong Bing. 2023. Is GPT-3 a Good Data Annotator?. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, ACL 2023 . Association for Computational Linguistic...
2023
-
[15]
Angela Fan, Mike Lewis, and Yann N. Dauphin. 2018. Hierarchical Neural Story Generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018 . Association for Computational Linguistics, 889–898
2018
-
[16]
Azul Garza and Max Mergenthaler Canseco. 2023. TimeGPT-1. CoRR abs/2310.03589 (2023)
2023 arXiv
-
[17]
Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen. 2023. MultiModal-GPT: A Vision and Language Model for Dialogue with Humans. CoRR abs/2305.04790 (2023)
2023 arXiv
-
[18]
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News Summarization and Evaluation in the Era of GPT-3. CoRR abs/2209.12356 (2022)
2022 arXiv
-
[19]
Xiao Han, Shuhan Yuan, and Mohamed Trabelsi. 2023. LogGPT: Log Anomaly Detection via GPT. In IEEE International Conference on Big Data, BigData 2023 . IEEE, 1117–1122
2023
-
[20]
Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. 2023. ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023
2023
-
[21]
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curi- ous Case of Neural Text Degeneration. In International Conference on Learning Representations
2020
-
[22]
Borg- wardt
Max Horn, Michael Moor, Christian Bock, Bastian Rieck, and Karsten M. Borg- wardt. 2020. Set Functions for Time Series. InProceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Pro- ceedings of Machine Learning Research...
2020
-
[23]
Weinberger
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger
-
[24]
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, Yi Ren, Yuexian Zou, Zhou Zhao, and Shinji Watanabe. 2024. AudioGPT: Understanding and Generat- ing Speech, Music, Sound, and Talking Head. In T...
2024
-
[25]
Harsh Jhamtani and Taylor Berg-Kirkpatrick. 2021. Truth-Conditional Caption- ing of Time Series Data. In EMNLP
2021
-
[26]
Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. In Pro- ceedings of the 38th International Conference on...
2021
-
[27]
Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen
Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y. Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen. 2024. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. In The Twelfth International Conference on Learnin...
2024
-
[28]
Price, Scott Cohen, and Christopher Kanan
Kushal Kafle, Robik Shrestha, Brian L. Price, Scott Cohen, and Christopher Kanan
-
[29]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, ...
2020
-
[30]
Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020. ACM, 39–48
2020
-
[31]
In IEEE Winter Conference on Applications of Computer Vision, W ACV
Answering Questions about Data Visualizations using Efficient Bimodal Fusion. In IEEE Winter Conference on Applications of Computer Vision, W ACV
-
[32]
Pengzhi Li, Yan Pei, and Jianqiang Li. 2023. A comprehensive survey on design and application of autoencoder in deep learning. Appl. Soft Comput. 138 (2023), 110176
2023
-
[33]
Yuxuan Liang, Haomin Wen, Yuqi Nie, Yushan Jiang, Ming Jin, Dongjin Song, Shirui Pan, and Qingsong Wen. 2024. Foundation Models for Time Series Analysis: A Tutorial and Survey. CoRR abs/2403.14735 (2024)
2024 arXiv
-
[34]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58t...
2020
-
[35]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023
2023
-
[36]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. ArXiv abs/1907.11692 (2019)
2019 arXiv
-
[37]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81
2004
-
[38]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regular- ization. In 7th International Conference on Learning Representations, ICLR 2019 . OpenReview.net
2019
-
[39]
Anita Mahinpei, Zona Kostic, and Chris Tanner. 2022. LineCap: Line Charts for Data Visualization Captioning Models. In 2022 IEEE Visualization and Visual Analytics (VIS). IEEE, 35–39
2022
-
[40]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021. IEEE, 9992–10002
2021
-
[41]
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2024. Recent Ad- vances in Natural Language Processing via Large Pre-trained Language Models: A Survey. ACM Comput. Surv. 56, 2 (2024), 30:1–30:40
2024
-
[42]
Soichiro Murakami, Akihiko Watanabe, Akira Miyazawa, Keiichi Goshima, Toshi- hiko Yanase, Hiroya Takamura, and Yusuke Miyao. 2017. Learning to Generate Market Comments from Stock Prices. In Proceedings of the 55th Annual Meet- ing of the Association for Computational Linguisti...
2017
-
[43]
van Panhuis, and Christos Falout- sos
Yasuko Matsubara, Yasushi Sakurai, Willem G. van Panhuis, and Christos Falout- sos. 2014. FUNNEL: automatic mining of spatially coevolving epidemics. In The 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2014. ACM, 105–114
2014
-
[44]
Rodrigo Frassetto Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019. Multi-Stage Document Ranking with BERT. CoRR abs/1910.14424 (2019)
2019 arXiv
-
[45]
Patil, Tianjun Zhang, Xin Wang, and Joseph E
Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. 2023. Gorilla: Large Language Model Connected with Massive APIs. CoRR abs/2305.15334 (2023)
2023 arXiv
-
[46]
Rodrigo Frassetto Nogueira and Kyunghyun Cho. 2019. Passage Re-ranking with BERT. CoRR abs/1901.04085 (2019)
2019 arXiv
-
[47]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Time Series Language Model for Descriptive Caption Generation Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable V...
2021
-
[48]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res. 21 (2020), 140:1–140:67
2020
-
[49]
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023. ToolLLM: Facilitating Large Language Models to Master 1...
2023 arXiv
-
[50]
Wataru Sakata, Tomohide Shibata, Ribeka Tanaka, and Sadao Kurohashi. 2019. FAQ Retrieval Using Query-Question Similarity and BERT-Based Query-Answer Relevance. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval . ...
2019
-
[51]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. In Advances in Neural Information Processing Systems 36: Annual...
2023
-
[52]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Rad- ford, Mark Chen, and Ilya Sutskever. 2021. Zero-Shot Text-to-Image Generation. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021 (Proceedings of Machine Learning Re...
2021
-
[53]
Xiaofei Sun, Xiaoya Li, Jiwei Li, Fei Wu, Shangwei Guo, Tianwei Zhang, and Guoyin Wang. 2023. Text Classification via Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023. Association for Computational L...
2023
-
[54]
Mohamed Trabelsi, Jin Cao, and Jeff Heflin. 2020. Semantic Labeling Using a Deep Contextualized Language Model. CoRR abs/2010.16037 (2020)
2020 arXiv
-
[55]
Pranay Kumar Venkata Sowdaboina, Sutanu Chakraborti, and Somayajulu Sri- pada. 2014. Learning to Summarize Time Series Data. In Computational Linguis- tics and Intelligent Text Processing - 15th International Conference, CICLing 2014 (Lecture Notes in Computer Science, Vol. 84...
2014
-
[56]
Davison, and Jeff Heflin
Mohamed Trabelsi, Zhiyu Chen, Brian D. Davison, and Jeff Heflin. 2021. Neural ranking models for document retrieval. Inf. Retr. J. 24, 6 (2021), 400–444
2021
-
[57]
Davison, and Jeff Heflin
Mohamed Trabelsi, Zhiyu Chen, Shuo Zhang, Brian D. Davison, and Jeff Heflin
-
[58]
Mohamed Trabelsi, Jin Cao, and Jeff Heflin. 2021. SeLaB: Semantic Labeling with BERT. In 2021 International Joint Conference on Neural Networks (IJCNN) . 1–8. https://doi.org/10.1109/IJCNN52387.2021.9534408
2021
-
[59]
Mohamed Trabelsi and Hüseyin Uzunalioglu. 2023. Absformer: Transformer- Based Model for Unsupervised Multi-Document Abstractive Summarization. In Document Analysis and Recognition - ICDAR 2023 Workshops (Lecture Notes in Computer Science, Vol. 14194). Springer, 151–166
2023
-
[60]
David Wadden, Ulme Wennberg, Yi Luan, and Hannaneh Hajishirzi. 2019. Entity, Relation, and Event Extraction with Contextualized Span Representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Pro- cessing and the 9th International Joint Con...
2019
-
[61]
Difeng Wang, Wei Hu, Ermei Cao, and Weijian Sun. 2020. Global-to-Local Neural Networks for Document-Level Relation Extraction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020 . Association for Computational Linguistics, 3711–3721
2020
-
[62]
Mohamed Trabelsi, Jeff Heflin, and Jin Cao. 2022. DAME: Domain Adaptation for Matching Entities. In Proceedings of the 15th ACM International Conference on Web Search and Data Mining (WSDM 2022)
2022
-
[63]
Yiyu Wang, Jungang Xu, and Yingfei Sun. 2022. End-to-End Transformer Based Model for Image Captioning. In Thirty-Sixth AAAI Conference on Artificial Intelli- gence, AAAI 2022, Thirty-Fourth Conference on Innovative Applications of Artificial Intelligence, IAAI 2022, The Twelve...
2022
-
[64]
Qingsong Wen, Linxiao Yang, Tian Zhou, and Liang Sun. 2022. Robust Time Series Analysis and Applications: An Industrial Perspective. In KDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022 . ACM, 4836–4837
2022
-
[65]
Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, and Doyen Sahoo. 2024. Unified Training of Universal Time Series Forecasting Transformers. CoRR abs/2402.02592 (2024)
2024 arXiv
-
[66]
Liang Wang, Nan Yang, and Furu Wei. 2024. Learning to Retrieve In-Context Examples for Large Language Models. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics . Association for Computational Linguistics, 1752–1767
2024
-
[67]
Hongju Yan and Hongbing Ouyang. 2018. Financial Time Series Prediction Based on Deep Learning. Wirel. Pers. Commun. 102, 2 (2018), 683–700
2018
-
[68]
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pe...
2021
-
[69]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi
-
[70]
Jingjing Xu, Caesar Wu, Yuan-Fang Li, and Pascal Bouvry. 2024. Transformer Multivariate Forecasting: Less is More? CoRR abs/2401.00230 (2024)
2024 arXiv
-
[71]
Xingxing Zhang, Furu Wei, and Ming Zhou. 2019. HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summa- rization. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019. Association for Computa...
2019
-
[72]
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference...
2021
-
[73]
Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, and Rong Jin. 2023. One Fits All: Power General Time Series Analysis by Pretrained LM. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023. A USING DIFFERENT PERCE...
2023
-
[74]
In 8th International Conference on Learning Representations, ICLR 2020
BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020 . OpenReview.net
2020
- [75]
-
[2017]
In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017
Densely Connected Convolutional Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 . IEEE Computer Society, 2261–2269
2017
-
[2020]
In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
Table Search Using a Deep Contextualized Language Model. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery, New York, NY, USA, 589–598
-
[2022]
In Proceedings of the Web Conference (WWW 2022)
StruBERT: Structure-aware BERT for Table Search and Matching. In Proceedings of the Web Conference (WWW 2022)
2022
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.