REVIEW 5 major objections 5 minor 1 cited by
TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read TempoGPT claims that quantizing time series into discrete tokens—the same representation as text—lets large language models actually reason about temporal data.
desk verdict Solid synthetic benchmark plus a plausible recipe for discrete time-series tokens, but the ablation conflates quantization with VQ-VAE pretraining and the evaluation never leaves the template distribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is quantization encoding plus a shared embedding layer. Quantization encoding uses a weight-shared 1-D convolutional encoder with patching and channel independence to convert a time series $X \in \mathbb{R}^{M \times N}$ into temporal embeddings, then maps each embedding to its nearest codeword in a fixed discrete codebook, producing temporal tokens $X_T \in \{\langle 0\rangle, \langle 1\rangle, \dots, \langle K\rangle\}$. These tokens are merged into the LLM's vocabulary ($V_2 = V_0 \cup V_1$) and the word embedding matrix is expanded accordingly ($W_2 = [W_0; W_1]$), so one shared embedding layer maps both text and time series to vectors. Its work is to give the two modalities an identical representation pattern, which the paper argues is what makes multi-modal alignment and logical reasoning possible; ablations removing either quantization or pre-training slow convergence and degrade reasoning accuracy.
What would settle it
Take a time-series reasoning test whose questions and reasoning steps are authored by humans from a different domain (or from the same circuit but with template families held out), train an otherwise identical pair of models—one with continuous temporal embeddings and one with quantized temporal tokens—and compare conclusion accuracy and logical reasoning accuracy. If the quantization advantage shrinks to zero or reverses when the data no longer follows the training templates, the claim that discrete temporal tokens generically improve reasoning is refuted.
Extended reading notes
Core claim
TempoGPT's central claim is that the reason multi-modal time-series language models (TLMs) underperform at complex reasoning is a representation mismatch: continuous temporal embeddings follow a different pattern than discrete textual tokens, so the alignment layer cannot meaningfully connect the two. The paper's remedy is quantization encoding: a frozen VQ-VAE-style encoder turns each patch of each variable into a discrete token from a temporal codebook, and these tokens are added to the LLM's vocabulary so a shared, expanded embedding layer maps temporal and textual tokens into one space. Because the whole vocabulary is discrete, the model adjusts a finite set of embedding vectors to align the modalities, rather than trying to align one continuous stream with one discrete stream. Experiments on five constructed tasks—trend analysis, trend forecast, fault judgment, fault diagnosis, fault analysis—show that TempoGPT reaches 82.4–83.3% average conclusion accuracy against 52.7–71.4% for continuous-embedding TLMs, and that quantization alone more than doubles the performance of some baselines on reasoning tasks. The paper also introduces deception rate (correct conclusion via faulty logic) as a metric and reports that TempoGPT lowers it from 13.3–16.0% to 2.7%, evidence that the model reasons from the actual temporal values instead of guessing.
Load-bearing premise
The load-bearing premise is that the constructed linear-circuit benchmark, whose questions come from the same rule-based templates that generate the training labels, is a valid measure of complex time-series reasoning; the paper tests no real-world or independently authored reasoning questions.
Editorial extensions
If this is right
- Any TLM built on a frozen encoder and LoRA-tuned LLM can adopt the recipe: train a VQ-VAE on the target domain, quantize patches into codebook tokens, and expand the embedding layer.
- Quantization is a positive component for reasoning even when it loses some temporal detail, because consistent representation with text matters more than exact numeric reconstruction.
- Larger base models benefit more under TempoGPT, reversing the trend observed for continuous-embedding TLMs, where larger models perform worse on reasoning.
- Pre-training with temporal-text pairs is still necessary: without it, trend analysis and forecasting accuracy collapse, showing that alignment and temporal perception are complementary.
- Fine-tuning method matters: full fine-tuning of a small model (GPT-2, 125M) approaches the performance of 3B-parameter models, so the quantization mechanism may allow lightweight models to compete.
Reading between the lines
- If the representation-mismatch story generalizes, similar reasoning gains should appear in other modalities where continuous features meet discrete text—for instance, replacing continuous image-patch embeddings with discrete codebook tokens in vision-language models; the paper does not test this.
- The white-box data-generation recipe is domain-portable (thermal, mechanical, chemical systems), but the paper's own template-bound evaluation suggests the gains may shrink on out-of-template or real-world reasoning tasks until the labels are diversified beyond rule-based generation.
- Codebook size and patch length are treated as fixed choices; tuning them as alignment hyperparameters, rather than reconstruction hyperparameters, is a natural next step the paper leaves open.
- Vocab growth is the hidden cost of this design—each variable adds tokens to the LLM vocabulary—so the approach trades embedding-matrix size for alignment; high-dimensional sensor arrays may hit practical limits.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes TempoGPT, a time-series language model (TLM) that quantizes temporal embeddings into discrete tokens via a VQ-VAE-style codebook and processes those tokens together with text tokens through a shared embedding layer, aiming for a consistent representation pattern across modalities. The authors also introduce a data-construction pipeline that generates multi-modal linear-circuit time-series data with rule-based, human-in-the-loop, and ChatGPT-diversified templates, yielding five reasoning tasks: trend analysis, trend forecast, fault judgement, fault diagnosis, and fault analysis. The model is trained in two stages (pre-training for alignment, fine-tuning for instruction following) and evaluated with three metrics: Conclusion Accuracy (CA), Logical Reasoning Accuracy (LRA), and Deception Rate (DR). The experiments compare TempoGPT against textual-prompt and continuous-embedding baselines on the constructed dataset, and the authors report state-of-the-art performance, improved logical reasoning, and a positive contribution from quantization via an ablation study.
Significance. If the causal role of quantization were cleanly established, the work would be a useful contribution to the TLM community: the shared-embedding, discrete-token design is simple, inexpensive to train at GPT-2 scale, and the released code/data facilitate reproduction. The central idea---that aligning the representation pattern of time series with that of text can improve multi-modal reasoning---is plausible and worth testing. However, the current evidence is weakened by a confounded ablation (Table 3), a fully self-constructed evaluation setting, and small manual evaluation sets (75 samples) without inter-annotator agreement, so the strength of the claimed causal effect is not yet supported.
major comments (5)
- [§5.4, Table 3] The ablation labeled '+Quantization' changes more than quantization: relative to 'Original', it adds VQ-VAE pretraining, freezes the encoder, and replaces the continuous embedding with codebook tokens. Thus the observed gains (e.g., GPT-2 Linear Fault Analysis CA rising from 42.0 to 75.3) could be caused by the pretraining or frozen-encoder regularization rather than by discrete tokenization. To support the paper's central causal claim, add a control that trains the continuous-embedding baseline with the same VQ-VAE-pretrained, frozen encoder but no codebook, or explicitly demonstrate that 'TempoGPT w/o quantization' in Table 8 is exactly that control.
- [§3, §5.1.1, §6] All experiments are conducted on a test set generated from the same rule-based templates used to create training labels, so the 'state-of-the-art in complex time series reasoning tasks' claim is overgeneralized. The paper should either evaluate transfer on an independent or established time-series reasoning benchmark (e.g., TimeSeriesExam, or zero-shot tasks from prior work) or substantially qualify the conclusion as applying only to the constructed electrical-circuit dataset.
- [§5.3.2, Table 2] The LRA and DR metrics, which carry much of the logical-reasoning claim, are based on manual grading of only 75 samples total (25 per task), and no inter-annotator agreement is reported. With 25 samples, the difference between DR 0% and DR 4% is a single response, and the reported advantage of TempoGPT over the continuous baselines is not accompanied by confidence intervals or a statistical test. The authors should report agreement statistics and, if possible, a larger or at least interval-estimated evaluation.
- [§5.2, Table 1] The text states that 'regardless of the base model employed, the performance of TempoGPT surpasses TLMs based on continuous embedding,' but Table 1 shows GPT-2 Linear achieving 69.7 on Fault Judgement versus TempoGPT (GPT-2) at 65.8, and Table 3 shows the same 'Original' continuous model at 69.7 versus '+Quantization' at 61.8. This overclaim should be corrected, and the discussion should acknowledge that quantization can hurt on some tasks and model sizes.
- [§4.1.1, §4.2, Appendix B.2] The proposed method depends on a 'predefined encoder and temporal codebook,' but the paper defers VQ-VAE details to reference [23] and provides no codebook size, codebook dimension, VQ-VAE training data, or hyperparameters. Since the quantization mechanism is the central contribution, these details are needed for reproduction and for assessing the sensitivity of the results to codebook capacity.
minor comments (5)
- [Abstract and §6] The word 'Specially' is used in the abstract and conclusion where 'Specifically' is intended; this should be corrected.
- [Table 8 and §5.4] Table 8's caption reads 'The detail concept about CA ...'; the phrase 'detailed results' would be clearer. Also, Figure 6's caption should explicitly define what is removed in 'TempoGPT w/o quantization' to avoid ambiguity with the Table 3 design.
- [Table 5] The claim that TempoGPT uniformly surpasses continuous-embedding methods is again contradicted by Table 5, e.g., LLaMA-3.2-1B (Attention) Trend Analysis at 99.3 versus TempoGPT (LLaMA-3.2-1B) at 95.8; the discussion should acknowledge such exceptions.
- [§5.1.2] The CA metric is computed by string matching against the reference conclusion; given that the conclusions are generated from templates, this is a reasonable choice, but the paper should state whether the matching tolerates paraphrases that are semantically equivalent.
- [References] References [48] and [49] point to model card URLs rather than stable archival versions; adding versioned citations or DOIs would improve reproducibility.
Circularity Check
No circularity: the paper's empirical claims are not equivalent to their inputs by construction, and no load-bearing self-citation chain is present.
full rationale
Walking the paper's claimed derivation chain, I find no step in which a prediction or first-principles result reduces to its own inputs by construction. The central architectural claim is that quantizing temporal embeddings into discrete tokens and routing both temporal and textual tokens through a shared embedding layer produces a consistent representation pattern (Section 4.1). This is a design definition, not a derived result: the paper does not claim to prove that consistency causes improved reasoning; it tests that claim empirically in Table 3 and Figure 6. The VQ-VAE codebook is trained separately on time series data (Section 4.1.1, referencing external work [23]) and then frozen, so the temporal tokens are not fitted to the reasoning labels. The reasoning labels themselves are generated by rule-based templates plus human-in-the-loop and ChatGPT diversification (Section 3), and the model is fine-tuned to predict them; the evaluation then measures held-out samples from the same task family. That is a self-referential benchmark validity limitation, not a circular reduction: the test answers are not contained in the training inputs by construction, and the model must still compute the underlying circuit quantities. Similarly, the ablation in Table 3 compares continuous-embedding baselines before and after adding quantization, and the reader is correct that this toggles VQ-VAE pretraining and frozen-encoder regularization along with quantization; however, that is a controlled-variable confound, not a case where the conclusion is logically equivalent to the premise. The paper also does not rely on load-bearing self-citations: references [22], [23], and [42] are external prior works, and no uniqueness theorem or ansatz is imported from the authors' own earlier papers. The manual LRA/DR scoring is performed by the authors (Section 5.1.2 and Appendix B.1), which is a subjectivity concern, but it does not make any derivation circular. Overall, the paper's empirical contributions are self-contained experiments rather than disguised restatements of their inputs, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Temporal codebook (size and training details)
assumptions (4)
- domain assumption The linear circuit simulation correctly models the physical relationships between voltage sources, loads, and current.
- domain assumption Rule-based and human-in-the-loop templates generate correct ground-truth reasoning labels.
- domain assumption A pretrained VQ-VAE provides semantically meaningful discrete temporal tokens.
- domain assumption Manual evaluation of LRA and DR is objective and reproducible.
Cite this review
Pith. "Pith review of TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding." pith.science (2026). https://pith.science/paper/WPRSI5XT
@misc{pith2026250107335,
author = {Pith},
title = {Pith review of: TempoGPT: Enhancing Time Series Reasoning via Quantizing Embedding},
year = {2026},
howpublished = {\url{https://pith.science/paper/WPRSI5XT}},
note = {Machine review of arXiv:2501.07335}
}
read the original abstract
Multi-modal language model has made advanced progress in vision and audio, but still faces significant challenges in dealing with complex reasoning tasks in the time series domain. The reasons are twofold. First, labels for multi-modal time series data are coarse and devoid of analysis or reasoning processes. Training with these data cannot improve the model's reasoning capabilities. Second, due to the lack of precise tokenization in processing time series, the representation patterns for temporal and textual information are inconsistent, which hampers the effectiveness of multi-modal alignment. To address these challenges, we propose a multi-modal time series data construction approach and a multi-modal time series language model (TLM), TempoGPT. Specially, we construct multi-modal data for complex reasoning tasks by analyzing the variable-system relationships within a white-box system. Additionally, proposed TempoGPT achieves consistent representation between temporal and textual information by quantizing temporal embeddings, where temporal embeddings are quantized into a series of discrete tokens using a predefined codebook; subsequently, a shared embedding layer processes both temporal and textual tokens. Extensive experiments demonstrate that TempoGPT accurately perceives temporal information, logically infers conclusions, and achieves state-of-the-art in the constructed complex time series reasoning tasks. Moreover, we quantitatively demonstrate the effectiveness of quantizing temporal embeddings in enhancing multi-modal alignment and the reasoning capabilities of TLMs. Code and data are available at https://github.com/zhanghaochuan20/TempoGPT.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
The authors organize LLM-based time series reasoning into three exclusive topologies (direct, chain, branch) crossed with four objectives, and use them to label 125 papers, benchmarks, and resources.
Reference graph
Works this paper leans on
-
[23]
Mingyue Cheng, Yiheng Chen, Qi Liu, Zhiding Liu, and Yucong Luo. 2024. Advancing Time Series Classification with Multimodal Language Modeling. arXiv:2403.12371 (2024)
arXiv 2024
-
[1]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2024. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. In Proceedings of the 12th International Conference on Learning Representations
work page 2024
-
[2]
Abdul Fatir Ansari, Lorenzo Stella, Ali Caner Türkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Rangapuram, Sebas- tian Pineda-Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C. Maddix, Michael W. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke- Schneider, and Yuyang Wang. 2024. Chronos: Learning the Langua...
arXiv 2024
-
[3]
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang. 2023. Pengi: An Audio Language Model for Audio Tasks. In Proceedings of the 37th International Conference on Neural Information Processing Systems
work page 2023
-
[4]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binko...
work page 2022
-
[5]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. In Proceedings of the 37th International Conference on Neural Information Processing Systems
work page 2023
-
[6]
Ane Blázquez-García, Angel Conde, Usue Mori, and José Antonio Lozano. 2022. A Review on Outlier/Anomaly Detection in Time Series Data. ACM computing surveys (2022), 56:1–56:33
work page 2022
-
[7]
Hatice Vildan Dudukcu, Murat Taskiran, Zehra Gülru Çam Taskiran, and Tulay Yildirim. 2023. Temporal Convolutional Networks with RNN approach for chaotic time series prediction. Applied soft computing (2023), 109945
work page 2023
Show all 56 references
-
[8]
Shun Liu, Kexin Wu, Chufeng Jiang, Bin Huang, and Danqing Ma. 2024. Financial Time-Series Forecasting: Towards Synergizing Performance And Interpretability Within a Hybrid Machine Learning Approach. arXiv:2401.00534 (2024)
2024 arXiv
-
[9]
Shervin Minaee, Tomás Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large Language Models: A Survey. arXiv:2402.06196 (2024)
2024 arXiv
-
[10]
Yong Liu, Tengge Hu, Haoran Zhang, Haixu Wu, Shiyu Wang, Lintao Ma, and Mingsheng Long. 2024. iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. In Proceedings of the 12th International Conference on Learning Representations
2024
-
[11]
Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. 2023. Are Transformers Effective for Time Series Forecasting?. InProceedings of the 37th AAAI Conference on Artificial Intelligence. 11121–11128
2023
-
[12]
Merrill, Mingtian Tan, Vinayak Gupta, Thomas Hartvigsen, and Tim Althoff
Mike A. Merrill, Mingtian Tan, Vinayak Gupta, Thomas Hartvigsen, and Tim Althoff. 2024. Language Models Still Struggle to Zero-shot Reason about Time Series. In Findings of the Association for Computational Linguistics: ACL 2024 . 3512–3533
2024
-
[14]
Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, and Rong Jin. 2023. One Fits All: Power General Time Series Analysis by Pretrained LM. In Proceedings of the 37th International Conference on Neural Information Processing Systems
2023
-
[15]
Yifu Cai, Arjun Choudhry, Mononito Goswami, and Artur Dubrawski. 2024. TimeSeriesExam: A time series understanding exam. arXiv:2410.14752 (2024)
2024 arXiv
-
[16]
Chengsen Wang, Qi Qi, Jingyu Wang, Haifeng Sun, Zirui Zhuang, Jinming Wu, Lei Zhang, and Jianxin Liao. 2024. ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual Data
2024
-
[17]
Zhe Xie, Zeyan Li, Xiao He, Longlong Xu, Xidao Wen, Tieying Zhang, Jianjun Chen, Rui Shi, and Dan Pei. 2024. ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning. arXiv:2412.03104 (2024)
2024
-
[18]
Jun Li, Che Liu, Sibo Cheng, Rossella Arcucci, and Shenda Hong. 2023. Frozen Language Model Helps ECG Zero-Shot Learning. In Medical Imaging with Deep Learning: MIDL 2023. 402–415
2023
-
[19]
Zhonghang Li, Lianghao Xia, Jiabin Tang, Yong Xu, Lei Shi, Long Xia, Dawei Yin, and Chao Huang. 2024. UrbanGPT: Spatio-Temporal Large Language Models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 5351–5362
2024
-
[20]
Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. In Proceedings of the 1st International Conference on Learning Representations
2013
-
[21]
Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam
Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In Proceedings of the 11th International Conference on Learning Representations
2023
-
[22]
Hyunseung Chung, Jiho Kim, Joon-Myoung Kwon, Ki-Hyun Jeon, Min Sung Lee, and Edward Choi. 2023. Text-to-ECG: 12-Lead Electrocardiogram Synthesis Con- ditioned on Clinical Text Reports. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023 . 1–5
2023
-
[24]
Gupta, and Jingbo Shang
Xiyuan Zhang, Ranak Roy Chowdhury, Rajesh K. Gupta, and Jingbo Shang. 2024. Large Language Models for Time Series: A Survey. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence . 8335–8343
2024
-
[25]
Nate Gruver, Marc Finzi, Shikai Qiu, and Andrew Gordon Wilson. 2023. Large Language Models Are Zero-Shot Time Series Forecasters. In Proceedings of the 37th International Conference on Neural Information Processing Systems
2023
-
[26]
Kashif Rasul, Arjun Ashok, Andrew Robert Williams, Arian Khorasani, George Adamopoulos, Rishika Bhagwatkar, Marin Biloš, Hena Ghonia, Nadhir Hassen, Anderson Schneider, et al. 2023. Lag-llama: Towards foundation models for time series forecasting. In R0-FoMo: Robustness of Few...
2023
-
[27]
Arik, Tomas Pfister, Yixiang Zheng, Wen Ye, and Yan Liu
Defu Cao, Furong Jia, Sercan Ö. Arik, Tomas Pfister, Yixiang Zheng, Wen Ye, and Yan Liu. 2024. TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting. In Proceedings of the 12th International Conference on Learning Representations
2024
-
[28]
Chenxi Sun, Hongyan Li, Yaliang Li, and Shenda Hong. 2024. TEST: Text Proto- type Aligned Embedding to Activate LLM’s Ability for Time Series. InProceedings of the 12th International Conference on Learning Representations
2024
-
[29]
Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen
Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y. Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen. 2024. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. In Proceedings of the 12th International Conferenc...
2024
-
[30]
Yong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang, and Mingsheng Long
-
[31]
Merrill, Vinayak Gupta, Tim Althoff, and Thomas Hartvigsen
Mingtian Tan, Mike A. Merrill, Vinayak Gupta, Tim Althoff, and Thomas Hartvigsen. 2024. Are Language Models Actually Useful for Time Series Fore- casting? arXiv:2406.16964 (2024)
2024 arXiv
-
[32]
Liangwei Nathan Zheng, Chang George Dong, Wei Emma Zhang, Lin Yue, Miao Xu, Olaf Maennel, and Weitong Chen. 2024. Revisited Large Language Model for Time Series Analysis through Modality Alignment. arXiv:2410.12326 (2024)
2024 arXiv
-
[33]
Yifu Cai, Arvind Srinivasan, Mononito Goswami, Arjun Choudhry, and Artur Dubrawski. 2024. JoLT: Jointly Learned Representations of Language and Time- Series for Clinical Time-Series Interpretation (Student Abstract). In Proceedings of the 38th AAAI Conference on Artificial Int...
2024
-
[34]
Mingyu Jin, Hua Tang, Chong Zhang, Qinkai Yu, Chengzhi Liu, Suiyuan Zhu, Yongfeng Zhang, and Mengnan Du. 2024. Time Series Forecasting with LLMs: Understanding and Enhancing Model Capabilities. arXiv:2402.10835 (2024)
2024 arXiv
-
[35]
Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montser- rat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng. 2023. Large Language Models as General Pattern Machines. In Conference on Robot Learning, CoRL 2023. 2498–2518
2023
-
[36]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings of the 40th International Conference on Machine Learning. 19730–19742
2023
-
[37]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings...
2021
-
[38]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022. BLIP: Boot- strapping Language-Image Pre-training for Unified Vision-Language Under- standing and Generation. In Proceedings of the 39th International Conference on Machine Learning. 12888–12900
2022
-
[39]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. In Conference on Computer Vision and Pattern Recognition, CVPR 2024. 26286–26296
2024
-
[40]
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. ImageBind One Embedding Space to Bind Them All. In Conference on Computer Vision and Pattern Recognition, CVPR 2023. 15180–15190
2023
-
[41]
Winnie Chow, Lauren Gardiner, Haraldur T Hallgrímsson, Maxwell A Xu, and Shirley You Ren. 2024. Towards time series reasoning with llms.arXiv:2409.11376 (2024)
2024 arXiv
-
[42]
Yiqun Duan, Jinzhao Zhou, Zhen Wang, Yu-Kai Wang, and Chin-Teng Lin. 2023. DeWave: discrete EEG waves encoding for brain dynamics to text translation. In Proceedings of the 37th International Conference on Neural Information Processing Systems. 9907–9918
2023
-
[43]
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. In Proceedings of the 35th AAAI Conference on Artificial Intelligence. 11106–11115. TempoGPT: ...
2021
-
[44]
Tianyu Wu, Shizhu He, Jingping Liu, Siqi Sun, Kang Liu, Qing-Long Han, and Yang Tang. 2023. A brief overview of ChatGPT: The history, status quo and potential future development. IEEE/CAA Journal of Automatica Sinica (2023), 1122–1136
2023
-
[45]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elic- its Reasoning in Large Language Models. In Proceedings of the 36th International Conference on Neural Information Proces...
2022
-
[46]
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017. Neural Dis- crete Representation Learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems . 6306–6315
2017
-
[47]
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. (2019)
2019
-
[48]
Meta Llama-3.2-1B
2024. Meta Llama-3.2-1B. https://huggingface.co/meta-llama/Llama-3.2-1B
2024
-
[49]
Meta Llama-3.2-3B
2024. Meta Llama-3.2-3B. https://huggingface.co/meta-llama/Llama-3.2-3B
2024
-
[50]
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. 2024. TinyLlama: An Open-Source Small Language Model
2024
-
[51]
Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al . 2023. Phi-2: The surprising power of small language models. Microsoft Research Blog (2023), 3
2023
-
[52]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the 10th International Conference on Learning Representations
2022
-
[53]
Openai GPT-3.5
2022. Openai GPT-3.5. https://chatgpt.com/auth/login
2022
-
[54]
Openai GPT-4
2023. Openai GPT-4. https://openai.com/index/gpt-4/
2023
-
[55]
Furong Jia, Kevin Wang, Yixiang Zheng, Defu Cao, and Yan Liu. 2024. GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting. In Proceedings of the 38th AAAI Conference on Artificial Intelligence . 23343–23351
2024
-
[56]
Pengyu Chen. 2021. Effects of the entropy weight on TOPSIS. Expert Systems with Applications (2021), 114186. Haochuan Zhang, Chunhua Yang, Jie Han, Liyang Qin, and Xiaoli Wang A DATA DETAILS Due to space limitations, we present the data from the fine-tuning stage in more detai...
2021
-
[2024]
arXiv:2402.02370 (2024)
AutoTimes: Autoregressive Time Series Forecasters via Large Language Models. arXiv:2402.02370 (2024)
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.