REVIEW 2 major objections 5 minor 124 references
Even the best LLM agent fully solves only 46% of jointly feasible trip-planning tasks; unstated traveler needs are the universal bottleneck.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 15:11 UTC pith:WGQIRQ4X
load-bearing objection Solid, reproducible joint-feasibility travel-agent benchmark; the 46% ceiling and D1 bottleneck are real inside the sandbox, but “unstated needs” is oversold relative to what D1 actually measures. the 2 major comments →
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On TREK’s feasible tasks, even the strongest evaluated agent produces a plan that is jointly constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and responsive to unstated persona needs on only 46.2% of cases (median 6.6%, floor 0.0% across 15 agents). Implicit-need satisfaction is the universal bottleneck: it remains the last wall for the frontier model and a top failure for every agent tested.
What carries the argument
Feasible itinerary synthesis, scored by the task-perfect rate: a single plan must pass every applicable one of nine deterministic dimensions (and hard gates) at once. The evaluator is rule-based with no LLM judge; gold references demonstrably score 1.0; 267 tasks are typed-infeasible (route/entity/budget) so correct refusal is first-class.
Load-bearing premise
That success on a synthetic, internally consistent travel sandbox—with unstated needs reduced to facility checklist matches against stated persona labels—certifies the real deployable skill of building executable, traveler-responsive itineraries.
What would settle it
If a frontier tool-using agent, under the same TREK harness and scorer, reaches near-100% task-perfect on the 533 feasible tasks—including implicit-need cells—and human raters judge those plans as realistic and followable at rates matching the gold, the claimed capability gap collapses.
If this is right
- Joint feasibility, not single-axis tool success, becomes the right unit for certifying travel (and similar multi-constraint) agents.
- Implicit persona needs must be treated as a first-class, hard requirement rather than soft preference flavor text.
- Typed refusal of impossible requests (route/entity/budget) is a scored skill, not an afterthought.
- Deterministic no-judge scorers paired with achievable gold make remaining gaps attributable to agents, not rubrics.
- More deliberation or token spend does not automatically buy more joint feasibility under strict tool schemas.
Where Pith is reading between the lines
- The same joint-feasibility template—synthetic consistent world, typed impossibility, deterministic conjunction scoring—could stress-test agents in logistics, clinical scheduling, or multi-leg procurement.
- If unstated needs stay the rising bottleneck as models improve elsewhere, progress may require explicit preference-to-resource mapping modules rather than longer chain-of-thought alone.
- A hidden-persona split (infer the traveler from indirect language, score the same facility checklist) would test whether the current D1 gap understates the real personalization problem.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TREK is a benchmark for joint feasible itinerary synthesis by tool-using LLM agents. It provides 800 multi-constraint travel tasks (533 feasible, 267 typed-infeasible) over a synthetic, internally consistent KB of 212,530 records across 375 cities and 13 personas, accessed via validated RESTful APIs. Scoring is fully deterministic across nine dimensions with no LLM judge; each task ships a human-verified gold that scores 1.0 under the same evaluator. Evaluating 15 agents, the authors report a top task-perfect rate of 46.2% on solvable tasks (median 6.6%, floor 0.0%), with implicit-need satisfaction (D1) as the universal bottleneck and B3 as a planner/non-planner watershed; reasoning and extra tokens do not clearly help. The release includes dataset, sandbox, evaluator, and agent code.
Significance. If the measurements hold under the stated operationalization, TREK is a meaningful advance over TravelPlanner, ChinaTravel, and concurrent travel benchmarks by coupling (i) a purely rule-based joint-feasibility scorer, (ii) a gold proven to hit the ceiling on all 800 tasks, (iii) typed route/entity/budget infeasibility as scored behavior, and (iv) a production-style tool sandbox with an efficiency axis. The 15-model study is carefully instrumented (per-dimension failure tables, aggregation robustness, truncation diagnostics, shared B3 travel-time helper). These are the right ingredients for a reproducible agent benchmark: bit-reproducible scores, an attainable ceiling, and falsifiable per-dimension failure modes. The work is significant for tool-agent evaluation even if the strongest interpretive claim about “unstated needs” needs tightening.
major comments (2)
- [Abstract; §1; §4.2 D1; §5.3; App. E.3] The headline scientific claim that “satisfying travelers’ unstated persona needs” is the universal bottleneck overreaches what D1 measures. Queries name persona keywords explicitly (e.g., “for elderly travelers”); the system prompt (§D.4) further instructs agents that these labels are requirements, to derive amenities, and to use amenity/facility filters; D1 then scores deterministic facility set-intersection (plus fixed star/rating tests) against KB fields (§4.2, App. D.3, A.10). Appendix E.3 correctly concedes D1 does not test persona inference from indirect language. The residual ~53.7% D1 failure for GPT-5.6 is still important—it is multi-cell stated-persona→facility booking under joint constraints—but abstract, intro, RQ2, and conclusion should rename and reframe D1 (e.g., persona-conditioned facility grounding) so the bottleneck claim matches the instrument. Without that edit, the
- [§5.1; §5.3; App. D.4 system prompt] Relatedly, the agent harness scaffolds the very capability D1 is said to isolate: §D.4 tells models which traveler types imply concrete booking needs and that plans are scored on matching amenities. That makes the frontier D1 failure more striking as a grounding/execution gap, but it weakens any claim that agents fail at discovering partly spoken needs. Please either (a) report an ablation with a minimal prompt that does not enumerate the persona→facility mapping recipe, or (b) explicitly scope the finding as failure under a scaffold that already reveals the mapping problem. Option (b) is acceptable if the framing in major comment 1 is fixed; option (a) would strengthen the paper.
minor comments (5)
- [§2.1; Table 1] Table 1 and §2 positioning vs. ChinaTravel/TravelBench/TravelEval is clear and useful; consider adding one sentence on whether any concurrent benchmark has since added typed infeasibility or achievable gold, to keep the “to our knowledge” claim durable.
- [§4.2; App. D.3] Free scoring parameters (D3 β=4, B2 60-minute dwell cap, B3 surface/air constants, efficiency γ and token budget B) are documented but not sensitivity-tested beyond the task-perfect 0.95 check. A short appendix sweep on β and the B2 cap would reassure readers that rankings are not knife-edge.
- [Figure 4; Table 15] Figure 4 heatmap and Table 15 are excellent; ensure color/print accessibility and that “top-two failure” is defined in the caption (rank by failure rate among applicable dimensions).
- [§6; Appendix F] Appendix F already handles synthetic-world, single-run, and Claude geo-block limits well. Move a one-sentence external-validity caveat into the main conclusion so readers who skip the appendix see the scope bound next to the 46.2% headline.
- [Abstract; §3.2] Minor consistency: abstract says needs are “only partly stated” while templates inject explicit persona keywords—align wording with the revised D1 framing.
Circularity Check
No significant circularity: TREK is a transparent benchmark whose gold/labels are perfect by construction by design, while the headline agent scores are independent empirical measurements.
full rationale
TREK does not claim a first-principles derivation or a fitted natural-law prediction. Its load-bearing empirical claim—that frontier agents reach only ~46.2% task-perfect on solvable tasks, with D1 as the residual bottleneck—is an external measurement of third-party models under a fixed harness and a deterministic scorer. Feasibility labels and gold itineraries scoring 1.0 are openly engineered to be correct/perfect against that same scorer (§3.3–3.5, §4.1), which is standard achievable-ceiling benchmark design, not a hidden reduction of a predicted quantity to its fit. Budget bands, typed infeasibility, B3’s shared travel-time model, and D1’s facility set-intersection are author-specified world rules disclosed in the paper; agents are scored against them, not used to redefine them. There is no self-citation uniqueness theorem, no parameter fit renamed as prediction, and no ansatz smuggled in via overlapping-author prior work that forces the result. Framing tension around calling D1 “unstated” needs (personas are named in queries) is a construct-validity issue, not circularity of the derivation chain. Score 0.
Axiom & Free-Parameter Ledger
free parameters (6)
- Budget slack bands (tight u∈[1.02,1.10], loose u∈[1.35,1.80]; infeasible u∈[0.55,0.92]) =
means/ranges reported; near-even 267 tight / 266 loose
- D3 overspend penalty β=4 =
β=4
- B2 dwell-time cap of 60 minutes =
60 minutes
- B3 travel-time model constants =
R=6371 km; surface/air constants as in Eq. (5)
- Efficiency penalties (γ=10 per surplus call, C_max=15, token budget B) =
γ=150/C_max=10; cap 15 billable calls
- Geometric-mean floor ε for category product =
small ε (unspecified numeric)
axioms (6)
- domain assumption A synthetic, internally consistent KB is a stronger basis for certifying executability than a scraped drifting corpus.
- ad hoc to paper Unstated persona needs are adequately operationalized as deterministic nonempty intersection of fixed facility sets (plus fixed star/rating tests for luxury/foodie) on booked KB resources.
- ad hoc to paper Only logically impossible persona combinations (party-size contradictions; fast-paced budget vs luxury) are excluded; other co-occurrence is jointly satisfiable.
- domain assumption A deployable plan requires simultaneous satisfaction of constraint, truthfulness, executability, and correct refusal—aggregated conjunctively (task-perfect) and via geometric mean of categories.
- domain assumption Identifier-level KB match (not mutable price/rating equality) is the right hallucination test; attribute fidelity is deferred to budget/other checks.
- domain assumption Standard tool-agent evaluation practices (temperature 0, fixed harness, provider token counts as cost proxy) suffice for comparative claims in RQ1–RQ3.
invented entities (4)
-
TREK task suite + synthetic multi-domain travel KB (212,530 records, 375 cities)
no independent evidence
-
Thirteen traveler personas with fixed facility/amenity/service requirement tables
independent evidence
-
Nine-dimension deterministic evaluator and task-perfect metric with hard gates
no independent evidence
-
Typed infeasibility labels (entity/route/budget) as first-class scored behavior (D4)
no independent evidence
read the original abstract
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must be physically traversable, the total must clear a budget, and the plan must serve a traveler whose needs are only partly stated. Existing agent benchmarks reward these properties one at a time and grade the final output with soft or LLM-judged rubrics, which cannot certify that a returned plan is executable and are neither reproducible nor auditable. We introduce TREK (Travel Reasoning and Evaluation Kit), a benchmark for feasible itinerary synthesis: producing a single plan that is jointly constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and responsive to the traveler's unstated persona needs. TREK comprises 800 multi-constraint tasks - 533 feasible and 267 provably infeasible with typed route/entity/budget causes - over a synthetic, internally consistent knowledge base of 212,530 records across 375 cities and 13 personas, served through a production-style tool sandbox of validated RESTful APIs. Every task is scored by a fully deterministic, rule-based evaluator with no LLM judge and ships a human-verified gold reference that scores a perfect 1.0 under that same evaluator, so the ceiling is demonstrably achievable and every remaining gap is an agent limitation rather than scorer strictness. Evaluating 15 LLM agents across nine constraint dimensions, we find that even the strongest (GPT-5.6) produces a fully-feasible plan on only 46.2% of solvable tasks, with a median of 6.6% and a floor of 0.0%; satisfying travelers' unstated needs emerges as the universal bottleneck, unsolved even at the frontier. We release the dataset, tool sandbox, deterministic evaluator, and agent code as a fully reproducible benchmark.
Figures
Reference graph
Works this paper leans on
-
[1]
AAA. 2017. Foodie Travelers Are Embracing the Culinary Travel Trend. https://newsroom.aaa.com/2017/04/foodie-travelers-embracing-culinary- travel-trend/. AAA Newsroom travel release; Accessed 2026-02-09
2017
-
[2]
AAA Foundation for Traffic Safety. 2020. Older Drivers and Advanced Driver Assistance Systems. https://aaafoundation.org/older-drivers-and-advanced- driver-assistance-systems/. Research brief; Accessed 2026-02-09
2020
-
[3]
Amazon AGI. 2025. The Amazon Nova Family of Models: Technical Report and Model Card. arXiv:2506.12103 [cs.CL]
Pith/arXiv arXiv 2025
-
[4]
Amazon Web Services. 2025. Introducing Amazon Nova 2 Lite, a Fast, Cost- Effective Reasoning Model. https://aws.amazon.com/blogs/aws/introducing- amazon-nova-2-lite-a-fast-cost-effective-reasoning-model/. AWS News Blog release announcement
2025
-
[5]
Hannah Aston. 2024. Travel Trends: Luxury Culinary Tourism. https://discover.hotelbeds.com/resources/insight/2024-travel-trends-luxury- culinary-tourism. Accessed 2026-02-09
2024
-
[6]
Avis Rent A Car. n.d.. Child Safety Seats — Avis Car Rental. https://www.avis. com/en/products-and-services/products/childsafetyseats. Accessed 2026-02-09
2026
-
[7]
Srinivas Billa and Xiaonan Jing. 2025. TravelBench: Exploring LLM Perfor- mance in Low-Resource Domains.arXiv preprint arXiv:2510.02719(2025). arXiv:2510.02719 https://arxiv.org/abs/2510.02719
arXiv 2025
-
[8]
Jiangxi Network Broadcasting and Television Station. 2020. Publishing the 2020 Post-COVID Self-Drive Tourism Report: Consumer Trends Analysis. https://cn. chinadaily.com.cn/a/202010/09/WS5f8020fca3101e7ce972844c.html. Published on China Daily Chinese edition; Accessed 2026-02-09
2020
-
[9]
Soumyabrata Chaudhuri, Pranav Purkar, Ritwik Raghav, Shubhojit Mallick, Manish Gupta, Abhik Jana, and Shreya Ghosh. 2025. TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics. ar...
Pith/arXiv arXiv 2025
-
[10]
Junle Chen, Wei Chen, Yehong Xu, Zhengjun Huang, Yuqian Wu, Zhoujin Tian, Kai Wang, Lei Wang, and Xiaofang Zhou. 2026. Trip+: Benchmarking Agents in Personalized Interactive Travel Planning.arXiv preprint arXiv:2606.21169(2026). arXiv:2606.21169 https://arxiv.org/abs/2606.21169
Pith/arXiv arXiv 2026
-
[11]
Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, and Lei Chen. 2026. TravelEval: A Comprehensive Bench- marking Framework for Evaluating LLM-Powered Travel Planning Agents. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ACM. arXiv:2606.01046 doi:10.1145/3770855.3817533
Pith/arXiv arXiv 2026
-
[12]
Xiang Cheng, Yulan Hu, Xiangwen Zhang, Lu Xu, Lide Tan, Zheng Pan, Xin Li, and Yong Liu. 2026. Beyond Itinerary Planning—A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 29200–29251...
Pith/arXiv arXiv 2026
-
[13]
Xiang Cheng, Yulan Hu, Lulu Zheng, Zheng Pan, Xin Li, and Yong Liu. 2026. GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Plan- ning.arXiv preprint arXiv:2605.25200(2026). arXiv:2605.25200 https://arxiv.org/ abs/2605.25200
Pith/arXiv arXiv 2026
-
[14]
Comcast Business Community Editorial Team. [n. d.]. The “Why” Behind WiFi: How Deploying WiFi is Good for Business. https: //business.comcast.com/community/browse-all/details/the-why-behind- wifi-how-deploying-wifi-is-good-for-business. Accessed 2026-02-09
2026
-
[15]
Condor Ferries. 2025. Family Travel Statistics 2025. https://www.condorferries. co.uk/family-travel-statistics. Accessed 2026-02-09
2025
-
[16]
Condor Ferries. 2025. Pet Travel Statistics 2025. https://www.condorferries.co. uk/pet-travel-statistics. Accessed 2026-02-09
2025
-
[17]
Condor Ferries. 2025. Solo Travel Statistics 2025. https://www.condorferries.co. uk/solo-travel-statistics. Accessed 2026-02-09
2025
-
[18]
Denise Curtin. 2018. THIS hotel is now offering the world’s first ever ‘Instagram Butler’. https://her.ie/life/hotel-now-offering-worlds-first-ever-instagram- butler-370134. Her.ie lifestyle article; Accessed 2026-02-09
2018
-
[19]
CyberPublicity. 2025. Nightlife Enthusiasts. https://www.cyberpublicity.com/ programmatic-advertising/hobbies-passions/nightlife-enthusiasts/. Accessed 2026-02-09
2025
-
[20]
Renfei Dang, Zhening Li, Shujian Huang, and Jiajun Chen. 2025. The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning Models. arXiv:2505.16448 [cs.AI] https://arxiv.org/abs/2505.16448
arXiv 2025
-
[21]
Tomas de la Rosa, Sriram Gopalakrishnan, Alberto Pozanco, Zhen Zeng, and Daniel Borrajo. 2024. TRIP-PAL: Travel Planning with Guarantees by Com- bining Large Language Models and Automated Planners.arXiv preprint arXiv:2406.10196(2024). arXiv:2406.10196 https://arxiv.org/abs/2406.10196
Pith/arXiv arXiv 2024
-
[22]
DeepSeek-AI. 2025. DeepSeek-V3.2. https://huggingface.co/deepseek-ai/ DeepSeek-V3.2. Model card
2025
-
[23]
Bin Deng, Yizhe Feng, Zeming Liu, Qing Wei, Xiangrong Zhu, Shuai Chen, Yuanfang Guo, and Yunhong Wang. 2025. RETAIL: Towards Real-world Travel Planning for Large Language Models. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computa- tional Linguistics. arXiv:2508.15335 doi:10.18653/v1/2025.emnlp...
Pith/arXiv arXiv 2025
-
[24]
Ben Duhig. 2022. McKinsey reveals 7 emerging traveller archetypes. https://www.linkedin.com/posts/benduhig_7-emerging-traveller-archetypes- all-travel-activity-7211269404188176384-JYBG. LinkedIn post. Accessed 2026-02-09
2022
-
[25]
Enterprise Rent-A-Car. n.d.. What is the Enterprise Pet Policy? https://www.enterprise.com/en/car-rental-faqs/us-general/car-rental- pet-friendly-policy.html. Accessed 2026-02-09
2026
-
[26]
Expedia Group. 2022. New research: 97% of honeymoon plans were thwarted by the pandemic, leading to the rise of the ’mega-moon’. https://www.expediagroup.com/media/media-details/2022/New-research-97- of-honeymoon-plans-were-thwarted-by-the-pandemic-leading-to-the-rise- of-the-mega-moon/default.aspx. Accessed 2026-02-09
2022
-
[27]
Food Inspiration Magazine. 2023. Trendwatch – Food tourism. https://www. foodinspirationmagazine.com/39-food-tourism/trendwatch-food-tourism. Ac- cessed 2026-02-09
2023
-
[28]
Four Seasons Hotels and Resorts. 2024. Hotels With Michelin Star Restaurants. https://www.fourseasons.com/magazine/taste/michelin-starred- restaurants/. Accessed 2026-02-09
2024
-
[29]
Four Seasons Hotels and Resorts. 2024. A Meal to Remember: Luxury Dining with Four Seasons. https://www.fourseasons.com/magazine/taste/michelin- starred-restaurants/. Accessed 2026-02-09
2024
-
[30]
GLM-5 Team, Z.ai. 2026. GLM-5: From Vibe Coding to Agentic Engineering. arXiv:2602.15763 [cs.CL]
Pith/arXiv arXiv 2026
-
[31]
Google DeepMind. 2026. Gemma 4: Byte for Byte, the Most Capable Open Models. https://blog.google/innovation-and-ai/technology/developers-tools/ gemma-4/. Release announcement; Gemma 4 31B Dense
2026
-
[32]
Atharva Gundawar, Mudit Verma, Lin Guan, Karthik Valmeekam, Siddhant Bhambri, and Subbarao Kambhampati. 2024. Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning.arXiv preprint arXiv:2405.20625 (2024). arXiv:2405.20625 https://arxiv.org/abs/2405.20625
Pith/arXiv arXiv 2024
-
[33]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al . 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948(2025)
Pith/arXiv arXiv 2025
-
[34]
Yilun Hao, Yongchao Chen, Yang Zhang, and Chuchu Fan. 2025. Large Lan- guage Models Can Solve Real-World Planning Rigorously with Formal Verifica- tion Tools. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Tech- nologies (Volume 1: Long Papers). Association for C...
Pith/arXiv arXiv 2025
-
[35]
Hertz. 2024. Exploring Hertz’s Premium Emergency Roadside Assistance Ser- vice. https://www.hertz.com/us/en/blog/resources/hertz-premium-roadside- assistance. Accessed 2026-02-09
2024
-
[36]
Hertz. n.d.. Child Car Seats — Value-Added Services. https://www.hertz.com/us/ en/products-and-services/value-added-services/united-states/child-car-seats. Accessed 2026-02-09
2026
-
[37]
Angelina Holt. 2023. A Road Trip Guide for Your Hotel Guests: Accommo- dation Essentials. https://www.hoteliga.com/en/blog/a-road-trip-guide-for- your-hotel-guests-accommodation-essentials. Hoteliga blog article. Accessed 2026-02-09
2023
-
[38]
Tzu-Heng Huang, Harit Vishwakarma, and Frederic Sala. 2025. Time To Im- peach LLM-as-a-Judge: Programs are the Future of Evaluation.arXiv preprint arXiv:2506.10403(2025). arXiv:2506.10403 https://arxiv.org/abs/2506.10403
Pith/arXiv arXiv 2025
-
[39]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. InInternational Conference on Learning Rep- resentations. arXiv:2310.06770 https://openreview.net/forum?id=VTF8yNQM66
Pith/arXiv arXiv 2024
-
[40]
Kao, Maryam Fazel-Zarandi, and Yuandong Tian
Da Ju, Song Jiang, Andrew Cohen, Aaron Foss, Sasha Mitts, Arman Zharmagam- betov, Brandon Amos, Xian Li, Justine T. Kao, Maryam Fazel-Zarandi, and Yuandong Tian. 2024. To the Globe (TTG): Towards Language-Driven Guaran- teed Travel Planning. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. As...
Pith/arXiv arXiv 2024
-
[41]
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Kaya Stechly, Mudit Verma, Siddhant Bhambri, Lucas Paul Saldyt, and Anil B. Murthy. 2024. Position: LLMs Can’t Plan, But Can Help Planning in LLM-Modulo Frameworks. In Proceedings of the 41st International Conference on Machine Learning, Vol. 235. PMLR, 22895–22907. arXiv:2402.01817 https://proceedings.ml...
Pith/arXiv arXiv 2024
-
[42]
Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, and Shreya Ghosh. 2026. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. InFindings of the Association for Computational Linguistics: ACL 2026. Association for Computational Linguistics, 40269–40292. arXiv:2510.21329 doi:10.18653/v1/2026.findings-a...
arXiv 2026
-
[43]
Katapult. 2023. Is now the time to improve theme park food and dining experi- ences? https://www.katapult.co.uk/is-now-the-time-to-improve-theme-park- food-and-dining-experiences. Accessed 2026-02-09
2023
-
[44]
Kimi Team. 2026. Kimi K2.5: Visual Agentic Intelligence. arXiv:2602.02276 [cs.CL]
Pith/arXiv arXiv 2026
-
[45]
Cheeun Lee and Eun Hak Lee. 2024. Evaluation of Urban Nightlife Attractiveness for Millennials and Generation Z.Cities149 (2024), 104934. doi:10.1016/j.cities. 2024.104934 Accessed 2026-02-09
arXiv 2024
-
[46]
Lee, Gina Cui, Jungkeun Kim, Yuri Seo, and Hyunji Chon
Jacob C. Lee, Gina Cui, Jungkeun Kim, Yuri Seo, and Hyunji Chon. 2021. Photo Taking Paradox: Contrasting Effects of Photo Taking on Travel Satisfaction and Revisit Intention. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3978205. SSRN Electronic Journal. Accessed 2026-02-09
2021
-
[47]
Sarah Lee. 2025. Accessible Travel for All. https://www.numberanalytics.com/ blog/accessible-travel-guide. Accessed 2026-02-09
2025
-
[48]
Charlie Leocha. 2025. Top Hotel Amenities That Travelers Really Want When Choosing Accommodations. https://www.travelersunited.org/top-hotel- amenities-travelers-really-want/. Accessed 2026-02-04
2025
-
[49]
Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023. API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguisti...
-
[50]
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. 2025. AgentBench: Evaluating LLMs as Agents. arXiv:2308.03688 [cs.AI] https://arxiv.org/abs...
Pith/arXiv arXiv 2025
-
[51]
LOVU Travel. 2025. 7 Hotel Amenities That Make Couples Book Direct. https: //business.lovu.travel/7-hotel-amenities-that-make-couples-book-direct. Ac- cessed 2026-02-09
2025
-
[52]
Travel Maestro. 2012. No-man’s Land: the Rising Trend of Women-Only Hotel Floors. https://www.covingtontravel.com/2012/11/no-mans-land-the-rising- trend-of-women-only-hotel-floors/. Covington Travel blog article; Accessed 2026-02-09
2012
-
[53]
Meta AI. 2025. The Llama 4 Herd: The Beginning of a New Era of Na- tively Multimodal AI Innovation. https://ai.meta.com/blog/llama-4-multimodal- intelligence/. Release announcement; Llama 4 Scout and Maverick
2025
-
[54]
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2024. GAIA: A Benchmark for General AI Assistants. In International Conference on Learning Representations. arXiv:2311.12983 https: //openreview.net/forum?id=fibxvahvs3
Pith/arXiv arXiv 2024
-
[55]
Mistral AI. 2025. Introducing Mistral 3. https://mistral.ai/news/mistral-3/. Release announcement; introduces Mistral Large 3
2025
-
[56]
MMGY Global. 2022. Portrait of Travelers with Disabilities: Mobility & Ac- cessibility. https://www.mmgyglobal.com/news/portrait-of-travelers-with- disabilities/. MMGY Global news release; Accessed 2026-02-09
2022
-
[57]
Moonshot AI. 2025. Kimi K2 Thinking. https://huggingface.co/moonshotai/Kimi- K2-Thinking. Model card
2025
-
[58]
National Civil Rights Museum. n.d.. Plan Your Visit — National Civil Rights Museum. https://civilrightsmuseum.org/visit/. Accessed 2026-02-09
2026
-
[59]
National Park Service. 2025. Socioeconomic Monitoring Visitor Sur- veys. https://www.nps.gov/subjects/socialscience/socioeconomic-monitoring- visitor-surveys.htm. Accessed 2026-02-09
2025
-
[60]
Hang Ni, Fan Liu, Xinyu Ma, Lixin Su, Shuaiqiang Wang, Dawei Yin, Hui Xiong, and Hao Liu. 2025. TP-RAG: Benchmarking Retrieval-Augmented Large Lan- guage Model Agents for Spatiotemporal-Aware Travel Planning. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 12392–12418. ar...
arXiv 2025
-
[61]
nmu. 2022. How To Make Travel a Kid-Friendly Experience at Your Ho- tel. https://www.thesolutionsdesk.com/how-to-make-travel-a-kid-friendly- experience-at-your-hotel/. TheSolutionsDesk (Guest Supply blog). Accessed 2026-02-09
2022
-
[62]
OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexan- der Prokofiev, Alexander Wei, Allison Tam, Ally Bennett, Ananya Kumar, An- dre Saraiva, Andrea Vallone, Andrew Duberstein, Andrew Ko...
Pith/arXiv arXiv 2024
-
[63]
OpenAI. 2025. gpt-oss-120b & gpt-oss-20b Model Card. arXiv:2508.10925 [cs.CL]
Pith/arXiv arXiv 2025
-
[64]
OpenAI. 2026. GPT-5.6: Frontier Intelligence that Scales with Your Ambition. https://openai.com/index/gpt-5-6/. Model release; GPT-5.6 Sol (frontier tier) via the Responses API
2026
-
[65]
OurAirports. 2026. Open Airport Data. https://ourairports.com/data/. Public domain dataset containing worldwide airport information. Accessed 2026-02-04
2026
-
[66]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe
-
[67]
Srikant Panda, Amit Agarwal, and Hitesh Laxmichand Patel. 2025. AccessEval: Benchmarking Disability Bias in Large Language Models. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Associ- ation for Computational Linguistics. arXiv:2509.22703 doi:10.18653/v1/2025. emnlp-main.1653
arXiv 2025
-
[68]
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv:2307.16789 [cs.AI] https://arxiv.org/a...
Pith/arXiv arXiv 2023
-
[69]
Qwen Team, Alibaba. 2025. Qwen3-Next-80B-A3B-Instruct. https://huggingface. co/Qwen/Qwen3-Next-80B-A3B-Instruct. Model card
2025
-
[70]
Reddit. 2025. LPT: In touristy areas, tourist trap restaurants may have high online ratings from clueless tourists. https://www.reddit.com/r/LifeProTips/ comments/1kgu263/lpt_in_touristy_areas_tourist_trap_restaurants/. Accessed 2026-02-09
2025
-
[71]
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024. LaMP: When Large Language Models Meet Personalization. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics. arXiv:2304.11406 doi:10.18653/v1/2024.acl-long.399
Pith/arXiv arXiv 2024
-
[72]
Sawgrass Marketing. 2023. Seven Travel Personas You Need to Know for Niche Hospitality Marketing. https://www.sawgrassmktg.com/blog/seven-travel- personas-you-need-to-know-for-niche-hospitality-marketing. Accessed 2026- 02-09
2023
-
[73]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessí, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: language models can teach themselves to use tools. InProceedings TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning KDD ’27, August 2027, of the 37th Internation...
2023
-
[74]
Jie-Jing Shao, Bo-Wen Zhang, Xiao-Wen Yang, Baizhi Chen, Si-Yu Han, Jinghao Pang, Wen-Da Wei, Guohao Cai, Zhenhua Dong, Lan-Zhe Guo, and Yu-Feng Li
-
[75]
Zijian Shao, Jiancan Wu, Weijian Chen, and Xiang Wang. 2025. Personal Travel Solver: A Preference-Driven LLM-Solver System for Travel Planning. InPro- ceedings of the 63rd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers). Association for Computational Linguistics. doi:10.18653/v1/2025.acl-long.1339
-
[76]
Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. 2024. Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge.arXiv preprint arXiv:2406.07791(2024). arXiv:2406.07791 https: //arxiv.org/abs/2406.07791
arXiv 2024
-
[77]
SiteMinder. 2024. Hotel amenities: Examples and ideas list. https: //www.siteminder.com/r/trends-advice/hotel-management-tips-ideas/hotel- amenities-property-services/. Accessed 2026-02-09
2024
-
[78]
Matylda Siwek, Anna Kolasińska, Krzysztof Wrześniewski, and Mag- dalena Zmuda Palka. 2022. Services and Amenities Offered by City Hotels within Family Tourism as One of the Factors Guaranteeing Satisfactory Leisure Time.International Journal of Environmental Research and Public Health19, 14 (2022), 8321. doi:10.3390/ijerph19148321 Accessed 2026-02-09
-
[79]
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024. Large Language Models are Inconsistent and Biased Evaluators.arXiv preprint arXiv:2405.01724 (2024). arXiv:2405.01724 https://arxiv.org/abs/2405.01724
Pith/arXiv arXiv 2024
-
[80]
Megan Sullivan. 2013. Pet-Friendly Hotels Prove Profitable. https:// lodgingmagazine.com/profiting-from-pets/. Accessed 2026-02-09
2013
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.