Pith. sign in

REVIEW 24 cited by

Challenges and Applications of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.10169 v1 pith:AD5S3JNV submitted 2023-07-19 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords applicationchallengesfieldlanguagelargemodelsalreadyapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) went from non-existent to ubiquitous in the machine learning discourse within a few years. Due to the fast pace of the field, it is difficult to identify the remaining challenges and already fruitful application areas. In this paper, we aim to establish a systematic set of open problems and application successes so that ML researchers can comprehend the field's current state more quickly and become productive.

Discussion (0). Sign in to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FastTPS: An Optimized Method for LLM Token Phase for AI accelerators

    cs.LG 2026-07 conditional novelty 6.0 of 10

    FastTPS accelerates LLM token-phase inference via reloading-free static KV-cache management, tiled fused RoPE attention, and interlaced-weight MLP fusion, yielding up to 6× speedup at 93% bandwidth on AMD NPUs.

  2. Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Unsupervised autoencoders on Ising configurations form magnetization then energy representations in two dynamical regimes, with recursive error flow fields sharing topology across layers.

  3. Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Across 74 OSINT/CTI AI studies, hallucination is widely named but end-to-end measured in only one non-reproducible system, so a human–AI co-pilot is the most defensible near-term architecture.

  4. Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy

    cs.SE 2026-03 conditional novelty 6.0 of 10

    MCP server faults form five empirical categories—server setting, server/tool configuration, server/host configuration, documentation, and general programming—confirmed by a 41-practitioner survey.

  5. Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A word-as-image pixel language model trained with next-token prediction reports lower perplexity than a token-embedding LLaMA on noisy and non-Latin-script text, though its noise evaluation holds tokenization fixed.

  6. PhantomHunter: Detecting Unseen Privately-Tuned LLM-Generated Text via Family-Aware Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    PhantomHunter detects text from privately fine-tuned LLMs by learning shared token-probability traits within LLaMA, Gemma and Mistral families, reporting F1 above 96% on held-out derivatives.

  7. Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems

    cs.AI 2026-07 conditional novelty 5.5 of 10

    Reasoning LLMs in ultra-long sessions exhibit latent semantic drift that inverts operator control; a fitted stability coefficient Ks detects the bifurcation and a latent-steering arbitrator is proposed to restore it.

  8. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    cs.AI 2026-08 conditional novelty 5.0 of 10

    For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.

  9. Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models

    cs.CR 2025-08 conditional novelty 5.0 of 10

    Backdoor text samples show smaller log-probability changes under mask-filling perturbations than clean samples, which enables zero-shot backdoor detection without the poisoned model.

  10. What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations

    cs.CL 2025-07 reject novelty 5.0 of 10

    Aya-23-8B appears to activate multiple related languages internally and concentrate code-mixing neurons in final layers, but the paper's own limitations undercut the claim that these are language-specific neurons.

  11. The Impact of Fine-tuning Large Language Models on Automated Program Repair

    cs.SE 2025-07 conditional novelty 5.0 of 10

    On three Java APR benchmarks, LoRA and IA3 adapters match or beat full-model fine-tuning for most tested code LLMs while training less than one percent of parameters.

  12. Hallucination Detection with Small Language Models

    cs.CL 2025-06 reject novelty 5.0 of 10

    A multi-small-model ensemble with sentence splitting, z-score normalization, and harmonic mean detects hallucinations in RAG answers with a reported 10% F1 gain over single-model baselines.

  13. Psychologically Enhanced AI Agents

    cs.AI 2025-09 conditional novelty 4.0 of 10

    MBTI personality prompts measurably change how LLM agents write stories and play strategic games, with self-reflection before communication supporting cooperative behavior.

  14. Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25

    cs.HC 2025-08 conditional novelty 4.0 of 10

    A workshop at deRSE25 produced seven user-interface sketches for LLMs that emphasize branching, context management, and user weighting, which the authors map onto their whiteboard-based interface concept.

  15. LOCOFY Large Design Models -- Design to code conversion solution

    cs.SE 2025-07 reject novelty 4.0 of 10

    A proprietary design-to-code pipeline is described with claimed high fidelity and LLM outperformance, but the evaluation is self-referential, unquantified, and unreproducible.

  16. Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques

    cs.CR 2025-07 conditional novelty 4.0 of 10

    A survey that maps LLM applications, vulnerabilities, and defenses across eight cybersecurity domains, but with significant citation and rigor problems.

  17. Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Distilled students at 43% to 50% of teacher size keep over 90% of teacher Exact Match on SQuAD and MLQA, though one-shot gains reverse on SQuAD test for Pythia.

  18. Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications

    cs.MA 2025-07 conditional novelty 4.0 of 10

    The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.

  19. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  20. Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems

    cs.CR 2025-07 reject novelty 3.0 of 10

    Multi-stage prompt inference attacks against enterprise LLMs are formalized and defenses are proposed, but the preprint gives no reproducible evidence for its central claims.

  21. Is It Time To Treat Prompts As Code? A Multi-Use Case Study For Prompt Optimization Using DSPy

    cs.SE 2025-07 conditional novelty 3.0 of 10

    A five-task case study shows DSPy prompt optimization can improve LLM accuracy on some tasks, notably contradiction detection (46.2% to 64.0%), but results vary and no code or data are released.

  22. From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs

    cs.CR 2025-06 conditional novelty 3.0 of 10

    LLMs can assist both attackers and defenders in cybersecurity, but context limits, hallucinations, and weak reasoning make them unsafe to deploy without human oversight and real-world evaluation.

  23. Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A survey applying the Kaddour et al. challenge taxonomy to GPT-4o and DeepSeek-V3-0324, concluding closed models favor safety while open models favor cost and customization.

  24. A Comprehensive Review of Human Error in Risk-Informed Decision Making: Integrating Human Reliability Assessment, Artificial Intelligence, and Human Performance Models

    cs.HC 2025-06 unverdicted novelty 2.0 of 10

    A review of human error research concluding that integrating AI and cognitive models into human reliability assessment can markedly improve predictive fidelity, but data scarcity and opacity remain barriers.

Pith tools