REVIEW 4 cited by
The Landscape and Challenges of HPC Research and LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, language models (LMs), especially large language models (LLMs), have revolutionized the field of deep learning. Both encoder-decoder models and prompt-based techniques have shown immense potential for natural language processing and code-based tasks. Over the past several years, many research labs and institutions have invested heavily in high-performance computing, approaching or breaching exascale performance levels. In this paper, we posit that adapting and utilizing such language model-based techniques for tasks in high-performance computing (HPC) would be very beneficial. This study presents our reasoning behind the aforementioned position and highlights how existing ideas can be improved and adapted for HPC tasks.
Forward citations
Cited by 4 Pith papers
-
CelloAI: Leveraging Large Language Models for HPC Software Development in High Energy Physics
A locally hosted RAG-based coding assistant improves kernel retrieval and porting coverage for HEP codebases, though no tested LLM correctly ports the hardest kernels.
-
Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling
ReAct-style LLM schedulers can balance multiple HPC scheduling objectives on 10-100 job workloads, though cloud API latency makes them unsuitable for real-time deployment.
-
HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks
Confidence-modulated per-response advantage weighting (HARGO) improves GRPO-style RL post-training on four heterogeneous HPC tasks, leading WinRate, data-race F1, and PLP similarity at 0.5B.
-
LLM4VV: Evaluating Cutting-Edge LLMs for Generation and Evaluation of Directive-Based Parallel Programming Model Compiler Tests
In a six-model comparison, DeepSeek-Coder-33B generated the most passable directive-based compiler tests (Pass@1=0.434) and Qwen2.5-Coder-32B judged test validity best (F1=0.735).
Discussion (0). Sign in to comment.