REVIEW 10 cited by
HyperCLOVA X Technical Report
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean, English, and code data, followed by instruction-tuning with high-quality human-annotated datasets while abiding by strict safety guidelines reflecting our commitment to responsible AI. The model is evaluated across various benchmarks, including comprehensive reasoning, knowledge, commonsense, factuality, coding, math, chatting, instruction-following, and harmlessness, in both Korean and English. HyperCLOVA X exhibits strong reasoning capabilities in Korean backed by a deep understanding of the language and cultural nuances. Further analysis of the inherent bilingual nature and its extension to multilingualism highlights the model's cross-lingual proficiency and strong generalization ability to untargeted languages, including machine translation between several language pairs and cross-lingual inference tasks. We believe that HyperCLOVA X can provide helpful guidance for regions or countries in developing their sovereign LLMs.
Forward citations
Cited by 10 Pith papers
-
Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations
Cooperative paging with keyword bookmarks and a recall() tool yields the highest answer quality among six long-context methods on LoCoMo across four LLMs.
-
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.
-
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
A new benchmark, Nunchi-Bench, shows that LLMs know Korean superstition facts but frequently fail to apply them in practical cultural contexts, and that explicit cultural framing beats prompt language alone.
-
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
A cost-effective recipe consisting of tokenizer extension, continual pretraining, FP8 training, and SFT/DPO post-training yields Korean-English bilingual 8B models with top Korean benchmark scores.
-
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
Selective-pivot-token training plus layer-averaging-with-recurrence reportedly gives 2.5x parameter compression on a small Korean LLM, but the efficiency claim lacks its decisive controls and the abstract advertises r...
-
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
A 102B Korean-English model, expanded from Llama 3 70B with LlamaPro and Masked Structure Growth and trained on 194B tokens, scores 64.74 on KMMLU and 83.34 on KorMedMCQA, roughly matching GPT-4 on Korean benchmarks.
-
KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models
A 3.3B-parameter NLLB model fine-tuned on Korean social media data beats GPT-4o and HCX-3 on BLEU for Korean grammatical error correction.
-
Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales
A benchmark of 809 yokai questions shows Japanese-centric LLMs, particularly Llama-3-based continual pretraining models, outperform English-centric models on Japanese folktale knowledge.
-
Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
A three-step 'Doppelgänger' conversation makes LLM agents drop their role and leak internal prompts, and a CAT defense prompt only partially stops it.
-
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
The submission's abstract promises an LLM safety survey, but the provided body is the opening page of an unrelated arithmetic-dynamics paper, so the artifact is internally inconsistent.
Discussion (0). Sign in to comment.