REVIEW 7 cited by
LLM-jp: A Cross-organizational Project for the Research and Development of Fully Open Japanese LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces LLM-jp, a cross-organizational project for the research and development of Japanese large language models (LLMs). LLM-jp aims to develop open-source and strong Japanese LLMs, and as of this writing, more than 1,500 participants from academia and industry are working together for this purpose. This paper presents the background of the establishment of LLM-jp, summaries of its activities, and technical reports on the LLMs developed by LLM-jp. For the latest activities, visit https://llm-jp.nii.ac.jp/en/.
Forward citations
Cited by 7 Pith papers
-
Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs
Instruction-hierarchy compliance in LLMs is asymmetric by language and position, and cross-language conflicts yield systematically higher compliance than same-language ones (Language Boundary Effect).
-
JETHICS: Japanese Ethics Understanding Evaluation Dataset
JETHICS provides 78,000 Japanese moral judgment examples, and evaluation shows GPT-4o averages about 0.71 accuracy while the best non-proprietary Japanese LLM reaches about 0.50.
-
The Emergence of Abstract Thought in Large Language Models Beyond Any Language
Across 20 open LLMs, shared multilingual neurons grow in number and per-neuron importance over release generations, which the authors interpret as evidence of language-agnostic abstract thought and use to guide neuron...
-
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
A Japanese safety dataset with reference answers improves LLM safety through fine-tuning and provides a benchmark that reveals large safety differences across 12 models.
-
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
Using a new LLC Score, translation and geo-culture cloze experiments on three small multilingual LLMs show that latent-language consistency does not reliably predict task accuracy, contradicting the paper's initial hy...
-
Cost of Reasoning in non-English Languages: A Case Study on Japanese
Japanese reasoning-language control is feasible with CPT plus GRPO, but incurs a capability cost and does not free-improve cultural Japanese performance.
-
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
A single BERT-scale model with a game-context token and LLM-assisted label transfer achieves toxicity detection comparable to per-game models while extending to seven languages.
Discussion (0). Continue with ORCID to comment.