REVIEW 19 cited by
Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we introduce Hunyuan-Large, which is currently the largest open-source Transformer-based mixture of experts model, with a total of 389 billion parameters and 52 billion activation parameters, capable of handling up to 256K tokens. We conduct a thorough evaluation of Hunyuan-Large's superior performance across various benchmarks including language understanding and generation, logical reasoning, mathematical problem-solving, coding, long-context, and aggregated tasks, where it outperforms LLama3.1-70B and exhibits comparable performance when compared to the significantly larger LLama3.1-405B model. Key practice of Hunyuan-Large include large-scale synthetic data that is orders larger than in previous literature, a mixed expert routing strategy, a key-value cache compression technique, and an expert-specific learning rate strategy. Additionally, we also investigate the scaling laws and learning rate schedule of mixture of experts models, providing valuable insights and guidances for future model development and optimization. The code and checkpoints of Hunyuan-Large are released to facilitate future innovations and applications. Codes: https://github.com/Tencent/Hunyuan-Large Models: https://huggingface.co/tencent/Tencent-Hunyuan-Large
Forward citations
Cited by 19 Pith papers
-
InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
InfiniteHBD embeds optical circuit switching inside each transceiver to build reconfigurable ring networks for GPU clusters, claiming node-level fault isolation at roughly one-third the cost of NVL-72.
-
Can Vision Language Models Understand Mimed Actions?
Vision-language models identify real actions with context far better than they identify mimed actions performed by 3D avatars, while humans are equally accurate on both.
-
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
SOAP and Muon, stabilized by per-step QR eigenbasis updates and KL-Shampoo covariance accumulation, beat AdamW on large-batch LLM pretraining up to 100M-token batches.
-
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
TSHA is a new 80,000-pair benchmark for indoor safety hazard assessment; current vision-language models score roughly 45-85, and fine-tuning on TSHA raised Qwen2.5-VL-3B by 18.3 points on TSHA's test set.
-
SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation
SegMoTE shows that adding token-level mixture-of-experts routing to a frozen SAM decoder can match or beat medical-segmentation models trained on far more data, using 0.15M curated masks and 17M trainable parameters.
-
Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
Sparse attention with chunk-aware sparsity growth and hierarchical frame/block selection accelerates autoregressive video diffusion at ~1.3x with VBench quality on par with dense attention.
-
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
PISCES post-trains text-to-video models using dual optimal-transport-aligned rewards (global quality plus token-level semantic) and outperforms annotation-based and annotation-free baselines on VBench and human evaluation.
-
Forecast then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
A training-free predictor-corrector method that accelerates Diffusion Transformers by solving a feature-ODE, achieving large compute reductions with modest quality loss.
-
DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
A new benchmark claims to be the first to test VLMs on both external and in-cabin driving risks, and reports a fine-tuned model far outperforming all baselines.
-
PanoLora: Bridging Perspective and Panoramic Video Generation with LoRA Adaptation
Fine-tuning a pretrained video diffusion model with LoRA rank 16 on about 1,000 synthetic videos produces panoramic video with good seam closure, but the claim that rank must exceed 8 degrees of freedom is not proven.
-
RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
RLMR dynamically adjusts penalties for constraint-violating creative-writing samples during GRPO training and reports gains in IFEval and human preference over fixed-weight baselines.
-
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
Grove MoE uses unequal-size adjugate experts with complexity-based activation to run 33B-parameter models at roughly 3.1 to 3.3B active parameters while matching larger open models in benchmarks.
-
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation
PER-based preference optimization (DPO, PPO, GRPO) reduces lyric-to-song hallucination in an audio language model, with the largest gains from DPO plus reject sampling.
-
ViSP: A PPO-Driven Framework for Sarcasm Generation with Contrastive Learning
A new sarcasm-generation dataset and a reward-optimized vision-language model that outperforms LLMs on benchmark metrics, though its main quality metric is the same model used to train it.
-
Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
An empirical study proposing a 'basic-refinement' split in MoE models, where shared experts generalize and routed experts specialize, with efficiency claims undermined by internally inconsistent numbers and an unvalid...
-
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
A field study of 308 generated RAIL licenses and 1.7 million HuggingFace models shows growing adoption of behavioral-use clauses, and the authors argue the next priority is building tools to track adherence.
-
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
An on-the-fly router that predicts speaker-specific adapter weights lets a speech foundation model adapt to dysarthric speakers with zero-shot, real-time processing, achieving the lowest reported word error rate on UASpeech.
-
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.
-
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.
Discussion (0). Continue with ORCID to comment.