AnyAudio-Judge introduces a rubric-based benchmark with 7920 samples and a trained evaluator model using SFT and GRPO on 105K CoT samples to assess and enhance instruction following in audio generation.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 2years
2026 2roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following
AnyAudio-Judge introduces a rubric-based benchmark with 7920 samples and a trained evaluator model using SFT and GRPO on 105K CoT samples to assess and enhance instruction following in audio generation.
- MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech