Pith. sign in

REVIEW 2 cited by

MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11541 v3 pith:AHS3DNFU submitted 2025-02-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords complexinstructionmethodalignmentdatainstruction-followingmethodsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Complex instruction-following with elaborate constraints is imperative for Large Language Models (LLMs). While existing methods have constructed data for complex instruction alignment, they all rely on a more advanced model, especially GPT-4, limiting their application. In this paper, we propose a Multi-granularity Self-Contrastive Training (MuSC) framework, to improve the complex instruction alignment without relying on a stronger model. Our method is conducted on both coarse and fine granularity. On coarse-granularity, we construct constraint-aware preference data based on instruction decomposition and recombination. On fine-granularity, we perform token-aware preference optimization with dynamic token-level supervision. Our method is evaluated on open-sourced models, and experiment results show our method achieves significant improvement on both complex and general instruction-following benchmarks, surpassing previous self-alignment methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models

    cs.AI 2026-01 conditional novelty 6.0 of 10

    Logic-structured rewards—averaging parallel constraints, decaying rewards after sequential failures, rewarding only the active conditional branch—improve instruction-following and transfer to reasoning.

  2. DecIF: Improving Instruction-Following through Meta-Decomposition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    DecIF generates high-quality instruction-following training data from scratch with meta-decomposition and response filtering, and SFT with it improves IFEval, Multi-IF, FollowBench, and LiveBench scores over prior syn...

Pith tools