pith. sign in

Bmmr: A large-scale bilingual multimodal multi-discipline reasoning dataset

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

citation-role summary

dataset 1

citation-polarity summary

fields

cs.CV 2 cs.CL 1

years

2026 2 2025 1

verdicts

UNVERDICTED 3

roles

dataset 1

polarities

use dataset 1

clear filters

representative citing papers

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

cs.CL · 2026-06-16 · unverdicted · novelty 7.0

ZPPO improves distillation to small vision-language models by using binary and negative candidate prompts plus a replay buffer for hard questions, outperforming standard distillation and GRPO on a 31-benchmark suite with largest gains at the 0.8B scale.

citing papers explorer

Showing 2 of 2 citing papers after filters.