Pith. sign in

Paper Citation Record · LEDGER

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

As of 18 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 13 inbound Pith citation observations for arXiv:2506.16141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16141 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:48:37.725055Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:35:11.261146Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1f1d5302-beb0-423c-96f4-79e6943e7596 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:32.974881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:32.974881Z digest=sha256:1350d9db9590737901cd98d633aab59ec3810a0eccea60ffb8b3b6f96fe326e2

Observation 24167e82-b909-4652-9040-1d7369cbefcb · outbound

This paper cites an unresolved cited work.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:48:41.128978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:33.018546Z digest=sha256:4312437d064c421ffcd019ae5f306d41078847804a12d3bf3a086cbcba49ce8e

Observation 96bc01f3-3303-4e46-94cd-badd1d66a95a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.083995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.083995Z digest=sha256:6ce0fa9178a9497dfb39806fe0b913bfa6638fdb976524d008a48caacdeecd60

Observation 07a8d6fe-123a-4013-9744-9ef7209e42e7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.164335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.164335Z digest=sha256:fb84992e210bbc21d6ef7e07d6f2d1e31780a0ec0b4c470f5382b540f52d0de5

Observation 7aedcf7f-19ce-4c21-afc1-14879ba2e61c · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.223889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.223889Z digest=sha256:9efc86eb2e7179c393d8b879894a70ec50cb673a846dabdfabd603b2e579bedd

Observation 87668594-5f01-4f8b-87d3-d9f12a73318e · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.343894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.343894Z digest=sha256:f95b273655298c48a4b8c224c3c3519abd1e455290729a7230fffd4f887b622a

Observation 7ba8fcbb-2e6b-49cc-a7f8-6b01284932df · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.451030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.451030Z digest=sha256:5d78eb483643b65d809da312334ae37bdfddb9ab5d88eb98f2a00e1ca3391cbc

Observation f1dcbb12-835c-47db-b815-97acdf0acfd0 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.585809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.585809Z digest=sha256:0bdb0864e7403ecbbfac4bb3e4d051ae4473abcb5d6cd0c930014659fc7d1661

Observation 4d460403-1db4-4434-a36a-5a883226b3fc · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.702470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.702470Z digest=sha256:87373096301499cddcd2ef76e6be2db369393726d8a488c7d57fe110974bc2c1

Observation 01639783-15c5-441b-99d6-50825f3a2196 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.790717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.790717Z digest=sha256:b22869757f37c3041832f33e3f0df20372587b1ba3e17f961c3a365a40131153

Observation 3e860905-a7be-4343-940b-cca4e68e3f50 · outbound

This paper cites Video-r1: Reinforcing video reasoning in mllms, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-r1: Reinforcing video reasoning in mllms, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.869684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:33.849002Z digest=sha256:8bca91cb123c44e0cc92f4c844d66aaf610164cdf6ffb37aa4c507db6e1bf595

Observation 1b678cb0-2ede-464c-a21e-6c7e4010fef8 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.921347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.921347Z digest=sha256:64a76adbace366ff972a4998a537f50d95e8575ff28f0231c19e4b37e82d6f59

Observation fed62ce0-a034-485a-bf36-63085f2fbbd6 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:33.993024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:33.993024Z digest=sha256:c0ca78f2160795ecabc19285782eea33ed439e638763cedaae144f8770fc5dc0

Observation caef7470-ac90-4e9c-8f9c-ee95ebbe62f4 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.652360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:34.042645Z digest=sha256:d5ce07a74a57605241b1475a56ccd4dc40b905ad05714bf19630cfc0182af70a

Observation 5ff65fe0-73c8-4310-8e5b-87392040dbc7 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.135877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.135877Z digest=sha256:0586d09f638812c3c566c0b956317be68babc4f2374f794a08817118ded6cfdd

Observation f9c40393-dab8-4541-82f8-5cca8d1246b1 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Proximal policy optimization algorithms, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.372247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:34.212172Z digest=sha256:c0a925b9b4d86de8a95218b71978bc5d2b7b83976a9052304bcb0f71761afff0

Observation ba36716a-bc1f-4057-bbe6-55210a24beae · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.320914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.320914Z digest=sha256:152cd40c2ec68e3c945592d658311d628af5c450d67bc3b2744b01d7d7776293

Observation e55eb96e-5451-4977-8298-c7ca82c2048c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.438236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.438236Z digest=sha256:febb8cee2f7e8c69fbec055f558c999a41a6f632bfaa5e082d7223e634c8d8e9

Observation 55dc7406-8b3a-440b-b8c9-44bd12456a59 · outbound

This paper cites Let’s verify step by step, 2023.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Let’s verify step by step, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.522538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.522538Z digest=sha256:1343e1483ab35b0dd5cbea77fa78ae934eea2744a7fb1e2df182c3976b0ef707

Observation 0924cf59-4514-4f59-8d31-deddba17e0f4 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Solving math word problems with process- and outcome-based feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.603403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.603403Z digest=sha256:7dbcf8d9eb827a283d8406247e1532dfe3d22d9cc2b5dfb723819fa2a65e939e

Observation 703c21e7-fa80-4e50-8baa-1a99f975d6ab · outbound

This paper cites Alphamath almost zero: Process supervision without process.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Alphamath almost zero: Process supervision without process

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:40.165451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:34.660272Z digest=sha256:41a87e05c9e8d2dbe5f1a4a8287634b2f1242389e94e09fcecdb9c94e3206940

Observation 210ad232-6d95-4711-a1dd-85c5fe2758f7 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.742575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.742575Z digest=sha256:f2206c62d4bf751a1128c1dc32d37a9b9bdbb2d260280d4d93c3ae5bd6f2429f

Observation 0feb8827-418f-418e-8da2-a7028c875bb9 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.838198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.838198Z digest=sha256:e03e0321aa6e7a448f15f940cfe93a752f448f10c2a611dc9274060828948d39

Observation 3e166373-3cab-4d7d-93c9-1559829a576c · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:34.894129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:34.894129Z digest=sha256:b5a9d3ac6d465f560530731f2d8fc466fffe8b06716cd5486d99da41c43a7f83

Observation 4249718d-264d-47b1-9722-2ef4abd5d1ae · outbound

This paper cites Evaluating mathematical reasoning beyond accuracy.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Evaluating mathematical reasoning beyond accuracy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.929136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:34.955288Z digest=sha256:7d18087cf0dfe4948cf053b4cbc4aa5a2d2ed7a38190126b0f85a800600a43e7

Observation 97e55492-936f-4c0e-a906-accf45d41bb2 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.066396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.066396Z digest=sha256:3223546971c7fcd3794c27c5869dce2bcc90cf65ccc005e1284c36b8faf41afd

Observation 82a53331-fc24-45c0-93c9-e2aad67d13f5 · outbound

This paper cites Warp: On the benefits of weight averaged rewarded policies, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Warp: On the benefits of weight averaged rewarded policies, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.715436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:35.168865Z digest=sha256:4db7ff6b03aecefbe4cb77e8bda43e041ea4dbb2322c8ab91cb13a077d64235c

Observation 3a7d7d8c-b443-443e-a03b-13792fd8bef1 · outbound

This paper cites Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Gtr: Guided thought reinforcement prevents thought collapse in rl-based vlm agent training, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.276903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.276903Z digest=sha256:a5ea86b767f9b552011fe83ccd5e01ba05296db3da152764446592db34de5f9d

Observation c747885d-dd23-4ad6-9ada-f9fdd4f839b3 · outbound

This paper cites Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mm-math: Advancing multimodal math evaluation with process evaluation and fine-grained classification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.435084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:35.459609Z digest=sha256:be8a92fb05e9212b5b556f0c9570a9a8439aa6cd11d39f63992e6fdb0e3fddbd

Observation 5527d3f9-a4c7-4a69-9141-87eaf4ece1dc · outbound

This paper cites Open-r1-video.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Open-r1-video

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.241098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:35.524033Z digest=sha256:5303d87bf95544fbc3a03832c2af26f2836d6442223a3f93bea6380a9578052f

Observation f4b54ff6-e282-4069-af2e-5c4dee1f3d62 · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.612776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.612776Z digest=sha256:63e8ef00f22feac1deb6b02b86b52da81f9a39db092a38b9c32cef75c359e1b9

Observation 4042c7ab-253b-407c-887f-31b9ae9694a2 · outbound

This paper cites Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:39.023697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:35.679709Z digest=sha256:622bebfdfd0ef807143bac040985853c2c354e04ea00e62d7f7c994cbffd6b6d

Observation e7c72394-096f-4d90-8344-71c77b1bb301 · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Dfew: A large-scale database for recognizing dynamic facial expressions in the wild

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:38.821362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:35.767077Z digest=sha256:2319341abbf81d3d0f3ab507e5048a3bd0514c4ab631fe193fafa94e686e943c

Observation a0df6ca4-fdf4-447a-850d-7e9688d04544 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.861496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.861496Z digest=sha256:5aece424bf64757df1bda0d848acd670638d88338873a3aa38339f99b861f3c6

Observation 22a0426f-ac6a-4cdc-933e-c641f0d5a4d0 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:35.947030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:35.947030Z digest=sha256:160bb3a84266ca235ee4035b3c0fcdd22bf2d432393613ae76f2ce389909d94e

Observation db357cd1-ea01-459b-945c-e6a9cfcee43c · outbound

This paper cites Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics ACL 2024, pages 8731–8772, 2024.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics ACL 2024, pages 8731–8772, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:48:38.593156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:48:36.039772Z digest=sha256:9881933e856939b07339ac204c5eadff3ac4df0090477bce4ec6d3b49f0a041f

Observation b6b4a07a-446e-465a-b0b2-fc16fed427ae · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.245848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.245848Z digest=sha256:d83fcce4bbb2fb72c294dff86887364d62c941ac7674d87225da86aefbecc36c

Observation 8120ea48-5a60-4ed9-b67c-264cbe53d991 · outbound

This paper cites Qwen2.5-VL Technical Report.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Qwen2.5-VL Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.371774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.371774Z digest=sha256:a3a94b7edccfd64f291dae43e66227b29d90cc0003d7c1747151147cd7f75cbf

Observation d1bee5ad-baba-454a-9bb7-b7384a0c64e1 · outbound

This paper cites GPT-4o System Card.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning GPT-4o System Card

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.504016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.504016Z digest=sha256:83d2a391611606494c964aa945daf3d11a75dc228efc2fdc4c9ede9128644b1e

Observation 52bd632d-3da6-4f60-8a1d-5a5eb892d476 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Llama-vid: An image is worth 2 tokens in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.600274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.600274Z digest=sha256:098bf62b3765956e05c350982a78a98da533b29d3ae19447bd54c496f0d32aae

Observation 9fa98f0e-5053-4c37-abe0-bb1c03d95e62 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.680253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.680253Z digest=sha256:4e86d658d8afaf83b547c1b0268921b3933bdf873dcba22516fe65294fb2cd61

Observation 0b8f8ce4-9400-4786-81bd-7c484317d5cc · outbound

This paper cites Long Context Transfer from Language to Vision.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.814956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.814956Z digest=sha256:9916842d8beb793a490b2eafe338651fa2de255728dfaa981504a3bd98bea493

Observation 94c136af-0e49-4990-a1f1-2c94e53f5172 · outbound

This paper cites Vila: On pre-training for visual language models.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Vila: On pre-training for visual language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:36.999749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:36.999749Z digest=sha256:9f9bc34ec29c93473bd1b0b19e8215c5838e20420ed5a25084d4d43dba15fb16

Observation 847d3506-46b5-4ed4-996a-c47b8e2c793e · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.181614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.181614Z digest=sha256:60c33d38c84a71e3b6fcd99192a4c67fa839746ccb69fffbb92aa9fe67062e60

Observation cd5f36df-dc5c-4131-8cbc-28b81e2c7ef5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.284629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.284629Z digest=sha256:5d2a45e01afe5db8b7c0195e17e2d1c1a445d828a0c9145a6cbbd9dd51a99561

Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.373080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.373080Z digest=sha256:b0dc3119af55b6a7c9ea839d48f945be8501f26d372052c82f3ebd2371f11f67

Observation a6198fb0-460d-4309-8991-d026ef49699f · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.464542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.464542Z digest=sha256:8be860b327fce235ad7792b35460e77400003f92dedc92c094f110fa80165758

Observation 7e2d0813-3723-406c-9f37-99613b9ef861 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.556314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.556314Z digest=sha256:9abbdb0a8b0071e93366a007ec74187d129e2e3dd10287c56ca9ae760424844b

Observation cfb50d06-db4a-4122-b1a9-f8f7b1410dc9 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.656075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.656075Z digest=sha256:0706305178ccde71ea7340f82c843d56682a6775b497aea4986da0d78bb8f9d9

Observation c93b68ad-4f32-4fad-90d7-69d5909b854a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.725055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.725055Z digest=sha256:e1b17bf09b84d3fa7cf94bc77941c4f2943540088334a4826239c984b91ebf3e

Pith citing papers

Observation 98c1c08e-f4f6-4914-8582-a873fff20c39 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:11.261146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:11.261146Z digest=sha256:40f6553dc6abe68efa6eb8db766e8b949b42edb8b2cb20a71aafbeff6c550d83

Observation ac8c59c9-589e-4ac1-9487-00d555b75867 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:03.419373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:03.419373Z digest=sha256:79c90c9161940a9c1af6d1a2fd4d5ca77af9cc547c2f9022af6b853dbbe121db

Observation ea14883e-98a6-4fb9-9ef4-a8371292709f · inbound

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation cites this paper.

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T00:15:23.806078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T00:15:23.806078Z digest=sha256:1d084ffe48b899e85e63a270d4e3c825964bbdf3e817c656947c389c67c21aa4

Observation ba0bfcdd-be2f-4489-b396-47970bb67ef4 · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:15.673718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:fe9702b22e6e94e5fa3d6e9b451bea27c594ea2e40be2ac482537b0e3e4db1ed

Observation e9c14120-8beb-4b12-8b0b-274c4bf4b9d0 · inbound

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization cites this paper.

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:27:51.300574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T19:24:05.375951Z digest=sha256:9b2027d133f0d743d6c77e3455fd75874d3d31af6427b274bc33c05445e9066a

Observation f3429f25-bb38-4d17-b2ff-70b78648a36c · inbound

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization cites this paper.

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.752486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T08:53:02.051453Z digest=sha256:3f33219bd58e59c2b08935fbc8929dc01057a5e98d30472e42babbed2d8755e8

Observation 9b56a2e3-91fe-42c4-a10f-3c2b1b66a312 · inbound

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding cites this paper.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.092273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:e1d94c5d5bcbd723479c1d336f4290fdd63e879363943eb4afd1168ae356f0f7

Observation 14e8d7db-96f7-45bf-99ff-3b55098fb9db · inbound

Touch-R1: Reinforcing Touch Reasoning in MLLMs cites this paper.

Touch-R1: Reinforcing Touch Reasoning in MLLMs GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.639487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T18:55:05.919049Z digest=sha256:4b63ab1e8eed608a9d00b8f3cc1e5af4935e4c82c37c6305ecf3999561a5c416

Observation 7dbb6506-44d7-4151-a07d-775cf052e55e · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.320768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:5e1e56c4f3a9f34ef683ed1f2f96e66491e6b4ed29d44f3e738cfe058171fe5a

Observation 1ef5e2ec-7c53-4b66-8092-847994b7ba01 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:44:36.816892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:c276ee6b815ffd7a11a2526286a82e89d5a56e0064cdcfd6365791f9f0f073d1

Observation d4b025b0-8a22-45d5-aae7-e2db5fb881f0 · inbound

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection cites this paper.

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.252360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:58:30.855338Z digest=sha256:055c153653adaae6030fe5e1177733ef42a41ff468ee2ef17befe2a7f71c22d5

Observation d83f06a3-2c79-4a41-b0bd-fc3c6771115c · inbound

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning cites this paper.

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:41.169558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T05:55:19.083517Z digest=sha256:93dc9e25e9438ed87ad277de19426de22caea5068852a768b2830d13460dc100

Observation ad9f6d09-1dc7-4dd0-9c3c-8b7df3220910 · inbound

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation cites this paper.

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:00.388339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:00.388339Z digest=sha256:c3298d91991acb1a8be2b8fc80d6de320b0a86650762fb18966720455fa1e28c