Pith. sign in

Paper Citation Record · LEDGER

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

As of 12 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2608.03972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03972 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:53:13.195296Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:47:41.755414Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T16:47:42.403064Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact6
  • verified fuzzy22
  • unresolved27
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71316cb8-5dfe-4eca-aeb6-84068d6591a4 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.524151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.932970Z digest=sha256:295c8892a3e6e1b44aee95e9f4fa86c5e2437e8903e0f28ba2050dafa1aac1a7

Observation a6cd4b45-4225-435a-9331-cba37575e76f · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning DAPO: An open-source LLM reinforcement learning system at scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.503525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.938125Z digest=sha256:aa5bb3920b50130363c571f3e9d02598d8e165cc85c55663642f32b331bf1c80

Observation 8dd9ea9a-6b98-4d28-9433-78731f14513b · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.942968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.942968Z digest=sha256:a2c00e04c464274b5c455133afb2029f121eced209817e2fe087ee19b3c2683e

Observation b9f173e5-5435-4412-8d60-8680b15edc49 · outbound

This paper cites Loong: Synthesize long chain-of-thoughts at scale through verifiers, 2025.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Loong: Synthesize long chain-of-thoughts at scale through verifiers, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.484558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.947974Z digest=sha256:08308f9fff929611e73e5f4fb9d25d4bd77bce754891c4337e4b8ccccb53f2d5

Observation 676f668a-52ec-46bb-8205-8d92a83901b9 · outbound

This paper cites Reinforcement mid-training, 2025.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Reinforcement mid-training, 2025

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-05T04:53:13.931848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.952698Z digest=sha256:4e92946ef3b8fe147f2465d3bd9f0504874af0f08c5473d678af3bb87d10dd1e

Observation 44f0d7fa-e63b-4d3f-95ab-b22290f112a8 · outbound

This paper cites Self-evolving multi-agent systems via textual backpropagation.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Self-evolving multi-agent systems via textual backpropagation

Reference 6

Resolution
verified exact
doi, observed 2026-08-05T04:53:13.269181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.957456Z digest=sha256:a9cc64e975041320989d79859090ae90f537307f40d9e55fd63ae80f5faf2acb

Observation 67a31706-7647-441c-8dc1-7555dd841db5 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Connectionism,.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning On-policy distillation.Thinking Machines Lab: Connectionism,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.967025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.967025Z digest=sha256:62a879977c6bbb1970a2fdfc90fc302aac945a1e67ec9981b2fbc0fd0efdbcc3

Observation 08623710-4d5f-4471-9ed1-b85223d145b3 · outbound

This paper cites EchoRL: Reinforcement learning via rollout echoing.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning EchoRL: Reinforcement learning via rollout echoing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.431786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.977027Z digest=sha256:fe8202f5141cf5d7a1ec91ec6a8ec33d0b919999851ae8f554ccd90fb7e67b62

Observation 52a20fe1-d4db-4942-bd10-d5997539c958 · outbound

This paper cites Learning to reason under off-policy guidance.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Learning to reason under off-policy guidance

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.403651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.981914Z digest=sha256:2271fccfb3338bd2f27f52a76898d1d450f653c2803f3636f9d57c9c9bc12a85

Observation b0ef0794-26c1-4934-9d36-ef04fc58c1de · outbound

This paper cites Ozdaglar.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Ozdaglar

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.384926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.986788Z digest=sha256:a052150875d3cbf1abc0cef39b1f2ab5e67367a1462fe3c74210d598651dd162

Observation b85e76e7-0456-4d29-bc9a-c56812d5dcb0 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.991139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.991139Z digest=sha256:6aa01a894d674b3eeea13bcfe58a79f314498b8e5261da8ff935ca6256562a86

Observation 3e10b2e6-8c8f-44a6-aa55-41da810147f7 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.995759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.995759Z digest=sha256:df09b52869a9e76f02ecf6c20b567e006f51e6efaaf1d5ddcee8b295cfc44ef6

Observation ed2e493a-78a5-41cf-b3d3-c6115491da56 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Measuring mathematical problem solving with the MATH dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.000441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.000441Z digest=sha256:3ce4aa2c8af11870f42540c287bb593baa070eb0f448dc942c4c69cccca69565

Observation f3f5af37-516a-4256-a71d-ad4ed9033892 · outbound

This paper cites Solving quantitative reasoning problems with lan- guagemodels.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Solving quantitative reasoning problems with lan- guagemodels

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.356660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.004741Z digest=sha256:537b2c109e8b3afa6f204259f838dbd999dd780018829cbc1536d819aafb7d3c

Observation 9a4ab40e-1c4d-42b2-abc4-82df9bd4935a · outbound

This paper cites OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.009142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.009142Z digest=sha256:be937bbe975d288205cd82f78ef4702b0af83eb10fd81ac6ac2593f381cd454e

Observation 459ab9fd-1d12-4713-a675-b2c08aa7424b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.014357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.014357Z digest=sha256:32e1bc96e4032a1fca959934d94fbc9c85d12d89d5952dd4914c9ee19bfb21a0

Observation 1112e77e-80b2-45ad-b511-b36b7e5fab67 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.019460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.019460Z digest=sha256:4e23b19835ce7e1e6bfcf98385e7c07dec83e8728e27d3f12d64a9193d41812f

Observation efe2d354-b221-4331-b963-3c3aa75c560c · outbound

This paper cites MMLU-pro: A more robust and challenging multi-task language understanding 12 benchmark.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning MMLU-pro: A more robust and challenging multi-task language understanding 12 benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.319942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.024162Z digest=sha256:aba9c1933b7873552722184d4d4f85fcea5108ca9642075562e433f0b41bb790

Observation 8f65f649-835a-4863-9ab0-ad4026eaee28 · outbound

This paper cites SimpleRL-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning SimpleRL-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.303448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.028821Z digest=sha256:6dc7a2e5e1740c70c6643ac96fcf013e6a5310812c851002c0cc3d48722b8156

Observation 6ebc4b3a-56c6-45b5-84ad-3dcc8aa093ca · outbound

This paper cites Open- reasoner-zero: An open source approach to scaling up reinforcement learning on the base model.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Open- reasoner-zero: An open source approach to scaling up reinforcement learning on the base model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.286659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.033764Z digest=sha256:189d0535f66bf4c43bc4f86b55dba9358509ef6ef67026270f0f361345c750dc

Observation b6cf11d9-1ffb-4e6d-b7ca-4fd5826cd031 · outbound

This paper cites Process reinforcement through implicit rewards.Transactions on Machine Learning Research,.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Process reinforcement through implicit rewards.Transactions on Machine Learning Research,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.268959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.038050Z digest=sha256:b7229ff7c74693ec8b2a99e99a84000e6512e116623258099d300cdbdda3bf04

Observation db9f4ef3-9ecd-428d-9fed-a5df98748910 · outbound

This paper cites Qwen2.5 Technical Report.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.046874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.046874Z digest=sha256:1c4cd784dd702fff8a610d2d0f825fc23fd8bef17ce8efa9f234d749a9ebfa8e

Observation 8e59e571-48aa-4a9a-8545-fbb844e7488b · outbound

This paper cites The Llama 3 Herd of Models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.051552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.051552Z digest=sha256:48f30c57c95ba8873e9a70e15a41ffdc4718286ce0905c963254f3befccc29dc

Observation c97db0fe-842d-4390-928a-c1a0eee4a97a · outbound

This paper cites Think outside the policy: In-context steered policy optimization.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Think outside the policy: In-context steered policy optimization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.237674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.057291Z digest=sha256:9abc181e8974d878187a00a5986d7af171764156c4af43f69a04dd15a4a92b26

Observation f4754bd8-ebb5-4826-81cc-08f8df5b9e12 · outbound

This paper cites SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.062069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.062069Z digest=sha256:d605c501acdd995a1fb8e46166b0ec03a5a4399df9676872ed6ad9f080f597c2

Observation c6f71012-67b5-4f5f-8950-493008b635b8 · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.066820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.066820Z digest=sha256:0bff9893dc847cf4c18559ded1687b99d33dbf1edb40b57dcdddae0947aaeb1a

Observation 4113cf82-1aa0-4f83-a1e1-30163a41d3b3 · outbound

This paper cites Don’t overthink it.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Don’t overthink it

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.220813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.072079Z digest=sha256:101a12fa800f88efd216b8e927287f8d25aad013ff50a5e81fcd338c903a8b12

Observation aa62ebab-5ecd-4e67-913b-e590d9786438 · outbound

This paper cites LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation- steering.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning LLaVA steering: Visual instruction tuning with 500x fewer parameters through modality linear representation- steering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.199622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.077044Z digest=sha256:09f4aea98f188740accde744948974c17668416d5ff511db6f6100e851de240e

Observation 00d862b3-3b99-465b-8966-3f520b032d86 · outbound

This paper cites PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.081608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.081608Z digest=sha256:cb64a27448fa65daec0954097c77f35447dc58f30a3216a0b28a0fcbf54bc201

Observation 11290c47-93f1-4ba8-a121-b8c902397195 · outbound

This paper cites Beyond nl2code: A structured survey of multimodal code intelligence,.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Beyond nl2code: A structured survey of multimodal code intelligence,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.169785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.091539Z digest=sha256:40636bfdbc6263e7ce9a02624bf1d256e7ad1ef5ae770e2be79977b8148ccf41

Observation 3cf0b7a4-e5fd-4080-a050-efd75a89652f · outbound

This paper cites SPOT! Revisiting Video-Language Models for Event Understanding.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning SPOT! Revisiting Video-Language Models for Event Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.101410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.101410Z digest=sha256:143d4861578ac01cda09d02b60307bcc97613d744ce6efd0683412d2ea46a73b

Observation d02da753-95da-47eb-86c2-22a74beb1c38 · outbound

This paper cites Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:53:13.646604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.106954Z digest=sha256:9fe6f3deafe7703f4be5d5175dc3688ffb17c09ff80e3362e0c32d3d5cd3e54f

Observation 614957f0-49a5-4284-9ca1-8e36fb0678d9 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.086535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.086535Z digest=sha256:13b55c07a9a4368c394f3e433bfffe3b03801c66022fc79151b7a349a5482d60

Observation caefa43e-66b5-448b-9654-1b652cf39833 · outbound

This paper cites KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.116166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.116166Z digest=sha256:66e1eb1d03f52d0a0b41a006859c1fdefbe31981fcb2e41e4682c03e90ebcf0b

Observation 641dc76e-2364-4908-a5ea-84cbcf94c2f3 · outbound

This paper cites Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:53:13.690777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.095917Z digest=sha256:9adbbdcce01817c412cdaccb99dcfb2352cfa290379f0b58a11d04eff239d0b7

Observation f9a3aeed-c64c-4e64-a40c-744ce16e7669 · outbound

This paper cites Aditya Prakash, Yizhou Sun, and Wei Wang.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Aditya Prakash, Yizhou Sun, and Wei Wang

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.125098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.125098Z digest=sha256:b65582b9a3d4bef4610f5b192226b501a3856772efe52891d47353016d5fa95c

Observation cd29fcac-7c1b-45c6-bf24-4b2c68e8208f · outbound

This paper cites HYPERION: Fine-grained hypersphere alignment for robust federated graph learning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning HYPERION: Fine-grained hypersphere alignment for robust federated graph learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.154663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.129286Z digest=sha256:5c8cdc1f29fb08de536a59f549cff57a300ae14251938f2a33c61c25effb3e55

Observation 86e77089-5ef2-4b20-ad58-6bfb328ad5cc · outbound

This paper cites Alignsae: Concept-aligned sparse autoencoders, 2026.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Alignsae: Concept-aligned sparse autoencoders, 2026

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-05T04:53:13.621870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.111636Z digest=sha256:e718e80c288bff6b8adf26f52e283f7b79cce784058bbfe111e046743388f248

Observation bf1166f2-560f-4ee7-86c6-b287bfd71e1e · outbound

This paper cites Can visual input be compressed? a visual token compression benchmark for large multimodal models, 2025.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Can visual input be compressed? a visual token compression benchmark for large multimodal models, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.138245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.138245Z digest=sha256:008fbe9aaea4b3dd4f8c0ab724673e07cd1f6e34f4a03a98e172b81f450d2172

Observation 33ee0171-b0d2-45d6-a8c8-35bc4f89def1 · outbound

This paper cites Ascd: Attention-steerable contrastive decoding for reducing hallucination in mllm.Proceedings of the AAAI Conference on Artificial Intelligence, 40(12): 10306–10314, Mar.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Ascd: Attention-steerable contrastive decoding for reducing hallucination in mllm.Proceedings of the AAAI Conference on Artificial Intelligence, 40(12): 10306–10314, Mar

Reference 40

Resolution
verified exact
doi, observed 2026-08-05T04:53:13.238886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.120867Z digest=sha256:ff7ab24164fd72753e563b5eaa96faaef64514e63da9dac91f271b7f05f8c87a

Observation 9f59fc7c-c5ff-4cb9-9a09-f4498dc40694 · outbound

This paper cites Let’s verify step by step.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Let’s verify step by step

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.121834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.148046Z digest=sha256:2f81e905464ecc6fec1f17d18cbf55db16d013638aa3c0060ef216a34f5b31c4

Observation 7480c158-ceb8-4a12-b726-2534a5dde6e2 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Star: Bootstrapping reasoning with reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.152614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.152614Z digest=sha256:800dc3f145a6dc8b864a28b6349233c53e0efbef46b9230d315d623084ddf4f3

Observation f23d5009-9e53-4424-922c-c96b54579004 · outbound

This paper cites Backdoor cleaning without external guidance in MLLM fine-tuning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Backdoor cleaning without external guidance in MLLM fine-tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.138677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.133448Z digest=sha256:7850aaed7649129b0848c7066c03b51385153092272bcf3bdb84c053dd3684e8

Observation d32c9b7a-0a3b-45a1-8917-2b78a3643fa0 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Minillm: Knowledge distillation of large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.074366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.161276Z digest=sha256:b2814241c2b382779e5c34d4165f9cb1fb8a85ff0b4a73ee6ebeef44823fadeb

Observation a51bb7c4-78cd-4b2f-99bc-40b975e53ee3 · outbound

This paper cites MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:13.142470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.142470Z digest=sha256:d10937a0eb3955deadd1338e8c35f2484d8c5be1c303f815cb2b3a811eae9ead

Observation 74b080b8-944d-4d8a-a72b-534a39c68219 · outbound

This paper cites Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.092400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.156977Z digest=sha256:752bc88cd5d6be650aaa6acc9e566e45d77db5dfe7d27758f566a6dced5abcdc

Observation 76c6f394-d6e1-4557-8a99-8089f7dcf0cb · outbound

This paper cites Generating Sequences by Learning to Self-Correct.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Generating Sequences by Learning to Self-Correct

Reference 50

Resolution
malformed identifier
no resolver link, observed 2026-08-05T04:53:13.165423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:13.165423Z digest=sha256:c8195597101df6be5d33f3bc3e95d21c2d1df26d836a288468ea83c641e50dfa

Observation 4a368c62-a29d-4e6e-9f93-ecdb5d9c6a58 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.057496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.170212Z digest=sha256:d8d68d7ea765ac4b7222edf897fa327998ea9a4c79b62f6ad020aeeb328d5877

Observation cc02cac2-914a-4afa-830f-3f0e1a453564 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.039323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.175101Z digest=sha256:c8144f009b95a10bbf40d13b9ab9d8b727945e7a09f2b841de604b94b4dd14ac

Observation e82dd163-107b-494f-8a7d-0843a8f6677a · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.021601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.179774Z digest=sha256:f2e2b7004c175311ca89111ac51ad14c9f2319ce18ca66da9dfdaf874324b3b2

Observation 615e3fea-8876-41e4-b695-975684456f7a · outbound

This paper cites 17 Figure 10Case Study: Reflective Reasoning (Success).

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning 17 Figure 10Case Study: Reflective Reasoning (Success)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.002582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.184559Z digest=sha256:8550768184c374bfe3cd50686876aa9026d980cae850f98dffb6335b90b3d0e1

Observation 165e5c50-9541-4901-b645-34bffda28b24 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:13.985398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.190617Z digest=sha256:9e8f1db138f2085e77a40676b50ddced3bf6c198a6d207f48be0f0456a3df91a

Observation 185dcd77-df9f-48aa-b89b-249267a1f57f · outbound

This paper cites The reference solution suggests there were some errors in the previous attempts. Recomputing:.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning The reference solution suggests there were some errors in the previous attempts. Recomputing:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:13.968802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.195296Z digest=sha256:a1d18d74b7d22290e5ad7e55c27a6502c88955482311f91dcf6d167f13ee2f4c

Observation c1872307-fff5-456f-8650-4dfb8e02e1e9 · outbound

This paper cites an unresolved cited work.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Unresolved cited work

Reference 483

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:53:14.461761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:12.962737Z digest=sha256:8749385eaf64fdcc35fb7295911477de9eff0d3b9fc20c526c1d9b426e64d5f7

Observation edd914fa-d6f8-45f0-8048-8e58b2befa49 · outbound

This paper cites https://thinkingmachines.ai/blog/on-policy-distillation.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning https://thinkingmachines.ai/blog/on-policy-distillation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:12.971795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:12.971795Z digest=sha256:0a317af69d3367d937750e0caa4472ccbb96ea31c6d51593df892c33deecd7ca

Observation 3fa1c871-1e68-4fbe-a32c-6e00ee3212e2 · outbound

This paper cites URLhttps://openreview.net/forum?id=9SkkifLopZ.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning URLhttps://openreview.net/forum?id=9SkkifLopZ

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:14.253295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T04:53:13.042559Z digest=sha256:90c6a2bb9c704e303a599cc80066fab1daedcaee2d846287c5c6156090474423

Pith citing papers

Observation ca766d31-089c-4d82-a559-29b1bcc5222c · inbound

OPD-V: Visual On-Policy Self-Distillation with Modality Balance cites this paper.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:47:42.408371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T16:47:41.755414Z digest=sha256:f03de8ceaa0793a6d128452f9f51ba334e577230c701501e22c59a4096e67793