Pith. sign in

Paper Citation Record · LEDGER

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

As of 17 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 13 inbound Pith citation observations for arXiv:2505.15804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15804 v3

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.409411Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.340790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.322058Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcdf50c0-212e-472b-95d0-367bc152ff3f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.818259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.818259Z digest=sha256:04c2c027541d108e13666f8125fb348ec01a0187f6e62e1e7aea60b629c712a1

Observation 841067a9-ed1a-46ba-a0ee-2455528740e4 · outbound

This paper cites Pixtral 12B.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.872572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.872572Z digest=sha256:dff7aa4f3a526cc1130bb7554c9becc1d099bd68e6b3039910ebc0d288b94ea6

Observation ca00110c-34f1-4ad6-9275-c01dd8fc5f2f · outbound

This paper cites Qwen Technical Report.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.949358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.949358Z digest=sha256:1f2018d409c44d9422e8a2413e86c56f8f6b56937a9e1b4c084ccde5e9ef198b

Observation ea2d13b7-7fab-409e-bd83-53ac10ced950 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.008408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.008408Z digest=sha256:b5c07cbe1e61e3e3b4f3ae23e81552cd0f925d3a6ea9bad11f8d1bf46755db06

Observation 212cd02f-0c13-4f1d-adda-1d442496413e · outbound

This paper cites Qwen2.5-VL Technical Report.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.082457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.082457Z digest=sha256:5673d964f19638f2357741a7b7197a95b9006adf9956781aa6268c777f76d6eb

Observation bd785b30-6e62-4a18-8f3d-bdbe30a01885 · outbound

This paper cites R1-v: Reinforcing super generaliza- tion ability in vision-language models with less than $3.https://github.com/Deep-Agent/ R1-V, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs R1-v: Reinforcing super generaliza- tion ability in vision-language models with less than $3.https://github.com/Deep-Agent/ R1-V, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.132137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.132137Z digest=sha256:c5b08ce61124dd5a931adc60efac0629dce395f4cfbf5846eda992ab7c6bcba6

Observation c57b4766-fc0e-46f8-b0e5-4965091fba79 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.240256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.240256Z digest=sha256:836defca7703869755e755746dfc2a045b319b7615f2b77533e1d21b55ff21a0

Observation 7fadc10a-2785-49c0-aedb-977e046774f4 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.254728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.254728Z digest=sha256:065186ca61abf2cf79cb8746b3957bd2ca67eaac14b4c877a748d6d11df9af79

Observation b1d57b38-53f5-49c4-a16d-20c220232ee3 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.300013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.300013Z digest=sha256:90a2438e46bddbcbae116880c867d59fc72eb9bd06e59d52bc02093e9d32d2f8

Observation 675891a1-c4a2-47a4-9012-e0581005d856 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.309847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.309847Z digest=sha256:92593f81f28d950bea75536fc82b65bcd78a7ce3afbfc088acee2f4e95d04b21

Observation a081f643-7ecc-473a-9653-2dedd8b9ab6a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.364691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.364691Z digest=sha256:29125f77b8c66eb711ca3c939ff62e8ec75eb3d5b40cf918b210623549ca960c

Observation 293ee2da-eb8e-4f1b-82cc-912464bb340f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.409185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.409185Z digest=sha256:224a5188802c449b19971faa45f4f816ff13ac893ee3a47faf443632a984755e

Observation a4f0255e-e449-42b3-8f9f-641679fe6bcd · outbound

This paper cites Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Sophiavl-r1: Reinforcing mllms reasoning with thinking reward

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.466292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.466292Z digest=sha256:8b9da608e94ed7f63cb3a5f8ad4a1c61cf654ea12ee661b8ca54da94fd17e4a9

Observation 1901736a-b61f-43ec-9751-af79e224f2da · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.500371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.500371Z digest=sha256:ecaefe00f0074c3cce54a0cfe21a0a34235cf11d7a16269c8df41aedf91f4a2c

Observation 50844e53-ff55-471e-b50c-1fd8da55cd7c · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.535470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.535470Z digest=sha256:fd49d6202a86ed7c61b86cf91f46e9a60fae9490bc3d704cee1f845304461e61

Observation 5d73b1f5-bc43-45ea-8adc-38778701ecfd · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.573252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.573252Z digest=sha256:11251babf377bbea4dc1a93452dbe15ee3775e2a5cb1e23267637a919213df2d

Observation f68052e0-3672-428a-a044-8e74c7dce9f9 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.605318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.605318Z digest=sha256:570e0117d89a7e1eef989ef0a1d5fd2ed710f5b2260200204aba2c689523fe18

Observation 37009068-0b55-4422-bbfa-5d26f070f066 · outbound

This paper cites The Llama 3 Herd of Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.644523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.644523Z digest=sha256:1ad4a048065f96ab416f68b513f02ee146b073529bda609ac1f3dcdf2cbe5060

Observation 9e9b71cb-00c7-43cd-b130-f5b43043dd0e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.685199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.685199Z digest=sha256:6ff7712ba7e028e0e0b17721e77d8f56966978c136f8633a6d4b7b39ba7ab4d2

Observation a7e2a54a-18fa-44ae-8c03-ba17156ca081 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.719514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.719514Z digest=sha256:1684467dea4f999ce61798969f91a2e74d915bd665236ca53cd623a1fe24c755

Observation 60642aa0-19f7-4be2-9edb-31f111399fa1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.755343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.755343Z digest=sha256:e493f3e6d4afcaabaf8e5d62a2c9a60d7fda5353d36bf4971b07e9e416bd3585

Observation 31f0ea88-2bad-4749-9571-c93c37a31657 · outbound

This paper cites Transformation driven visual reasoning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Transformation driven visual reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.794764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.794764Z digest=sha256:f74b0367ef6274dbceafece827a4a69c5afc0b35a22dc10ae6359cdfcf6a58c4

Observation 917f6bee-c864-4648-aedc-76a5bb7d8451 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.837834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.837834Z digest=sha256:187b0d118e095013476242cc9720ab60b7ddf7a0562e436027da0d1e8bad7bd0

Observation f85bfca2-c3c0-4ca3-addc-00913ef40ebb · outbound

This paper cites GPT-4o System Card.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs GPT-4o System Card

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.877810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.877810Z digest=sha256:d61a6eb7e80f5ed50802213d2ef32a862b801da1df720d02e622f522b5f0d65a

Observation d30f8d17-a1a3-411d-9322-29d2704458c1 · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.919164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.919164Z digest=sha256:488a2dee78b5b40854193fdb08e77e7dd05d0dea584c76d83ad5e0e791a3d20a

Observation 20186f00-b6df-410e-90fb-6bb84a7c64b3 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.956284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.956284Z digest=sha256:aa18df4f50a8cd597df357ab7d60c127dce1ac22e795fd9acbde0e5afce66565

Observation 0a109d7e-59ea-4cba-8521-f65ccbe551f3 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.027467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.027467Z digest=sha256:dc41a83aea2ca4f7fae40383eba291b6e159ea05fd992c209603f4d9fc812c6c

Observation 5e666ddb-be36-425a-86fd-cf359b6216de · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.075306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.075306Z digest=sha256:86f14830aded5fb58a92e310103727a752f765f82a169259ec84a8117706cdd3

Observation d582b926-e633-4030-af2b-ea3b7dd4edb3 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.096080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.096080Z digest=sha256:91cec2f02d41603267d3460fe5693be6918178fecc35fbc39512af8933f10e21

Observation caddb07a-2955-4f0f-b28c-415281230c70 · outbound

This paper cites Temporal Sampling for Forgotten Reasoning in LLMs.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Temporal Sampling for Forgotten Reasoning in LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.129693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.129693Z digest=sha256:c1acfac0f86843d098dd97fd86e5cdd4a51ac8ac37c1098057f43eb8f44070f0

Observation fa3718f8-5a0e-4349-bf35-85bc59e9dbb5 · outbound

This paper cites Sws: Self-aware weakness-driven problem synthesis in reinforcement learning for llm reasoning, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Sws: Self-aware weakness-driven problem synthesis in reinforcement learning for llm reasoning, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.134756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.134756Z digest=sha256:96d3854d5f913293c4b804ea06e4afc9099264d50b84ebd1a29f808a479f11c8

Observation 44d8611a-5f30-48b5-bde4-b16373eea687 · outbound

This paper cites Improved baselines with visual instruction tuning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Improved baselines with visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.139301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.139301Z digest=sha256:9f7bf4fb0c0ed688043f29d103c8c7ff572854c5ecb5a2682849d5e23f4021c8

Observation eec8507e-f0c7-4735-a8fe-825ecadb582b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.143883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.143883Z digest=sha256:f84cadf7291611675183a19353c57809822a92f55dc940d1d7ed4223b5afb37a

Observation 58f1947f-c76c-4444-8630-cb9fb76e6631 · outbound

This paper cites IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.147646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.147646Z digest=sha256:899f052ea687e09bf31530e749b946f0668a2ea3f58de3cae867e7a0e25fbef4

Observation fdf7e325-6210-4536-bd71-a9c819e8432e · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.152250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.152250Z digest=sha256:99f62d899597b2141c39732983cf064639b718976f2b9455e12656dd13f8a16e

Observation 25d24bf1-19d6-4345-88ee-c4c100b35c1f · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.156102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.156102Z digest=sha256:fcc346fd57fdb799c6251dd8dd66954420efbfb831878f49e7ab75e83b8200e0

Observation 0c8831b1-3d61-4d0e-8dc8-4c15ce522f78 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Docvqa: A dataset for vqa on document images

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.160189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.160189Z digest=sha256:85394ca5f5a76ce06a618cef05939b3e4df3a3f56b95bf75bb348f03c4e65d33

Observation cb33604e-ab4d-4479-992f-5a4abef9a15e · outbound

This paper cites ivispar– an interactive visual-spatial reasoning benchmark for vlms.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs ivispar– an interactive visual-spatial reasoning benchmark for vlms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.164590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.164590Z digest=sha256:ab7265119a22a3d298aaca3a1adb25e5b9c43b020f694e069e2b503a7141a7c9

Observation f5b910e8-6cbf-4963-97ab-65f3af45f796 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.168926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.168926Z digest=sha256:0b446cb3fdc947f122b69a94de5894531c455260e18550cb8644ce4a6b6df228

Observation e0443644-5b99-45d8-aca1-fcc9c1fdd016 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.173319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.173319Z digest=sha256:552a68a03a6e1354cd38b4ae38921e0f59fd3f13679397d6c6d1a45c0bce1c36

Observation 2bf1dcc4-e754-41a0-83dd-12fce5eed07f · outbound

This paper cites Learning transferable visual models from natural language supervision.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Learning transferable visual models from natural language supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.177791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.177791Z digest=sha256:feb4543eb696bfb8f25094b505093ca794c096a9bed4e39c086f30cae253a4e8

Observation 735331ce-7784-402a-accb-e2dadbf6ddb1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.182014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.182014Z digest=sha256:e6a9789e4e1bc5645bade90c0697044de05e6a35a3ef00f908c5967944170c12

Observation c8e46b05-bd85-4e0f-b50b-6367451fa9d2 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.185632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.185632Z digest=sha256:87fe2406a29d6ae65bb76b6d157bc964f63f493ffec61f53aaf69f997d4043d2

Observation dd0adcd0-9995-410b-bd76-a4e7d4c46316 · outbound

This paper cites Maniplvm-r1: Reinforcement learning for reasoning in embodied manipulation with large vision-language models, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Maniplvm-r1: Reinforcement learning for reasoning in embodied manipulation with large vision-language models, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.189660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.189660Z digest=sha256:83852ffd2ea6e1db730705426a59b52d2b9109d2dd859355ed8b37a3feb394e5

Observation e12c1220-c471-427b-a623-b9aab9688493 · outbound

This paper cites Generative multimodal models are in-context learners.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Generative multimodal models are in-context learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.193280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.193280Z digest=sha256:29867cb19cb6246ba19d114e931c6c741d082a3c98a1e53850b8291c73b2f0eb

Observation 1cb7b0e2-9c83-43a7-9cce-1d876cd68385 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.196929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.196929Z digest=sha256:b487049dbbb84c0c6d65e55f6d8d4e6d0643f3cc5003fe01353f55db00e3eb96

Observation bcf325dd-e680-401c-9b78-362e0746a881 · outbound

This paper cites LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.201162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.201162Z digest=sha256:ea7b76fb84b052e057e94008d6113980cad09f57520e999bb8f7b39d291b5b0a

Observation 44a1aefd-278f-4292-b814-a934fda0d30e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.205991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.205991Z digest=sha256:7dfabf2e63773d53b16baadaf0b2a0cd8e96bdbc0880f1e67014cc71b705fcc4

Observation 4d813ab3-513b-41b2-98d7-a677c7f9c0d0 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.210026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.210026Z digest=sha256:aea5029edd3d74623a9f88d9bd2b1a328ff81ec09e7d09e912b392e5baf88c84

Observation 8b76ae6b-8b34-42b9-8a24-c21f19a89c97 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.214143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.214143Z digest=sha256:6eef919f08b6959c829a7a8c291ab48655a2eb5c4ffe98e93ab73717d9abb937

Observation 95cbe8c0-37a0-440d-834d-f0fb00ad8cf2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.218603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.218603Z digest=sha256:e245d793529bea3c6ac9254a332a5cfe7cbbafadab6f237ec2170119f87d730c

Observation 489e5f84-5ebd-474e-ba7e-77519bc44865 · outbound

This paper cites Solidgeo: Measuring multimodal spatial math reasoning in solid geometry, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Solidgeo: Measuring multimodal spatial math reasoning in solid geometry, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.631470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.222331Z digest=sha256:6dd19a175bb352558961038167408ed24df31b47606d031e9386be3d5c29bb14

Observation dc84bc8d-8da3-4aa6-a097-e5ddbc760436 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.226384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.226384Z digest=sha256:908512eb8856be8a711e02407f77233b0178adc5840268c664d3fa6ba8c7b82f

Observation 444d3501-957b-412e-94e2-d863ecb8ec32 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LVBench: An Extreme Long Video Understanding Benchmark

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.230745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.230745Z digest=sha256:1275815698efd2a143bb17d4f124940d66a2a5cf98df60988c61a4a2533cd748

Observation 232b0cc2-325e-4178-9a77-6c71c261ab59 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Emu3: Next-Token Prediction is All You Need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.234497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.234497Z digest=sha256:2eafb2c18f0005b2c954807e2aae3e31b6d66479b4a9f59423cfe8214b0c384a

Observation b6686393-a273-418f-857c-304c2427e61a · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.243110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.243110Z digest=sha256:87983c7a297b9a8421ecb07ebb01832184fa024b6cb22dd9943717a76a78953a

Observation 164c55ef-ee9c-4dad-b383-3b7f377d05c5 · outbound

This paper cites CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.248418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.248418Z digest=sha256:0c433b6803ab53396e078c052bc42307a9657d0b48d96f2cb0a91bbbb4e4d1ff

Observation 0d081f85-7c69-48b1-8c65-71e19db85da3 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.253671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.253671Z digest=sha256:7043b8d5a8b721255ec894509c11adbd89736321b5539a6fb6d8a2b33e940d35

Observation e8c930ac-37a6-4783-9ef2-ee21b8180066 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.258330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.258330Z digest=sha256:db5c20e7dcf73182e728279a982d19af78a2933a16673abb2ad02dd3f61d5cb5

Observation 51606719-fb28-47d4-ad15-36ff7ef06158 · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.262101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.262101Z digest=sha256:6b75a2e3019e19475724b984c2237e01a64d91ebb9d7a7858cdd549dae4b14d6

Observation ddc43479-a4c1-4e89-9009-4e231dfb0a77 · outbound

This paper cites Qwen2.5 Technical Report.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen2.5 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.266320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.266320Z digest=sha256:89cbb5d9bec0b10357c9b39b702e1719e4658b471de73bf0507cd6978d1e9c47

Observation a6d0a0dc-20a9-4b37-b7ad-11ab42174270 · outbound

This paper cites DeepCritic: Deliberate Critique with Large Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepCritic: Deliberate Critique with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.269555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.269555Z digest=sha256:fad8b680f0b7bfdf8c8fa0149a526b703e42be18449f164b2902e5372624a3ae

Observation 35e4b4e5-238b-4731-a32d-1a80201f4aea · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Towards thinking-optimal scaling of test-time compute for llm reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.273216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.273216Z digest=sha256:7bdc7f2886dcce353b245dce8ee6b021d5cfc50a7ebb63b09abf775e9fdb8344

Observation 901e1fc9-dabd-4011-8231-b168b2820735 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.281622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.281622Z digest=sha256:b3de8d7600a584d1c554a5567b7a20e231f8db8d9ad250bbdc3918a3b2371a27

Observation e23c3376-8607-4cfd-b530-4cdb6610b899 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.285220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.285220Z digest=sha256:00066177ae6f9e7261032077fdd7cf1dfe8515896032539910fb4d3c15a3047b

Observation 519b80c0-16a8-4165-b300-b17059383537 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.289839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.289839Z digest=sha256:ad78fb2279a54b2ca6c03797c882e18833326e54c11d541e772f183be01fdc0d

Observation e1520b89-e6ad-4e5d-95f4-2128216ee265 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.293769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.293769Z digest=sha256:e60438fd286f0228b7ec0dee778bd54e523b9eecffd2a58b11ba36ca5210a944

Observation ba8898d5-c2c4-4162-b806-bda2c55114b5 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.297149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.297149Z digest=sha256:38fe77a03fcac5f0632cc26f056f6478a6b0db13af8ae1bf86e42f7d7aec5b5a

Observation 4d24babf-3c90-4ee8-8d36-656b4965946f · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.301134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.301134Z digest=sha256:5b7021d8a72cd45d9439ba6514435f803bd49301dad5d638d9f476c3aba6e9de

Observation f67027c9-4458-48db-9a08-8c0e0be4206f · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.305127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.305127Z digest=sha256:7f2f8ef65b6d03374c2b8b0a70382d8efccd8a6a251eea252d4d994fd6c3d27f

Observation 096c5082-9f0f-4141-9b5d-e7cca97dc84f · outbound

This paper cites We employ Qwen2.5-VL-7B as our base model and utilize vLLM [26] as the training framework.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs We employ Qwen2.5-VL-7B as our base model and utilize vLLM [26] as the training framework

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.605647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.309967Z digest=sha256:73274ddb0fac7627a356d09e61c29e8aacf570f12e28682dd6bccc84b333ca8c

Observation 244a5890-5c75-4fe9-899d-5cf984edde73 · outbound

This paper cites cube", "sphere.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cube", "sphere

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.592351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.314386Z digest=sha256:cb420e97cecb46cc3b4b7a01eaba811b0632e63b7e77d8e768ebe419420fb250

Observation db0c6a97-a730-43a0-a38f-f16d909cd050 · outbound

This paper cites large" (radius = 6),.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs large" (radius = 6),

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.581594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.318647Z digest=sha256:e0be782ed07b5f04753bc8765c7ccb5761fd15dc310bb5356efb3de921fdbbc9

Observation 18b1c8d7-ee0a-4965-9833-d33fe965697d · outbound

This paper cites gray", "red.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs gray", "red

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.571255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.322551Z digest=sha256:fd966c1eed33ac89e486bc22958f7f2100c0ba5daabcc16bb3a254d12b10607f

Observation c323fc22-ff28-4c8c-91f5-c36482bf9bc0 · outbound

This paper cites rubber",.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs rubber",

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.560229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.326696Z digest=sha256:9c1a43436d51b07b25c3c395edb522bc9c6d8903a64c1b1467a0747d5239246e

Observation 873d6753-d86f-4d12-84b2-2cbb8531affb · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.548789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.331251Z digest=sha256:b8aaec91cf4ceadcf956d3a348b0fb89be266f66a71ff3a879af89ee9186f8ea

Observation 11a8e450-aaaf-4057-9dbc-be23bb60b513 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.537164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.335528Z digest=sha256:e8d3cd4f4e4258deabf65c37056efe02d8f08983ebc58f272755851f83509764

Observation 0f7bfbb6-b6d6-44cc-870d-d6db07df822b · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.525219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.339882Z digest=sha256:48547ea73e8271ea1e1cfd4b0ca7f5f12cd1f988dd0a2d936a2c24a0a0c46142

Observation 065dac49-2284-4996-aa2a-0163cf1d7964 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.513574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.344086Z digest=sha256:a58736d3917da4a20a6334eb2c5c72715ce24df01c7cfb1c659377f51201c8ec

Observation 4a827d13-3c5d-461a-9659-77a7cccc9283 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.496484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.347883Z digest=sha256:8af35768ba7e31248fc63890d3b13dfc48ffd544475c2494d2835c08ad06ad4d

Observation e93abd86-47ce-4df2-8ffb-6a464d6f5781 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.483698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.351668Z digest=sha256:73e5ede9e275d5da6e49d3ecd7324881f2072795eae6893f0d1a72cfdecbb25b

Observation e66144b6-0ba6-492e-b06e-c3d0a3b6f0b9 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.471860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.355256Z digest=sha256:cf62224e400cce0ccaf21a034d0f2315382cead64e2ae714a289b6bafb7ed8c0

Observation cfb56c9a-21dc-4c78-a0d7-f11722c27f22 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.460565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.359404Z digest=sha256:83852996cccc56dcaa8dab3003004eba4d8e1e6e0e707799d50380f73477474e

Observation ab463e98-fd8e-4f31-a007-2314901443fc · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.440298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.363581Z digest=sha256:dc556a99b7d173667355ac83eaaf3898309ec20102258292a73e0a18e6909a8a

Observation 29dbecf5-bbf3-42fb-9721-69d2c6ec808e · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.426319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.367119Z digest=sha256:ddf200ccc6024a51c4af3fe1feda449c86bbf3fe6e25afa5bc67149fd456660e

Observation 61be8813-fb52-4503-92fa-cd11e7f434f1 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.411313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.370418Z digest=sha256:67c3515cf40b131e245ade1b54a90e78ed44b02bc2de435b931d5dcd5356d501

Observation 183f12b8-9608-4408-9b86-3b50e85ac1a1 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.397344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.374440Z digest=sha256:140c9bbd2a083fc9328f0dcb1b82d9976f1ccf830850884edc27d4a8bcda6a92

Observation d5820b93-e0f5-4ddf-a0a6-ed01e031abe1 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.382882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.378262Z digest=sha256:4bd5aa81bd79b86571a1644ec2eb43a7950388d468b1cbaad6b193d99352cdde

Observation a87e34da-cf3f-4178-8d74-735548c44d4d · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.371422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.381824Z digest=sha256:8af407e45d460df9e559da9e6d62d16193cafaf8a7f63cbee44842a59331ede2

Observation 9c93ac4e-86e4-4858-8a80-58a2ab6fa953 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.359597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.386297Z digest=sha256:eb8267c544f3b0f068a7102bbdeaee8c77b13b0876e20f442a5511e16c9c1e5e

Observation ca901aa9-f866-445d-be0f-d3bf6cc630ed · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.347914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.390123Z digest=sha256:12bb04c82ca368257823f3118c0c63f6167c6a8fc4d1ec8cfc246b89bc87ed84

Observation 808d1f98-d957-4809-8e18-8245b27918ed · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.334435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.393717Z digest=sha256:9768810f4efa2ed4c947154a3c0b5f5ebba64aee687dbb7de9a2d3bc18cf1c48

Observation 46eb0fbf-5a8a-4bda-aa91-b8eae4758204 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.319013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.397927Z digest=sha256:60f5366ed49f7dac535e2822b5b68440ac3f5ec40057369e05e45e469c942b8b

Observation 20b72e1b-ea82-4c20-8442-c828b79e361e · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.307761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.402199Z digest=sha256:c1e1936b8a2bea9f0de6aea709341046688d85b844871b3d10531ada8f5a3e12

Observation 9a894983-d8ce-402e-b519-68e6b63834c0 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.294524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.405954Z digest=sha256:362526530647ef24310629e72c8be7a1d7076348819bd97fb620d2d1ba6af066

Observation 6335bd32-dda7-42fb-b6b2-ff2428698243 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.245130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:15:42.409411Z digest=sha256:dfbab74a1ddd8e9152f4a44ba7a093027fa02399db5d130b21c3d617d0a2d4c6

Pith citing papers

Observation 4cb3a065-2600-492c-a285-090f70fc77d3 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 297

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.521889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:58e8028abf9da6d64ef6eabcb335097e44090b066bee83d95e16a159ebf2d90e

Observation cad82a6b-df4e-4c32-8a85-f002afae1123 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.404330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:b3e34bcef014bee13a71b444328c0048bde80e43e8a2abb7fa07c559645885d0

Observation af4c9a95-d7ed-4833-b603-264f4541f191 · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 273

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:13.340790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:13.340790Z digest=sha256:29a4ad2be44bdc281ad50db561328868fce2b1232aeb87f30a3361a9a0f16957

Observation 23c52071-9374-4517-9e4b-b5b62d315cf2 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 233

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.322627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:69ae29b878d3f3fdb681204c7d3571877883d204653f2fc151a408d023d186b6

Observation c0c118ee-ba2c-4943-91d0-197b451d7758 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.690136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:0a7c50b126368f3330d825dcc4c9598953e7db32d1ea79104b360983633456dd

Observation c5ef8303-4565-4db1-a785-e21994e09dc6 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.499729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:f89bb5034d0d49c45d705c621509a5bb712ba7342c1cb0d7b93a21154fee95db

Observation 0b1a2ca8-b126-4ea0-a620-aba6a117bbea · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.413324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:63157f4c6bcbaa4ce66ac458e770893d192e098b511fd6579963d854f7561244

Observation f5fe546e-47f0-4f33-b2df-f9ef43d8d2e8 · inbound

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking cites this paper.

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T16:55:20.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:55:20.099628Z digest=sha256:93f8cdd08767829b54c65134d46b6eb4ab3b97accd180361c50caf91154f11cd

Observation e5b02838-3488-422e-b562-c080880d485d · inbound

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs cites this paper.

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.985595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T21:35:12.859669Z digest=sha256:16cd6696602dcda989bc76c62cef1371ca9e82c31382223d63f12a027065ba43

Observation 9ae53715-7a4f-4529-9a94-fbda2d4735b4 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:52.525638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:1e10cd5cd002cee17ef302bdb5f6ca19d495427b121a7e94508f41434615ca40

Observation dd904fd6-2380-4591-afc2-19ec70eb6245 · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.573891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:dd8df465495f9d053b376b7272f217461c6d727d630d871330ff7d431255f26b

Observation a1bcdc9a-4721-4290-a3e5-7842709a2325 · inbound

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval cites this paper.

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:50.543718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:42:23.224746Z digest=sha256:04d17ab14f94103c40cd9b02949304238bb021cbf0417bc59c33421ba42de95c

Observation 76164de6-9435-4e6a-a1ef-2695f34cf3de · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 192

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.323884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:f9b790eb82ac2362e17315993e335a33c19e2b67e1edf3792c4c62c6c901acb7