Pith. sign in

Paper Citation Record · LEDGER

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

As of 7 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 12 inbound Pith citation observations for arXiv:2505.15804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15804 v3

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.409411Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T16:55:20.099628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.322058Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved90
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcdf50c0-212e-472b-95d0-367bc152ff3f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.818259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.818259Z digest=sha256:c8faf9239b8595cb8224ae8f43d6164912e9b8076fa508f1d658eaacea0ebf36

Observation 841067a9-ed1a-46ba-a0ee-2455528740e4 · outbound

This paper cites Pixtral 12B.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.872572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.872572Z digest=sha256:92d75fe02a212a63ddf76035ae61514f5fd19d61f37363839c5073e7a12a262a

Observation ca00110c-34f1-4ad6-9275-c01dd8fc5f2f · outbound

This paper cites Qwen Technical Report.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:40.949358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:40.949358Z digest=sha256:9f9346b7b341b20fdab958b182fc279f59943528881120ecbf5b982e7fc27123

Observation ea2d13b7-7fab-409e-bd83-53ac10ced950 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.008408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.008408Z digest=sha256:e8de5c09d351623408cc621c79ddbabd44f2095e5e5c03577bc4d0b17165db7e

Observation 212cd02f-0c13-4f1d-adda-1d442496413e · outbound

This paper cites Qwen2.5-VL Technical Report.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.082457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.082457Z digest=sha256:3501b343c5b844c4fed114e07fe48ed105a58dfe1dbe273b0a4bb82865fa630c

Observation bd785b30-6e62-4a18-8f3d-bdbe30a01885 · outbound

This paper cites R1-v: Reinforcing super generaliza- tion ability in vision-language models with less than $3.https://github.com/Deep-Agent/ R1-V, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs R1-v: Reinforcing super generaliza- tion ability in vision-language models with less than $3.https://github.com/Deep-Agent/ R1-V, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.132137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.132137Z digest=sha256:a93ae6923369df2af62cd4f3e44591d80db92086bd53ba03131c25c8337ee76f

Observation c57b4766-fc0e-46f8-b0e5-4965091fba79 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.240256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.240256Z digest=sha256:596a0b29a1c550e3bca2586efca6265299ed6c147d42a474ba90c20ab502c610

Observation 7fadc10a-2785-49c0-aedb-977e046774f4 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.254728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.254728Z digest=sha256:ad067760f9350eb9506a55d50606ac8ef7e86a1fc087ea2c4b80f3a9ad4278e2

Observation b1d57b38-53f5-49c4-a16d-20c220232ee3 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.300013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.300013Z digest=sha256:6a6d7b568b5fd12bf66c055568946aa88df2e9b394e98ab1d8b71208f01cf5b9

Observation 675891a1-c4a2-47a4-9012-e0581005d856 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.309847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.309847Z digest=sha256:6446d43d588e297b686d2997ffbc3d61f67138d6edc447f6f0db554d1ae07d6c

Observation a081f643-7ecc-473a-9653-2dedd8b9ab6a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.364691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.364691Z digest=sha256:41d627425fec5376d3c87d77b395a566a3be367f2b10cd2443cb27455da5aee4

Observation 293ee2da-eb8e-4f1b-82cc-912464bb340f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.409185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.409185Z digest=sha256:8387f861ae0575073577e8433d867b031af12588273841c6016ce77ed525c1d8

Observation a4f0255e-e449-42b3-8f9f-641679fe6bcd · outbound

This paper cites Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Sophiavl-r1: Reinforcing mllms reasoning with thinking reward

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.466292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.466292Z digest=sha256:b03393b9aa4376da1d8d4118d0934e55d5174cadd13ee170ffbce933ed0e2aca

Observation 1901736a-b61f-43ec-9751-af79e224f2da · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.500371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.500371Z digest=sha256:009509a6ce1ed042e1c858221f3dc9f744806735704c0c3589b9f164ba21750b

Observation 50844e53-ff55-471e-b50c-1fd8da55cd7c · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.535470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.535470Z digest=sha256:b26cd38fd8391610560dd0337db38195f956cf148f564217d91889176d7f7dd8

Observation 5d73b1f5-bc43-45ea-8adc-38778701ecfd · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.573252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.573252Z digest=sha256:5886151600b5b74788818624897d4684a946d5fb74d9b5960ea81d2481b447fa

Observation f68052e0-3672-428a-a044-8e74c7dce9f9 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.605318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.605318Z digest=sha256:e2af54964d3f4e88a9605cef6aa3e26e331e37fbbd9fe53ce8fb20da25148631

Observation 37009068-0b55-4422-bbfa-5d26f070f066 · outbound

This paper cites The Llama 3 Herd of Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.644523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.644523Z digest=sha256:53cb912ed9aa4f5c069110011cb8eb70d471f9b00522cacd520d35312f3211eb

Observation 9e9b71cb-00c7-43cd-b130-f5b43043dd0e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.685199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.685199Z digest=sha256:dbae49909b4903c96a74a4fdd7f28908bb318f0ca57667a282f08afb47753ffb

Observation a7e2a54a-18fa-44ae-8c03-ba17156ca081 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.719514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.719514Z digest=sha256:9d5450fab4aa0134544fab8b302b8d702ed0efba53c50b7b6c47332bf141ee5f

Observation 60642aa0-19f7-4be2-9edb-31f111399fa1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.755343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.755343Z digest=sha256:911d6e7024bc3f29c1171f3c803037429e89b067b4c3273d4da221fca6c86bb2

Observation 31f0ea88-2bad-4749-9571-c93c37a31657 · outbound

This paper cites Transformation driven visual reasoning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Transformation driven visual reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.794764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.794764Z digest=sha256:813c7c974c7bb247bfb641ac90d1f71b7040e7b8c658d7907ff90e7c92a7008a

Observation 917f6bee-c864-4648-aedc-76a5bb7d8451 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.837834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.837834Z digest=sha256:ed91ba7e09b3f1536659a4edf93a0525a8bedcc4df95cb2c8625d3022fb24905

Observation f85bfca2-c3c0-4ca3-addc-00913ef40ebb · outbound

This paper cites GPT-4o System Card.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs GPT-4o System Card

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.877810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.877810Z digest=sha256:967c8cfcfe85d32c956e0326b8d2831ea6c42d0aa1f5eed7fb3c4fd8a56fe757

Observation d30f8d17-a1a3-411d-9322-29d2704458c1 · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.919164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.919164Z digest=sha256:8ad2e53a05cfa4c552e357fefd1b2c35558a6f974fae2616ef8bff895beb3b42

Observation 20186f00-b6df-410e-90fb-6bb84a7c64b3 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.956284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.956284Z digest=sha256:51c3ecf11bd24342a2e814a6544a52f6b97794a37a881022f1e1bb140c633ab3

Observation 0a109d7e-59ea-4cba-8521-f65ccbe551f3 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.027467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.027467Z digest=sha256:c83b5777746fa82e505c4fa47733264a401b09cf7e155938622a6b4064eb2bb3

Observation 5e666ddb-be36-425a-86fd-cf359b6216de · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.075306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.075306Z digest=sha256:e3a32aaab8151bed7f34271dadd96f20fe4ee430a3456e66942ef4a78a1a3d1a

Observation d582b926-e633-4030-af2b-ea3b7dd4edb3 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.096080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.096080Z digest=sha256:868b30134a65f59fd6a5d26dfa71220ba52979691980130c30ab206c90a0a424

Observation caddb07a-2955-4f0f-b28c-415281230c70 · outbound

This paper cites Temporal Sampling for Forgotten Reasoning in LLMs.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Temporal Sampling for Forgotten Reasoning in LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.129693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.129693Z digest=sha256:e48207d48c148f24814217f7bdbc341023ed7e0b4007ede6139f9d084ca4f4c5

Observation fa3718f8-5a0e-4349-bf35-85bc59e9dbb5 · outbound

This paper cites Sws: Self-aware weakness-driven problem synthesis in reinforcement learning for llm reasoning, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Sws: Self-aware weakness-driven problem synthesis in reinforcement learning for llm reasoning, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.134756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.134756Z digest=sha256:f47430fe2493db7d9488f12fd27eea976f70c151b4c2bd1cc8035d60e8f2d352

Observation 44d8611a-5f30-48b5-bde4-b16373eea687 · outbound

This paper cites Improved baselines with visual instruction tuning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Improved baselines with visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.139301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.139301Z digest=sha256:114e0911b4dc2269c79e7cf5765f9698d7610f9c4a74d167fd65f7d49191d57c

Observation eec8507e-f0c7-4735-a8fe-825ecadb582b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.143883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.143883Z digest=sha256:9600d93a2e0b74197a0865aa218a882a8fb749715abd47542cbc2aa051bf9eb9

Observation 58f1947f-c76c-4444-8630-cb9fb76e6631 · outbound

This paper cites IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.147646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.147646Z digest=sha256:be7f2e8e7fdf8bf5cbbaa3f3e9e711ff1a3990cf4ef5cb491d4cd50fadbdb542

Observation fdf7e325-6210-4536-bd71-a9c819e8432e · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.152250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.152250Z digest=sha256:80867ed7b5cb36db6dbb5dcd02bbcbe943aea891f0084527c19b0cc931bcc72b

Observation 25d24bf1-19d6-4345-88ee-c4c100b35c1f · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.156102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.156102Z digest=sha256:1dfffece4be838e1716a045588726520ee69893fd5bc3e1949880e53d88bb9db

Observation 0c8831b1-3d61-4d0e-8dc8-4c15ce522f78 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Docvqa: A dataset for vqa on document images

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.160189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.160189Z digest=sha256:74b28895e46023601d5bfc53b16bf4c00ede8400e9738bace25d7ecfca056fad

Observation cb33604e-ab4d-4479-992f-5a4abef9a15e · outbound

This paper cites ivispar– an interactive visual-spatial reasoning benchmark for vlms.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs ivispar– an interactive visual-spatial reasoning benchmark for vlms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.164590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.164590Z digest=sha256:8541a6620c14aae66f8aa9e6d7a18c211f362fc43cb77c64c252f26c7bb07545

Observation f5b910e8-6cbf-4963-97ab-65f3af45f796 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.168926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.168926Z digest=sha256:be52e69cc73415e88046181db0ad5cae40f5363d40c55c0ad03a0f95ea321ae2

Observation e0443644-5b99-45d8-aca1-fcc9c1fdd016 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.173319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.173319Z digest=sha256:87bc026c29e3eb9d84fc3fb18f975c3bdd2d173f60f7fcb6a711bba8c5b2e36e

Observation 2bf1dcc4-e754-41a0-83dd-12fce5eed07f · outbound

This paper cites Learning transferable visual models from natural language supervision.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Learning transferable visual models from natural language supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.177791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.177791Z digest=sha256:a6ecc9315efe45cb96e588fb271d6e80fdb53506261217b6e15cf4440988bfd9

Observation 735331ce-7784-402a-accb-e2dadbf6ddb1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.182014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.182014Z digest=sha256:cc1e77673b69a85f26ccd209c9f8a6f41d664d66d52804ce02a9577c92aa003e

Observation c8e46b05-bd85-4e0f-b50b-6367451fa9d2 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.185632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.185632Z digest=sha256:47fe10c8733344d87c21c30b03df216a859a644d34c8ba26722945f58dd05777

Observation dd0adcd0-9995-410b-bd76-a4e7d4c46316 · outbound

This paper cites Maniplvm-r1: Reinforcement learning for reasoning in embodied manipulation with large vision-language models, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Maniplvm-r1: Reinforcement learning for reasoning in embodied manipulation with large vision-language models, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.189660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.189660Z digest=sha256:ed638ee47522b1ad268f50835a2bc5fd3178cb630226e147e72cec90fb9ce969

Observation e12c1220-c471-427b-a623-b9aab9688493 · outbound

This paper cites Generative multimodal models are in-context learners.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Generative multimodal models are in-context learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.193280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.193280Z digest=sha256:63b6906bb2fc180392d3f54e2ec16d2333733e468c38482620997a462b868176

Observation 1cb7b0e2-9c83-43a7-9cce-1d876cd68385 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.196929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.196929Z digest=sha256:b4560e2b0bc2c8a1ec7dd99c7595e8e40b4e6d6cedfb5ec8f46461d2c1c00a8e

Observation bcf325dd-e680-401c-9b78-362e0746a881 · outbound

This paper cites LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.201162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.201162Z digest=sha256:cdb3d2c5d2402a9ac8f1cd77896a870c1606d75ff095debf8e35b624246d3266

Observation 44a1aefd-278f-4292-b814-a934fda0d30e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.205991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.205991Z digest=sha256:9a5b88be4a7ac121634bc7d36844c18dcf5e1c947761d92ea72fe32154810408

Observation 4d813ab3-513b-41b2-98d7-a677c7f9c0d0 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.210026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.210026Z digest=sha256:38af2e226cacffbef294381e2d97691d375b54edd518a068f8d350946f40318e

Observation 8b76ae6b-8b34-42b9-8a24-c21f19a89c97 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.214143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.214143Z digest=sha256:785e30bc1d8dee7754f49f1ef95f4a06cab4d2224a9b1b8558628f41c2e05176

Observation 95cbe8c0-37a0-440d-834d-f0fb00ad8cf2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.218603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.218603Z digest=sha256:1597223353aeceeed0041b4e809fd47e5290da600bc65362569dfe482d51a862

Observation 489e5f84-5ebd-474e-ba7e-77519bc44865 · outbound

This paper cites Solidgeo: Measuring multimodal spatial math reasoning in solid geometry, 2025.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Solidgeo: Measuring multimodal spatial math reasoning in solid geometry, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.631470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.222331Z digest=sha256:54ba036dd9e25aa357e11479acb488bb9a19488a6e6a15a56e87d456c9f5ac72

Observation dc84bc8d-8da3-4aa6-a097-e5ddbc760436 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.226384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.226384Z digest=sha256:5d3f5af6a0040ee6d42a806fabf252181da4b04c802b340972a4d3419f7220a8

Observation 444d3501-957b-412e-94e2-d863ecb8ec32 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LVBench: An Extreme Long Video Understanding Benchmark

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.230745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.230745Z digest=sha256:274fc29e01c669ffc3f9ae214124dd15b8c4ba83e013248ceec5028c5f4979da

Observation 232b0cc2-325e-4178-9a77-6c71c261ab59 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Emu3: Next-Token Prediction is All You Need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.234497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.234497Z digest=sha256:47f5df0d9470ead6cdbf6e677f4c4a54a6cb1f2eb1194437af85469b7ada2359

Observation b6686393-a273-418f-857c-304c2427e61a · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.243110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.243110Z digest=sha256:f6a2dada7f1ce983e59ba70cb8d6ed22208a5804990aa80429472ff85f2eea31

Observation 164c55ef-ee9c-4dad-b383-3b7f377d05c5 · outbound

This paper cites CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.248418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.248418Z digest=sha256:ecdb77a923d57fda68c982b8fec27e4b8120091475a984a4ff72ad92dce825fc

Observation 0d081f85-7c69-48b1-8c65-71e19db85da3 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.253671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.253671Z digest=sha256:8324ff7c8e7da2f7fb808d87880b359f693d834d4b1400b8173a4ba4334f95e9

Observation e8c930ac-37a6-4783-9ef2-ee21b8180066 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.258330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.258330Z digest=sha256:e6117d31927e0dc179aed517146adfa440b7aa8a138d061cd3ec55b544d22708

Observation 51606719-fb28-47d4-ad15-36ff7ef06158 · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.262101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.262101Z digest=sha256:1cfaa8ee65c8a7a0f1c87717f5fe629d9f578835141662cd1f52adccd092b8e9

Observation ddc43479-a4c1-4e89-9009-4e231dfb0a77 · outbound

This paper cites Qwen2.5 Technical Report.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Qwen2.5 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.266320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.266320Z digest=sha256:4ed9559fc7614a453155d84c6e6b2132ca1844c745fbaacc0c54800113b7e9f0

Observation a6d0a0dc-20a9-4b37-b7ad-11ab42174270 · outbound

This paper cites DeepCritic: Deliberate Critique with Large Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs DeepCritic: Deliberate Critique with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.269555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.269555Z digest=sha256:3fc836bdd328443c69f6afd2018a6e6b57e73571b4d6552bbca1ca248204a691

Observation 35e4b4e5-238b-4731-a32d-1a80201f4aea · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Towards thinking-optimal scaling of test-time compute for llm reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.273216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.273216Z digest=sha256:28c8a5467cb7460ece5a6d10589ba73958ba67dceddee0f0e2b32a459287e8b1

Observation 901e1fc9-dabd-4011-8231-b168b2820735 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.281622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.281622Z digest=sha256:ca44154e79353dd14bf873b1cd17d288b47e96bf40d172d3ded00c0717409787

Observation e23c3376-8607-4cfd-b530-4cdb6610b899 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.285220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.285220Z digest=sha256:0dd798879a61f688f0cc94f09d234afdddba9627498dd0decc6c4bf68fe59e7a

Observation 519b80c0-16a8-4165-b300-b17059383537 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.289839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.289839Z digest=sha256:e4ec37190e3fe364cd6714e7a5eb4159a60a193567658838da32a3ca65311886

Observation e1520b89-e6ad-4e5d-95f4-2128216ee265 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.293769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.293769Z digest=sha256:6652f764181ddc2c3d6e4a51ee34f9c0b8990e0ab2543ef531e711a85efe4812

Observation ba8898d5-c2c4-4162-b806-bda2c55114b5 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.297149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.297149Z digest=sha256:5408929c858b9cc0f8bcea6431c73e3d25c767c2d56db0bed01017a860986374

Observation 4d24babf-3c90-4ee8-8d36-656b4965946f · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.301134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.301134Z digest=sha256:86ca4acdd0936823e2b4358946a174b744b69cd181a1e1f7d7173403b0052b56

Observation f67027c9-4458-48db-9a08-8c0e0be4206f · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.305127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.305127Z digest=sha256:7c17ae11f7bb9110dd8058c6d4d8d12dfaf2ee5234d6d58d75571cce9d145d5d

Observation 096c5082-9f0f-4141-9b5d-e7cca97dc84f · outbound

This paper cites We employ Qwen2.5-VL-7B as our base model and utilize vLLM [26] as the training framework.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs We employ Qwen2.5-VL-7B as our base model and utilize vLLM [26] as the training framework

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.605647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.309967Z digest=sha256:907bfeb21f54afe90903d5fd54e92455f49f0012b512b47ffb171c3f300ddc9f

Observation 244a5890-5c75-4fe9-899d-5cf984edde73 · outbound

This paper cites cube", "sphere.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cube", "sphere

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.592351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.314386Z digest=sha256:14f3462e1245e680d760267cb2d6ae6f9106ba0dc959dda55022e8d333caa4db

Observation db0c6a97-a730-43a0-a38f-f16d909cd050 · outbound

This paper cites large" (radius = 6),.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs large" (radius = 6),

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.581594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.318647Z digest=sha256:2fb1a56aae657069d943096c372d1b9b6dd487ce31b7b971fd6bde42c3253e15

Observation 18b1c8d7-ee0a-4965-9833-d33fe965697d · outbound

This paper cites gray", "red.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs gray", "red

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.571255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.322551Z digest=sha256:988e9850f511d1ce42ad8c7a600d8531b07eec3f64b0a0b2e1800e889974d834

Observation c323fc22-ff28-4c8c-91f5-c36482bf9bc0 · outbound

This paper cites rubber",.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs rubber",

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:43.560229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.326696Z digest=sha256:38592587e95841e16d7019ec0e87428ffdbc6fbe23cb120edc9a9999d65409b3

Observation 873d6753-d86f-4d12-84b2-2cbb8531affb · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.548789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.331251Z digest=sha256:e807ff766209bd27b62f3907b25a1992e345c7feda44c5c7060be0837e16f7d2

Observation 11a8e450-aaaf-4057-9dbc-be23bb60b513 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.537164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.335528Z digest=sha256:6a93446f9874b051541aafc508bdc8d44a7bfdb34604c93ff04520e296551593

Observation 0f7bfbb6-b6d6-44cc-870d-d6db07df822b · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.525219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.339882Z digest=sha256:a37ab825c63bb3baf36815f0f62a474e66f9a1de64b3b78a3ec3234e070f3c09

Observation 065dac49-2284-4996-aa2a-0163cf1d7964 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.513574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.344086Z digest=sha256:eddbd9961943d7453c59fcd7f7157671e0ccd4672b7902a740e10471f06a3d7a

Observation 4a827d13-3c5d-461a-9659-77a7cccc9283 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.496484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.347883Z digest=sha256:16043b1da1e1eb06b2a0018e6f7e73691bb6d3ce1f2efd73d29cb11f6ccbe93b

Observation e93abd86-47ce-4df2-8ffb-6a464d6f5781 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.483698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.351668Z digest=sha256:0e5469136f9cb96dc3309098e8e16015f2e37d70a71bce6e74c769a5b97c6e47

Observation e66144b6-0ba6-492e-b06e-c3d0a3b6f0b9 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.471860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.355256Z digest=sha256:88d29102317ac9922c5d8ed4604ca9b6a859197fcbd9250a7394a45673743f1e

Observation cfb56c9a-21dc-4c78-a0d7-f11722c27f22 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.460565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.359404Z digest=sha256:027794e83af7eceb50c608a61d8942cfa34145a1361825e6d45b38cf846936d1

Observation ab463e98-fd8e-4f31-a007-2314901443fc · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.440298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.363581Z digest=sha256:b31b1d13a3d0c41bf40e27f635397d9df50a911a965a09529dedb87bc2a1103c

Observation 29dbecf5-bbf3-42fb-9721-69d2c6ec808e · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.426319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.367119Z digest=sha256:7025518dc87b19a60657c78d53eaca6a7638e0d3694523f2f709a2d629ae41b9

Observation 61be8813-fb52-4503-92fa-cd11e7f434f1 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.411313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.370418Z digest=sha256:b2cc81952cda983c0f32e60c46d3f8480d8b34c12ca4b1b336e109d8d58da0d3

Observation 183f12b8-9608-4408-9b86-3b50e85ac1a1 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.397344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.374440Z digest=sha256:49bdbd947a544da791940c27a8ada5a1adb48e5a37811975a02200f2b916a8c7

Observation d5820b93-e0f5-4ddf-a0a6-ed01e031abe1 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.382882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.378262Z digest=sha256:53c191cf3854f4fd18485edee0f392f743be9358832a45a9eec2bd492f69559f

Observation a87e34da-cf3f-4178-8d74-735548c44d4d · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.371422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.381824Z digest=sha256:f4587aecc3dbdad3f064ea783e9437f124a7a1aef734b76c9f9c1350d30d318d

Observation 9c93ac4e-86e4-4858-8a80-58a2ab6fa953 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.359597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.386297Z digest=sha256:b26de141179c435e2e1a50c4fa97b52396fbe3ec0d01d4c297054ad46daae537

Observation ca901aa9-f866-445d-be0f-d3bf6cc630ed · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.347914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.390123Z digest=sha256:5028e93598002cc244e1335440ea3cbbdbd82a1ab95992ddd5cacf48f8f21cb7

Observation 808d1f98-d957-4809-8e18-8245b27918ed · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.334435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.393717Z digest=sha256:512f3c3ba436203d312b86ff1be2bc6fda14770d4ae345248d8a46c57beaed18

Observation 46eb0fbf-5a8a-4bda-aa91-b8eae4758204 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.319013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.397927Z digest=sha256:59e58e430aa4506a22620fa78d332ac5a3c1f7c7085d0462283d6d8b0f17b2e0

Observation 20b72e1b-ea82-4c20-8442-c828b79e361e · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.307761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.402199Z digest=sha256:85da3a9c6020c0d8f3bf9b3f4444f790d5587230b04fdf650f265080e2387080

Observation 9a894983-d8ce-402e-b519-68e6b63834c0 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.294524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.405954Z digest=sha256:2cec3ead86fc68f4cd2591509f47cb6879304b99114ac188d49af8a3a48e8dbd

Observation 6335bd32-dda7-42fb-b6b2-ff2428698243 · outbound

This paper cites an unresolved cited work.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:15:43.245130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:15:42.409411Z digest=sha256:e375e33a9c084ae17dd58da0034da15d99182c76f9cd1b81875bc3bf48aedf16

Pith citing papers

Observation 4cb3a065-2600-492c-a285-090f70fc77d3 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 297

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.521889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:26f4676ff3c008690ccb44437e31ac9c73ab0fa9d0c07936050c8825bc1d0904

Observation cad82a6b-df4e-4c32-8a85-f002afae1123 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.404330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:50e2207be583a7c8c8dcd2edfc125f114244dcadd266279a17bca2ffc80009da

Observation 23c52071-9374-4517-9e4b-b5b62d315cf2 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 233

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.322627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:dd0c90577d788dc7b5e56546c5f9e6414fb2e6b61243f43c028bbeb6414bfd73

Observation c0c118ee-ba2c-4943-91d0-197b451d7758 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.690136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:e07fcb2de04fdb3b2b95c9cd914c62ce6a41776a155569a352637e9effcbf416

Observation c5ef8303-4565-4db1-a785-e21994e09dc6 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.499729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:2d02f320c76bd9a0b45c04b2c88fda5aa7566a4324b5ad175c0e5232ac5fe257

Observation 0b1a2ca8-b126-4ea0-a620-aba6a117bbea · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.413324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:af641476e2257477d549eee867b53286e774417ffa140ab28ea08d072dea02c7

Observation f5fe546e-47f0-4f33-b2df-f9ef43d8d2e8 · inbound

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking cites this paper.

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T16:55:20.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:55:20.099628Z digest=sha256:1bad59997c72c65d7ddbaffd9c02576feebd0808b5d79bd8e5d1a964c40aa800

Observation e5b02838-3488-422e-b562-c080880d485d · inbound

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs cites this paper.

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.985595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:35:12.859669Z digest=sha256:236003006ea20cef80aa63893fadb06b87d986f6083d9cc942357a26b12d1236

Observation 9ae53715-7a4f-4529-9a94-fbda2d4735b4 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:52.525638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:b3e2de11e944d62a4f9cf85b28b3bfe400e1054f9723a68422845b86349fcb35

Observation dd904fd6-2380-4591-afc2-19ec70eb6245 · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.573891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:1ed77414207557e0e95873b18866a56d1a413156d6149b0e7b8e063565315d8c

Observation a1bcdc9a-4721-4290-a3e5-7842709a2325 · inbound

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval cites this paper.

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:50.543718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:42:23.224746Z digest=sha256:0c9a33d141d5b91a6020a8f2f80ba1cab6bd4cd07d8ab33e0dbb397ce261af37

Observation 76164de6-9435-4e6a-a1ef-2695f34cf3de · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 192

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.323884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:488bd780c849cbb318741155a0798f7aa874bd109d7e5b40d2d8611299f27561