Pith. sign in

Paper Citation Record · LEDGER

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation

As of 20 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2505.12744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12744 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:33:12.777054Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:32:29.151581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:09.870570Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9437be8-6633-42fe-93ba-a99950227dc9 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.553051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.553051Z digest=sha256:4314697508a2c10cf166383ddb6aad327fa4b940db195813163597acf846c657

Observation 054d140b-ea7f-4dca-a1bc-a6c150433fb9 · outbound

This paper cites Qwen2.5-VL Technical Report.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.558552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.558552Z digest=sha256:414df1ad0698b5d6a3ce585501f75dfd341650356cdf0032e9b12afd82140ea8

Observation 1e16e946-49ef-4832-850b-1edf634bf5c5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.562465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.562465Z digest=sha256:bdfa7312b04933490f0de938611927658b57bd9e08688eb47bb214ae022bc54a

Observation da33bb5f-c837-4fb8-a835-2512447dfc05 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.567255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.567255Z digest=sha256:f4406e79c68cf393c769ef836fab227cbdde15b4e5f20a4793d10c91e29403ea

Observation 5cfd8230-3e87-42a8-9973-5e58e278c02c · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.571906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.571906Z digest=sha256:6e1e82707cb491a832c5bbe1e7edfc67dc44d414a04db2f4fdcc53ea3e1f6202

Observation 606d6697-a9b2-4e10-8309-64266fbc35a7 · outbound

This paper cites IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.577205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.577205Z digest=sha256:0a8d870532df0eb941cdf25a5cf0f68de9f28eaf1e779c2c70e7267bc273b76a

Observation 111b15a2-ae6e-41b3-8f13-6e8d87168424 · outbound

This paper cites Moto: Latent motion token as the bridging language for robot manipulation.arXiv preprint arXiv:2412.04445, 2024.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Moto: Latent motion token as the bridging language for robot manipulation.arXiv preprint arXiv:2412.04445, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.582124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.582124Z digest=sha256:8bd3a01dfa620f9586a3747b781e6b0cb974553ddc125f9579f17495bc7a550d

Observation 3f5ea101-12cc-4bdc-906f-f89c6caf2e26 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.587177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.587177Z digest=sha256:a93a5fcf21213c99a4dd27c3a622f6ed9810a7ec8578e04ce172535189e365ef

Observation e276d1f4-9344-43a9-b333-59525faf413a · outbound

This paper cites Dynamo: In-domain dynamics pretraining for visuo.Motor Control, 2024.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Dynamo: In-domain dynamics pretraining for visuo.Motor Control, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:33:13.519431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:33:12.592029Z digest=sha256:c21138c64ffad940a6d9fe86d19e3dc5371ab427fe86bfb7591e9d22dea60bf8

Observation 25f9d2c0-de3e-4921-9b46-32627781030a · outbound

This paper cites Fineclip: Self-distilled region-based clip for better fine-grained understanding.Advances in Neural Information Processing Systems, 2024.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Fineclip: Self-distilled region-based clip for better fine-grained understanding.Advances in Neural Information Processing Systems, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:33:13.505543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:33:12.597140Z digest=sha256:91eaa0dcb5770d2c961ecc059275c380bef2d0633257f2e3833102e914867c9e

Observation 2310d7c0-ac87-4591-9314-9448a7f5b668 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.601561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.601561Z digest=sha256:882be57af2b4730fa4b3cdb3b10d06b7b85e3ca9a0a3653e1c48e424f35357e2

Observation 58f55d5e-666f-48ef-849e-e3f2976303b2 · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.606383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.606383Z digest=sha256:fde8df171816763fe8f65e35ae7e5c7e0c1820ef5380b6f0dbb95be8ffbd2c9e

Observation 7a626588-9adc-4552-b445-32193d6d50aa · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.610886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.610886Z digest=sha256:cbb684804ff95aba84a9a10afce47fec1440633fab6cda490b037d4d0a08fd00

Observation b14d0579-2e9c-457a-aec7-578d4479316b · outbound

This paper cites Deep residual learning for image recognition.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Deep residual learning for image recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.615147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.615147Z digest=sha256:313b766d023e35d3e1956e4a07c0c8aeeb7ea815fa126ed28b739a385d583ea8

Observation 9bbd22c5-1ada-438a-81e9-17f19e201937 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.619297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.619297Z digest=sha256:4265141b4ca2c55b35b6971245557a37c85fc56eb25597ea5b346cbd3203e188

Observation 50121185-7d97-45e8-9ee5-c5b68b326891 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.623781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.623781Z digest=sha256:e238dc11c97c74a68b5fd901c31f313a50d40e1f248f21af93587e0d862172f9

Observation 986b2236-b0fb-41ab-87a0-c89adc56dad2 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation An Embodied Generalist Agent in 3D World

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.628364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.628364Z digest=sha256:dc013f3965940fa3800b221fd0854256ae21fc5f8823d1de9e2a3539a10e7735

Observation 46f8da5e-a78c-424f-a55f-ea76046ce8cd · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.633560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.633560Z digest=sha256:5271f9072dc26aff9871335c18a1fef2dd0356ca6b810fa991f282cb29a9389a

Observation 303ff475-e242-476a-85db-5a372d88be0a · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.638487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.638487Z digest=sha256:6156421cd13f6598621a042d8945f437782eacc849ab576c782d5729e9732828

Observation 2a5dfefd-cce6-435a-b488-0ad208ff35e0 · outbound

This paper cites OpenAI o1 System Card.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation OpenAI o1 System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.643545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.643545Z digest=sha256:5db2ceba61702e7737e3edac31cdde0db1f1d91fff47ae098f464aca45b44dab

Observation 3fef3789-952a-4476-8729-aec616554738 · outbound

This paper cites macmillan, 2011.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation macmillan, 2011

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.648186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.648186Z digest=sha256:642d3e9fd2e0b05e7dacc991251cfec58a04d12225936ec13e18b1a2fed0cf50

Observation 6e4172b8-3e57-4640-bdb3-cf3dc4763842 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.652973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.652973Z digest=sha256:9ad293838630ce8549af5b08df5e3059b369d77d224a0ae7532892a434a9fd03

Observation 5f107e40-f37e-45ca-825c-956417c6268c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Lisa: Reasoning segmentation via large language model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.657610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.657610Z digest=sha256:6d1aef233bb9980e3028b11e23669afa245fa3bd77351fb68e848fe313bdb118

Observation bb1253bf-5e96-4e42-8e72-246894153c22 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.661855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.661855Z digest=sha256:6723c35bbb5ebcda904ed7d7f7994187373ba00610c65f802a6a8a76bc2a621b

Observation 337af9f0-41c6-439d-950f-33fc041554ac · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.666548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.666548Z digest=sha256:f48dd6627dfd21fbb215202dacd707a351fdbddd25d1ddc391bf773d0614ab91

Observation 8dd7c1a0-a8b9-41dd-9eb4-2746304ca2a7 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Open-vocabulary semantic segmentation with mask-adapted clip

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:33:13.450848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:33:12.670722Z digest=sha256:e0a3990c75b23fdfe120ddf83c56482fd36f1f8a96adac7411b937b0d127e7e2

Observation d25ab36c-4505-44c6-b94d-37d449624964 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Code as policies: Language model programs for embodied control

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.674400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.674400Z digest=sha256:6d349975e98e215cfebdcf2f14c1102fad9dbe516f2023408db2fff8c0b48599

Observation 8f31f9f2-cb43-4c08-a1ca-2e79ef004195 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.678103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.678103Z digest=sha256:11ca7c12a0adde1e3f24677cb78b0799db6743964c91127884e7749a74a7e810

Observation 20cc83de-0d40-4efb-ad29-71bdc6ae1a2f · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.682078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.682078Z digest=sha256:0f786f684ac7630edb23de667d51a6cbbd6bb6868f1a03b3f92a73f1376d3fd3

Observation 0c920773-b5c1-4072-be56-67ce2a73af87 · outbound

This paper cites Decoupled Weight Decay Regularization.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.686035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.686035Z digest=sha256:d60f217126a05494ee3b8e2b709eb4c91eb4cf22fc941fa1a24a7f9b058f160e

Observation 92af7a2e-7f78-43c7-899f-37ac6c3d585c · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.689769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.689769Z digest=sha256:426b14b4ea41142378decc002b3c4b9e8fdd24522461a6768f7f24810382a6e8

Observation 21cc9d64-dee7-4f4a-8b34-c857a59247c3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.693826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.693826Z digest=sha256:ed1c5f38e635b87f30a2618869fc7df21170088044a47e8b8abc8ae195f88a44

Observation 79013272-2282-41c1-b221-9800cec958a6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.697547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.697547Z digest=sha256:3558ea9ffa523411f4d9488121db82ec732cdef91979dade111155764ee94936

Observation 5fe58cd1-3f50-4324-92e2-8239cfe36c68 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.702319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.702319Z digest=sha256:d7954a6d3f971e8e051a53e20610c28bccbc8d27535eda1f215a59281423b3a7

Observation f5edd807-34a1-45e9-b295-35521afbcc3a · outbound

This paper cites Computing euler angles from a rotation matrix.Retrieved on August, 6(2000):39–63, 1999.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Computing euler angles from a rotation matrix.Retrieved on August, 6(2000):39–63, 1999

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:33:13.413930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:33:12.706679Z digest=sha256:cddabfec3115918b8a0bd5c0e7ce78086da404b1d37d16b13434609ea9e395d3

Observation 28c2bee9-4443-448e-a252-fb7b397b257f · outbound

This paper cites GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.710846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.710846Z digest=sha256:ed16316b59f3073a67da0d0a17062617ecad03cdfdd483dacece358d86b7507e

Observation c746766d-ba61-47ec-9324-b55f220a4ce6 · outbound

This paper cites Overcoming Support Dilution for Robust Few-shot Semantic Segmentation.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Overcoming Support Dilution for Robust Few-shot Semantic Segmentation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.715396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.715396Z digest=sha256:faaab0ba155b4eb0339a810b2d377e9e7d654de7583777abcd5b8a2635ddbe60

Observation 81f954b5-859b-49ca-abc9-faf06bfd477e · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Octo: An Open-Source Generalist Robot Policy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.720092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.720092Z digest=sha256:19b276d9c760f8c06525c06c98ddd9514af648cc9ba625683d668f0d84c2b56f

Observation f6b118d3-7c43-4508-b74f-24c7dfd2b14e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.725029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.725029Z digest=sha256:b8a20aa2656419fe6f7919d10ca6b869a9701319c320371ea47597c0fff06184

Observation 7f1396e7-c4ec-4a2d-b5c2-051ccc2a59a8 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Bridgedata v2: A dataset for robot learning at scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.730223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.730223Z digest=sha256:cea6eff5f1b47bcff85c29b5178ec3d8ca7207ac84deea3ecd95660c4147b769

Observation 0c61ae86-8fcc-43e6-bd73-3bcb10a1334d · outbound

This paper cites DART-LLM: Dependency-Aware Multi-Robot Task Decomposition and Execution using Large Language Models.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation DART-LLM: Dependency-Aware Multi-Robot Task Decomposition and Execution using Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.734688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.734688Z digest=sha256:74a826f0b3f2f626755ee00965820baff14a1d03ccd95d439e0ff8841a4863ac

Observation 256228c2-1388-40e6-a0b1-5a3249fde096 · outbound

This paper cites Q-learning.Machine learning, 8:279–292, 1992.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Q-learning.Machine learning, 8:279–292, 1992

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.740133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.740133Z digest=sha256:0772567aa07d2188ea48efbe9abdcca27776688ee0251177deb882044a97326b

Observation b3871ed7-281b-4b1f-9f3f-5ee4d1a88593 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.744820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.744820Z digest=sha256:05abc1305fb357fbeda098b45fff592197051db3adf7f5ce62999e9d9d5a684b

Observation 10a290b1-691d-4070-bb64-2b506ace1192 · outbound

This paper cites Qwen2.5 Technical Report.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Qwen2.5 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.749687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.749687Z digest=sha256:6f2675689d30af42a454bbcb8b9a41e7edde0a50c7e2f4c6c30eaf1db75a256a

Observation 7ad22a85-1b4c-49bd-8103-98a1074fff65 · outbound

This paper cites DeepCritic: Deliberate Critique with Large Language Models.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation DeepCritic: Deliberate Critique with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.754141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.754141Z digest=sha256:1d1d1850e2c898c3cdf7111a8cf6b37fb31e2f9f2226f5eb8d3cd4efaf48d28a

Observation 4dad9383-d3d0-4063-95be-fb7038cdaa81 · outbound

This paper cites Latent Action Pretraining from Videos.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Latent Action Pretraining from Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.758791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.758791Z digest=sha256:d73057d576abd79a7f0953db221548719cd097c61af6e5cd571c8c0a019077d9

Observation d1b71996-a53e-46b1-aec7-a7ac379ab7bd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.763387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.763387Z digest=sha256:159a748ac1cac5029b12931e51b01d908fc2d8d907d8c9f6ef2c709288fa2dd0

Observation 5bfb273f-f52f-427f-a7e4-fd445edbdc30 · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.767830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.767830Z digest=sha256:9022a54238057a0d8bbdeffcdfaece5b3cc72d80c09ad2dc5563253cbfb5d99f

Observation c5dada62-7421-440c-9c94-90048fbaa2e1 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.772083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.772083Z digest=sha256:8ff252d7b5cb7b56c19ccce1445452da1ae4bd615f6e5b7cbb80b5e863399262

Observation a3ad2b89-8e9e-4b9e-9570-0e14eb8bb404 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.777054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.777054Z digest=sha256:28d455eafd5d96446c4bcc9a501dca8802e3bcdbba4c6ac1252ce11129ef2fe8

Pith citing papers

Observation 97ee8c20-695a-4f66-bf59-d38a765773a6 · inbound

Mixture of Horizons in Action Chunking cites this paper.

Mixture of Horizons in Action Chunking Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:32:29.151581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:32:29.151581Z digest=sha256:a60eda6235d67b501b449e618ee7fb1dc7b96c9af39534fd1930ba093821bef9

Observation 7d9088c2-a2cb-4738-b39a-86e8688c1840 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:29.619640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:29.619640Z digest=sha256:9ce08fd51a83f8f46fe9b68e23745272e5b4abbfc5ab406d2c4676b60ea4a9ef

Observation fc62dda3-bc48-4f55-87eb-adba98446f84 · inbound

Learning Action Priors for Cross-embodiment Robot Manipulation cites this paper.

Learning Action Priors for Cross-embodiment Robot Manipulation Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:00:09.872094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T19:09:56.409766Z digest=sha256:e6e3d3cffa3b59a2e57df45ecefc9d6e36ac0d6e1271056793d91fb8f0dc2df6