Pith. sign in

Paper Citation Record · LEDGER

RoboBrain 2.0 Technical Report

As of 15 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 54 inbound Pith citation observations for arXiv:2507.02029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02029 v5

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:47:31.149194Z

measured 143 of 143 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:36:21.485253Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T12:04:50.478072Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39906c01-6c21-4128-949a-80cee1babd2f · outbound

This paper cites Webdataset: High-performance data loading for deep learning, 2020.

RoboBrain 2.0 Technical Report Webdataset: High-performance data loading for deep learning, 2020

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.532740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.532740Z digest=sha256:6575b776a61bf37ab9f1260a71f8b51d5f194bf3530c5fb466c5afe35221bd70

Observation 677d2b77-6640-43f0-a663-90604ecedada · outbound

This paper cites Claude sonnet 4.

RoboBrain 2.0 Technical Report Claude sonnet 4

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.647128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.647128Z digest=sha256:c730e468db92577dbb440bdc0652d33830b6c365bfc68d399379d36fff2775a7

Observation 2d2c5fc0-8b1e-4052-ab4c-0b1b6c11e429 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

RoboBrain 2.0 Technical Report Scanqa: 3d question answering for spatial scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.724399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.724399Z digest=sha256:442dd8e983d9f5619356cf72dfce27449e77bea7af001817081a6a41680665a2

Observation 244d8de8-3194-4fdd-a734-be9509806f03 · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

RoboBrain 2.0 Technical Report Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.803395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.803395Z digest=sha256:32e4e2c442139177db49fe3211b00a87a36f640ee2bce0f86c5660ca46ccdb9d

Observation a17216aa-7d8f-4945-8134-3654cc2ca257 · outbound

This paper cites Qwen2.5-VL Technical Report.

RoboBrain 2.0 Technical Report Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.861229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.861229Z digest=sha256:b6f3473df7d7de0e0c07a1733ddc6a2ba386da4828399f00f26082e886c79e32

Observation f72dd003-b3e8-4de9-9b7d-1fb15a1c66ad · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

RoboBrain 2.0 Technical Report AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:26.979701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:26.979701Z digest=sha256:53ccb9903563e2f86b86358472aa3ddf8b988d93ee5c49ca055af4f261c8bc89

Observation 0849cc23-5289-46b8-a50a-c9c1a710953a · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

RoboBrain 2.0 Technical Report Sharegpt4video: Improving video understanding and generation with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.050956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.050956Z digest=sha256:d5bac771c946fd86749207ee3a51aaaf30270d4c992c560eb4142191d4221d1d

Observation 79f26d2a-7a1c-41b2-8ff4-82c0113f6133 · outbound

This paper cites an unresolved cited work.

RoboBrain 2.0 Technical Report Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.208513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.208513Z digest=sha256:b0d5b6b392376f30f2c88783c99766d7a3623d02e802a2cbd74ba99a5ba60501

Observation ecf0f690-2bb9-40be-8e40-9e22bf23891c · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

RoboBrain 2.0 Technical Report EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.281733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.281733Z digest=sha256:3181f249ebf2a72309840720f558ee11402bda9f791fa103b8b3fb402486901d

Observation cc131c5a-11b1-4c06-a342-0008b3b18a14 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

RoboBrain 2.0 Technical Report Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.373818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.373818Z digest=sha256:e033fc2e6cd168b1f1f44913ad8d68760beb894563557004c342222c9ee94f71

Observation 5bbf8db0-ef9e-43af-bb5a-fb3b6f9a74cc · outbound

This paper cites Llm agents for education: Advances and applications.arXiv preprint arXiv:2503.11733, 2025.

RoboBrain 2.0 Technical Report Llm agents for education: Advances and applications.arXiv preprint arXiv:2503.11733, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.447799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.447799Z digest=sha256:62e96b73f488fdf835117926df33f5afc95145df6aacbea245ae36708d6f6508

Observation 525205e9-feb2-4aad-b8e7-96342ed3f4d9 · outbound

This paper cites Flagscale: A unified meta-framework enabling adaptive heterogeneous computing for the llm ecosystem.https://github.com/FlagOpen/FlagScale, 2024.

RoboBrain 2.0 Technical Report Flagscale: A unified meta-framework enabling adaptive heterogeneous computing for the llm ecosystem.https://github.com/FlagOpen/FlagScale, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.519289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.519289Z digest=sha256:57e988f64c5462dfa4977576cf69e1cf77fb55bb655c766802277781bed586e2

Observation b6cfd4bc-253f-4779-a74d-c800be6a9dbc · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

RoboBrain 2.0 Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.628877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.628877Z digest=sha256:36e23e548d17d90a1888db7f638065c2b34851f2f2685bd4f2c0affcd18f2379

Observation 16e40d10-1250-4016-899c-0ada7d423795 · outbound

This paper cites Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models.

RoboBrain 2.0 Technical Report Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.780921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.780921Z digest=sha256:88652f0f2cb001d1e8537b77abee6bf5cf5f97c0a39c1fe53121cb2cf4ff8667

Observation d40506a7-e4ba-4a11-b952-94aac606ad7c · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

RoboBrain 2.0 Technical Report Blink: Multimodal large language models can see but not perceive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.854641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.854641Z digest=sha256:8c2ea58ada83ef1be1fdcb08b17f81cb63695ee58b89ec4d26b257898d3a4db3

Observation e4ef93e2-9505-4ca9-be1e-d5455f83c727 · outbound

This paper cites Gemini 2.5 pro preview: even better coding performance.https://developers.googleblog.com/en/ gemini-2-5-pro-io-improved-coding-performance/, 2025.

RoboBrain 2.0 Technical Report Gemini 2.5 pro preview: even better coding performance.https://developers.googleblog.com/en/ gemini-2-5-pro-io-improved-coding-performance/, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:27.976037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:27.976037Z digest=sha256:17de8ccdea83b958e3452e1f512de7c5c6e15e5f3bd3d679d7278f6597a17d83

Observation d50f7c05-705e-4c0b-b809-91381ded576b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RoboBrain 2.0 Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.046364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.046364Z digest=sha256:63b1ab2f97a8da06baf0e953f82b3de92681ab54ce3c4e4dc31da1a85760ab86

Observation c67bc056-83d2-4264-83f5-d7a24b974d02 · outbound

This paper cites Seed1.5-VL Technical Report.

RoboBrain 2.0 Technical Report Seed1.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.125248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.125248Z digest=sha256:17558524d46fe72949955ff92871b59651e94a5adfd364ca75d0d03fe036eae0

Observation 21c77a9d-7977-47bc-b437-29abdcb4fa9b · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

RoboBrain 2.0 Technical Report Lvis: A dataset for large vocabulary instance segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.278055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:28.224406Z digest=sha256:cc0319d7b733b1c1966a896dc4b4789b228ea3d2a3ea8fe0d990101818bbcce6

Observation 079ee21b-024d-408f-aa54-591abe156fb4 · outbound

This paper cites FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation.

RoboBrain 2.0 Technical Report FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:47:31.781244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:28.353756Z digest=sha256:9a95eb9aa61269335c34169e8d56843905ff88818ea01b663979550ddb430e49

Observation 77a87e03-afaf-4699-9570-43234d43b22e · outbound

This paper cites A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry.

RoboBrain 2.0 Technical Report A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.421111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.421111Z digest=sha256:a4f82617deb30fade7f94f32626cc63fedb852ec06207d95dc304fff8bdfffd6

Observation dfec7ac0-3e6a-4cda-81e7-d2da9b7c947c · outbound

This paper cites GPT-4o System Card.

RoboBrain 2.0 Technical Report GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.505007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.505007Z digest=sha256:e5456f29c7a8165a5d3e2812bbd42673800b0e2c36a22dd4742bec8dcdde004d

Observation 1f6677a5-f6ce-49cf-a2f9-6ca964fee7fe · outbound

This paper cites Robobrain: A unified brain model for robotic manipulation from abstract to concrete.

RoboBrain 2.0 Technical Report Robobrain: A unified brain model for robotic manipulation from abstract to concrete

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.572627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.572627Z digest=sha256:0da39771a90eb5365fc973a4768bb10a6de4a91880066d1dc91d7bda79626cc9

Observation 62194061-0669-4ef6-9159-3dd198d0baf1 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

RoboBrain 2.0 Technical Report Imagic: Text-based real image editing with diffusion models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.632024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.632024Z digest=sha256:5467f579e62cbc95a987c228504bcd5cb0e98c640989fc6c3ac11788cb9d5de0

Observation 45924ed9-0e18-4060-a563-080f6458ba87 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

RoboBrain 2.0 Technical Report AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.735777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.735777Z digest=sha256:bb4aca77e53c1cb6c927c97ba82c8146cf4ba663d1ea20dd6dc6abeab0994fac

Observation 56cb0289-22aa-40af-8685-0e92e1d25712 · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

RoboBrain 2.0 Technical Report Reducing Activation Recomputation in Large Transformer Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.846985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.846985Z digest=sha256:fc48c4b1d2df5120b448827ecd2921f7b2909a99e0c445ce0467cc1df44efcd2

Observation be4130e0-c720-492f-8576-ab834e2836ea · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

RoboBrain 2.0 Technical Report Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.932959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.932959Z digest=sha256:e0f54502bfcf968a70e97864a85dc569e7c0805c4e4266e346f58b6e9d1de2c3

Observation 6c47d7e7-270a-4983-8b48-7fed382aa73e · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV, 2020.

RoboBrain 2.0 Technical Report The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV, 2020

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.251143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:29.014207Z digest=sha256:eb08353d76d60fb2cda4d81bc9491f7ca0ffb538c178054be509e92390491cb4

Observation 1c4e674b-8d1d-4997-91ff-05bdece189d2 · outbound

This paper cites Cubify Anything: Scaling Indoor 3D Object Detection.

RoboBrain 2.0 Technical Report Cubify Anything: Scaling Indoor 3D Object Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.062663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.062663Z digest=sha256:b01125403900c125911bf4c3d5695363837e3cdee4bc697a7056f23dd7485ef6

Observation 93f1fd8f-a9be-4e03-8992-7ad952de31e7 · outbound

This paper cites Energon: Scaling megatron-lm training with data and expert parallelism, 2023.

RoboBrain 2.0 Technical Report Energon: Scaling megatron-lm training with data and expert parallelism, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.241849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:29.159824Z digest=sha256:131759763cff4094ff969c8f9cde0884f504f30226a068328934e9056eb3aa18

Observation 36b5ab3a-f716-4ab1-9df3-cf863982ad9d · outbound

This paper cites DeepSeek-V3 Technical Report.

RoboBrain 2.0 Technical Report DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.218019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.218019Z digest=sha256:37b96f938a4e48f6a31d0b5dd62f50a5227a7dca2b99e9ef70b6587e72b34638

Observation 5658a705-190b-40ed-a03e-de457d93e25d · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

RoboBrain 2.0 Technical Report Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.322899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.322899Z digest=sha256:e4111f6797c6321aa3b069ac0dc9443ccf68ed945ee01f886bf2d67f1d904c61

Observation 8c48df8e-bd7d-49b2-96e2-18730c36fc9c · outbound

This paper cites Improved baselines with visual instruction tuning.

RoboBrain 2.0 Technical Report Improved baselines with visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.449571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.449571Z digest=sha256:3a3037f2e70d3ba7662d4df5f95145231ea89f5c60400a3451af4cc59a99c95d

Observation 92ea8b59-4c88-4c0a-b179-f028ac161a1e · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

RoboBrain 2.0 Technical Report Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.226010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:29.528347Z digest=sha256:31e931be430ff269559dd9cc5b60f86d853197b3b74ba8f150c76a3b59ecd865

Observation 3dc12c3f-bc74-4b62-a834-a1ec1e5d00bc · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

RoboBrain 2.0 Technical Report Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.553830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.553830Z digest=sha256:51e06d37df9dff674c085995f2f7b265bdf7132e3e2673e3d957ff41d30a5c46

Observation 92387108-7980-4289-8454-b5ce4b3b9582 · outbound

This paper cites Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces.

RoboBrain 2.0 Technical Report Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.605186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.605186Z digest=sha256:ddcf625cc3227aea8da381da3b18566a997f434b60811f8dacecc912ce0c800e

Observation 9f4eaa05-3c48-4295-867b-5f016dce07cd · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

RoboBrain 2.0 Technical Report GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.712258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.712258Z digest=sha256:aabe58e832d81d1eb4578e2dd9d20e37d68aa03cf08d92dae12ac96d87e56b5b

Observation b7d85442-4198-47da-8d4a-051b4a91c961 · outbound

This paper cites MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations.

RoboBrain 2.0 Technical Report MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:29.800018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:29.800018Z digest=sha256:d74660315061ec57064671fe3e3a1d144363136f3bf7bf46e1afa70c21da95c5

Observation c2252a69-9cda-4174-8228-8eca3d874cff · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

RoboBrain 2.0 Technical Report Sqa3d: Situated question answering in 3d scenes

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.209299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:29.963352Z digest=sha256:2b2ad412b0a8cdc546a228666653cc8667bfeba5483e4e94d171f2744db8b4a3

Observation 3235eac1-294a-40fc-b1e2-14ef618d2217 · outbound

This paper cites Mixed Precision Training.

RoboBrain 2.0 Technical Report Mixed Precision Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.084151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.084151Z digest=sha256:4c233ecaf5fd2bd06d122ce48e6754e035ae9eec44692a7c6493162f9e6937ac

Observation 6dbc3042-abe1-4d06-8e6e-1d41dbed98ad · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

RoboBrain 2.0 Technical Report Ocr-vqa: Visual question answering by reading text in images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.158464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.158464Z digest=sha256:292b598cba97d46af11976437eda74c048ed639cb4127317b68046a8ab103ed7

Observation 06a9e879-b5b4-4a56-b5ee-1a82b12ed39b · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

RoboBrain 2.0 Technical Report Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.192373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.293811Z digest=sha256:1fb38c0a57c2a617c76a89f7e749c5986d1a898dca9f0530b14e527157a7352d

Observation 722dc628-b96a-42c5-82ef-02ce5b32cc0e · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism, 2021.

RoboBrain 2.0 Technical Report Megatron-lm: Training multi-billion parameter language models using model parallelism, 2021

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.181339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.452725Z digest=sha256:ae85c18c271a5ed93bd64c14fc3f1274defa2fa69400f4c37ba98c6a2407fda5

Observation 66980da7-aeb2-4399-ba2f-9c20f51c687d · outbound

This paper cites GPT-4 Technical Report.

RoboBrain 2.0 Technical Report GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.603921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.603921Z digest=sha256:9a63eaaf696d35b4ea243b797b5945c17530d066360f70ccf818805f30cff828

Observation 93ee9142-f32d-4fbb-834a-00c070996b00 · outbound

This paper cites Gpt-4v(ision) system card.

RoboBrain 2.0 Technical Report Gpt-4v(ision) system card

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.169931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.755771Z digest=sha256:6eeedc51b659f2b1d788836ea33e4b1ee625f5aa8313f7316849fad1fd6f0370

Observation 434bef0a-bcf3-4392-ae1b-d79ef4c3c70b · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

RoboBrain 2.0 Technical Report SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.915731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.915731Z digest=sha256:4955a5e520887ba19c2f5c5e79aca79aa159d6f6e8c8b46b5f810f0182b49c45

Observation 98422b66-815e-4ab5-a764-b357fc839249 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

RoboBrain 2.0 Technical Report Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.976655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.976655Z digest=sha256:d35ceb2ac6b1b4a76e00265e7c110487825ba776d0171fa766cf2a58997437cd

Observation 63d5768d-ce84-43c8-a18b-371b80610a31 · outbound

This paper cites Unidepthv2: Universal monocular metric depth estimation made simpler.arXiv, 2025.

RoboBrain 2.0 Technical Report Unidepthv2: Universal monocular metric depth estimation made simpler.arXiv, 2025

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.152636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.980578Z digest=sha256:14eff7f766e239bd6b5b1c3ea80408c23c635235182f65abdceb26f7dcb62617

Observation 3fc90b62-ee1e-4288-a698-b003809ff63d · outbound

This paper cites Cuda memory management, 2023.

RoboBrain 2.0 Technical Report Cuda memory management, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.142246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.984067Z digest=sha256:375b21a98512de6c6f0596b82fa7d922469cd3f959fd93f5a1c525d366191c0c

Observation 63904301-a55f-459e-978d-2f56f8d39dae · outbound

This paper cites Qwen2.5-vl: Multimodal llms from alibaba, 2025.

RoboBrain 2.0 Technical Report Qwen2.5-vl: Multimodal llms from alibaba, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.132077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.987819Z digest=sha256:2cfffe765511f1e41023eedc73dc096b47c273dfa2255bd9dd285e2d74a459b8

Observation 9a1d25a4-97f9-4bb7-89d2-fb32017536ff · outbound

This paper cites Paco: Parts and attributes of common objects.

RoboBrain 2.0 Technical Report Paco: Parts and attributes of common objects

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.122618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.991348Z digest=sha256:9c22f718cbe4eb570daad48ededb351b661941ac8bd96d2f497bd3dec9bc79f6

Observation efccb56a-0994-459d-8329-e462582d17b7 · outbound

This paper cites Sam 2: Segment anything in images and videos.ICLR, 2025.

RoboBrain 2.0 Technical Report Sam 2: Segment anything in images and videos.ICLR, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.111805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:30.995693Z digest=sha256:bff7efdd5dcf87285d36a301db4706366c269a2ae1b03fdd1939604011f0072c

Observation 7d82e8a4-157c-4daa-95df-f7e611df51ae · outbound

This paper cites Sat: Spatial aptitude training for multimodal language models.

RoboBrain 2.0 Technical Report Sat: Spatial aptitude training for multimodal language models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:30.998852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:30.998852Z digest=sha256:3ec8bc6f1ee67dd98c8506825b7e7619f0eed9cfe537cf433102e6adb4df5465

Observation a0a5c116-98b8-404e-95fa-a049e3bef303 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

RoboBrain 2.0 Technical Report A-okvqa: A benchmark for visual question answering using world knowledge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.002436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.002436Z digest=sha256:fec078eed000b8edaae54a2885e44f4acaa53e51a47581b3a79e3a9fe6d2a26a

Observation 1c2c044b-5aef-42ff-879d-e3859db2d95b · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

RoboBrain 2.0 Technical Report Robovqa: Multimodal long-horizon reasoning for robotics

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.095280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.005433Z digest=sha256:93d3df5cff42d5ff094c5086a2f77b07a310dfd7ff9f2c618ed87a6474131bae

Observation 47522736-b458-4690-b46a-4353ca7b761f · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

RoboBrain 2.0 Technical Report Hybridflow: A flexible and efficient rlhf framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.008891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.008891Z digest=sha256:ec246a7bc10bc913b8eed9b4819535b3300d09bed3287fa09720f206e9c1ca40

Observation fd1a35aa-1e42-47d4-aab2-bd4d01433f9c · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

RoboBrain 2.0 Technical Report Emu edit: Precise image editing via recognition and generation tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.013713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.013713Z digest=sha256:5e90de44a8d772ba71d7b6d191bb2d363a3dd1fa82a33af9769ae0b244abec2e

Observation 9dcd3b63-22cc-4889-9413-54d1c3612abb · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

RoboBrain 2.0 Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.016748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.016748Z digest=sha256:1aa2f6ae44334ec4c43c49d705cfb9f026ff49b37d12fa338b852016fbaadd26

Observation 75e5e9dd-0aca-4622-bf7b-ed12be3f80a2 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

RoboBrain 2.0 Technical Report Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.020517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.020517Z digest=sha256:da31453f9739c2835ca23c54b4c13760fadb2c50bea77306c369e9bcc9141d1d

Observation 8e506e8c-a861-404e-a997-9e38d5b02cf5 · outbound

This paper cites MultiModalQA: Complex Question Answering over Text, Tables and Images.

RoboBrain 2.0 Technical Report MultiModalQA: Complex Question Answering over Text, Tables and Images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.024566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.024566Z digest=sha256:716c35ed66d734b5e45e2dee726726302221be2c1d3913c70c4e52f89f83b90b

Observation dc5661de-b82b-4453-97b1-4facdefd8cd9 · outbound

This paper cites RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration.

RoboBrain 2.0 Technical Report RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.028205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.028205Z digest=sha256:e0efc94b582b3cd28019b7a0538c7f21de346eece94ec2b8c2575e99c81e4ee6

Observation 94ff3ac6-ef97-4b1c-af79-3ed3e0d694df · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025.

RoboBrain 2.0 Technical Report Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.031498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.031498Z digest=sha256:6f05abfa852db8a837b7a22aef70036cb1fe734e502d4f91b9d3b1852aac0a58

Observation 1ecd60d6-94d8-4d48-b674-fe45f3c604b2 · outbound

This paper cites Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology, 2025.

RoboBrain 2.0 Technical Report Video understanding with large language models: A survey.IEEE Transactions on Circuits and Systems for Video Technology, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.034168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.034168Z digest=sha256:eb28f3022a64ec292a793722cf5debe62563f1bb92d8da9c5a77271936e7db50

Observation be8963c2-8f60-44b4-96df-23fb6d9a322e · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

RoboBrain 2.0 Technical Report Gemini Robotics: Bringing AI into the Physical World

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.038680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.038680Z digest=sha256:47cf2d100fe056ecedf956f84018f6b8a086308e666968b63f92d0284d58959f

Observation cac9d90b-871c-408d-b7af-81af1a63183d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

RoboBrain 2.0 Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.043230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.043230Z digest=sha256:cb1bb57068811f92429389497257c47da6dd0bdec8a1eb925f93be96e4fa21b7

Observation 7f133fbf-58a4-4744-bb1d-709678cb675d · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

RoboBrain 2.0 Technical Report Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.047827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.047827Z digest=sha256:a5040c1d8184cd5caaff89130d61e881ddd94017f4417cc9eb732ff99092a449

Observation d475f2c7-393b-4a55-a68f-386d731bd9ff · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS, 2024.

RoboBrain 2.0 Technical Report Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS, 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.052452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.052452Z digest=sha256:6b701fe25c4e29f62822f98d1bffd6b4499b6463d993c0e70d7db0810633fa41

Observation 50253336-2da7-445b-9c93-ff970d64e1fb · outbound

This paper cites verl: Volcano engine reinforcement learning for llms, 2024.

RoboBrain 2.0 Technical Report verl: Volcano engine reinforcement learning for llms, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.054021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.060947Z digest=sha256:ed6eecaeb3f6b13d8f71709303ad914b8f1d88923ff98c4eb506e36d5bfa9001

Observation 913dd2e6-7038-44db-8f2c-675451c9a2b9 · outbound

This paper cites Rio: 3d object instance re-localization in changing indoor environments.

RoboBrain 2.0 Technical Report Rio: 3d object instance re-localization in changing indoor environments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.044195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.068156Z digest=sha256:b89c7f199933a73a8dfe16b5783b9b4403a0be918426b6261f9728406abbc706

Observation 355ff154-54f6-468d-9d2a-cf1a9b7c12af · outbound

This paper cites HAQ: Hardware-Aware Automated Quantization with Mixed Precision.

RoboBrain 2.0 Technical Report HAQ: Hardware-Aware Automated Quantization with Mixed Precision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.071952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.071952Z digest=sha256:91a3617c307e7a73767ba22cdf9c81d1b040df2c27e1f4d18da3307bc3aca3a8

Observation fb068da7-b4c8-4806-8f81-2d16971dd52a · outbound

This paper cites GUI Agents with Foundation Models: A Comprehensive Survey.

RoboBrain 2.0 Technical Report GUI Agents with Foundation Models: A Comprehensive Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.078436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.078436Z digest=sha256:1e9526bc662b5ac2a08d2253e8964df617a83d9b88c99d3ddb0a7215f1265659

Observation 9b6cb17f-e9e5-48a1-9047-bc052a47239d · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

RoboBrain 2.0 Technical Report Internvideo2: Scaling foundation models for multimodal video understanding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.082918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.082918Z digest=sha256:dd266d70f87155cc557dafbd363b4f26c626d4fc6b420dc5e2abea504abbdeb1

Observation 2b6ac7c0-0881-4f2b-babd-4425db39a016 · outbound

This paper cites Qwen3 Technical Report.

RoboBrain 2.0 Technical Report Qwen3 Technical Report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.085978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.085978Z digest=sha256:a358bd0c7d3a2b7c27c3120af1a3db16dccd2f0164abea48073ae52cc9efe0bf

Observation 58be3216-1d8d-45c0-ab0e-5d2d615eeecf · outbound

This paper cites Magma: A foundation model for multimodal ai agents.

RoboBrain 2.0 Technical Report Magma: A foundation model for multimodal ai agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.092418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.092418Z digest=sha256:d2ef2d01dc0d4a56223a4c12ee5332d58fec9bc4c039e19bd36032a58f7e26a2

Observation bfbd99d3-e43e-460a-b0bb-7c491d0731b7 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

RoboBrain 2.0 Technical Report Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.100133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.100133Z digest=sha256:6626ffc1d1754117306849e13c0f09c7caf5d76637fe39bab311927763af03bb

Observation 8d4193a7-446a-441a-9e79-8037f04ca699 · outbound

This paper cites Modeling context in referring expressions.

RoboBrain 2.0 Technical Report Modeling context in referring expressions

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.013979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.103154Z digest=sha256:ab0d5d6f68a8baf8847fb6569d6033a9d8feb852c61aab9576bf3ffbef44117a

Observation 467bcb34-25e3-48b1-a861-3e2ead12dff4 · outbound

This paper cites Robopoint: A vision-language model for spatial affordance prediction for robotics,.

RoboBrain 2.0 Technical Report Robopoint: A vision-language model for spatial affordance prediction for robotics,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:32.005019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.106194Z digest=sha256:82fd3270f84df95940e522bb6ba7ddbe81f38fbaef3ed308937d18c67ae9ceb3

Observation 77eb9349-8346-4a05-9614-cfb18c4f8996 · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

RoboBrain 2.0 Technical Report Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.113444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.113444Z digest=sha256:61ef56ca61b313d2c059e577eedcbc0f43f1547c903560f80572bf25e12d7c03

Observation 9cbe1e21-b169-43ae-a279-94f17ba1a9e2 · outbound

This paper cites Recognize anything: A strong image tagging model.

RoboBrain 2.0 Technical Report Recognize anything: A strong image tagging model

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.995896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.116414Z digest=sha256:121a684bfd842dd562936201cd1b954f994a395a08ce5d64ca3b37a285071947

Observation 1c3f9caa-788f-4c8b-aae9-61c6d5482610 · outbound

This paper cites Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection.

RoboBrain 2.0 Technical Report Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.119288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.119288Z digest=sha256:8a0ae36bfd23515248e095f1da271ae5ddc364b3d0382ca78a9b3a92521c371a

Observation 4c8f8020-3653-4925-9c84-bdfc2bee5e80 · outbound

This paper cites Roborefer: Towards spatial referring with reasoning in vision-language models for robotics.

RoboBrain 2.0 Technical Report Roborefer: Towards spatial referring with reasoning in vision-language models for robotics

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.124480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.124480Z digest=sha256:fc5a1a1f30483c2047a0ac7e421a0ce52bee2130aea27c8e256f8c9830aa59ea

Observation 9ce09e21-fdc9-4c29-839f-f1f87cea34fd · outbound

This paper cites Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities.

RoboBrain 2.0 Technical Report Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.986318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.127650Z digest=sha256:88974be0ab4e2be89a8b21111a6f926eeeeba4ac382288bdcb38c0207cee3bc9

Observation 152ebfed-9f71-45d7-862a-04f50438a944 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

RoboBrain 2.0 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.131012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.131012Z digest=sha256:4037a6db0e0ba9b63dc9aaad86d52b2fd8a32cf381c75dd0d78891ad581bcc2a

Observation baa29bac-ba56-488b-9d9a-b09757409be3 · outbound

This paper cites Please point out the orange box,.

RoboBrain 2.0 Technical Report Please point out the orange box,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.976540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.133782Z digest=sha256:6b61d1374bd7c197d9005fb0efcbd572af1d248ff106086eaea64865c66a6603

Observation e041b661-cad0-4a2b-a86a-5917f5656bc7 · outbound

This paper cites an unresolved cited work.

RoboBrain 2.0 Technical Report Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:47:31.966825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.138713Z digest=sha256:28acd3f5b0998c3e5183acb1788a85b49cf6f09f12d2a2f574c73950f7acb9a7

Observation f38b44b1-b4f7-48c1-b0cc-271732b62447 · outbound

This paper cites Never use variable names as the action arguments, use the value instead.

RoboBrain 2.0 Technical Report Never use variable names as the action arguments, use the value instead

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.956781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.142243Z digest=sha256:dcc7a2ae883ef74d79825b01c47563feff9749fbaf957ccea78e7d7a17dbbaf1

Observation 6b57e33b-0328-4d46-861d-ece54d711147 · outbound

This paper cites If no tool call is needed, use final_answer tool to return your answer.

RoboBrain 2.0 Technical Report If no tool call is needed, use final_answer tool to return your answer

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.946710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.145508Z digest=sha256:228036e1d97f98189eaf8709d16a063683ba3501447607e77e88ffd84ce31197

Observation bfa14531-053c-4059-b7db-43deb12b9c8f · outbound

This paper cites # Now Begin! If you solve the task correctly, you will receive a reward of $1,000,000.

RoboBrain 2.0 Technical Report # Now Begin! If you solve the task correctly, you will receive a reward of $1,000,000

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:47:31.936002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:47:31.149194Z digest=sha256:3da42863616b864c1506beaa19c295cea447e1c13b49a1fb108ef061efcc5dc0

Observation 3fd2d4a8-16fb-4e9f-8587-06bfe8fdcf83 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

RoboBrain 2.0 Technical Report RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:31.110057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:31.110057Z digest=sha256:64d939d450f3e599cda6a6921e460959c1b2c67ef3d18fe2d59433bbcab7b8ff

Pith citing papers

Observation 1e619059-6be7-4144-bfef-f0c04d9be0bc · inbound

Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes cites this paper.

Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes RoboBrain 2.0 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:34.405669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:34.405669Z digest=sha256:4e581287e56c0251e88727469eebf70dfd6170bd52aab1a8ce96431bdc2ca8ce

Observation a8104a18-b6bd-4ad5-b0a8-412a6d027c3a · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning RoboBrain 2.0 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:09.771943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:09.771943Z digest=sha256:7d4e9b52fe6aec4458aa231fbc8a978f19d6cdf63b39b95ca068abff9498938f

Observation 721b8a4e-f81f-432b-8093-86e6492b33a4 · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.254972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.254972Z digest=sha256:7c57c99b5c4efd93e862adbc670be33558a55d1b8e8fdfe1265ff7d63c3808bd

Observation 5d270656-4c34-4b04-988e-df566723113e · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy RoboBrain 2.0 Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:40.028188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:aa1cd3b09b711e7ba5d255f120b166ff510e144a66002e888a3ae5f4d56e8b6a

Observation c6fb03d1-eca1-4379-84f1-a1909f7b60b0 · inbound

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models cites this paper.

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:10:48.948121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T03:09:09.713822Z digest=sha256:0684d471fde463096ef1ae52631b1ca9d1759af9ca59d1f6ef13cfac2856dcc4

Observation 312253bc-4d71-48ed-bf49-7c0cd97e49b6 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report RoboBrain 2.0 Technical Report

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.697672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:b213e10095a2ce914e1a52fcb0a6980a6d480a3e534724fe2d772d0b19fe836c

Observation 1310b23c-6acf-40e2-a448-6ee664d4bf35 · inbound

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation cites this paper.

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation RoboBrain 2.0 Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:12.127420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T20:16:16.133013Z digest=sha256:c83b033bb86f3f097c6cb75e4ef2185a19ee59618e7865679592b82079223b86

Observation d3cd3b4f-b244-4860-961c-2e6fc1eb8c58 · inbound

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents cites this paper.

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents RoboBrain 2.0 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:42:56.199932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:42:56.199932Z digest=sha256:9a65265ec0c61ce893af19cf69486c434a09f494a560bff42511a3a2539cff7c

Observation 59b7420e-52ac-42cf-b44c-0629262cf52b · inbound

Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes cites this paper.

Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes RoboBrain 2.0 Technical Report

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.084109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T22:06:34.356607Z digest=sha256:d7e37c52163e24e2249fa634e21cc7fb18bcfd0a83ce59bc054366a24693e5ba

Observation 1e942416-964a-4ecf-a20c-2f39435d7dca · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics RoboBrain 2.0 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.144120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.144120Z digest=sha256:32c024c6926f08a07674d5cdc9bb45a2ef2dbe31a34d99c78d36243af4279c1f

Observation 40d4743d-ee15-45b4-8475-052663c11788 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.865225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.865225Z digest=sha256:2c0d90c3d22648b3d5d195d43a6ace10b47c0a2b8341f5a6f2a8f1a28f7e59f5

Observation 3500d414-05ab-4a48-b91e-94189105986f · inbound

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations cites this paper.

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations RoboBrain 2.0 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T23:42:44.515158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:42:44.515158Z digest=sha256:a15dbc08cfec5884e43484a1d88bec7dd2b0f6fecd472a4dddd4bb09ade07a91

Observation 3c464154-1166-491c-bd68-a4bb32f19043 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints RoboBrain 2.0 Technical Report

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.391514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:91ec39c5083b5c7bcf8e52b6bb1359d05d8061830d502e640bb9322a3bd4a725

Observation 792e59bc-433d-4cc6-84e5-57b6a9705f94 · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning RoboBrain 2.0 Technical Report

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.143333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:ac639fca33a89ce67275f57ae13d754f0d757c31480c20ab3a4837ffcabb6761

Observation 116bf40b-0e40-460e-9ac7-9c301b178980 · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks RoboBrain 2.0 Technical Report

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:58.595352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:57c38983a4a86800a11e03b67faf706db64047f562748fdfdd56caaef9ed4c10

Observation 184893ee-00ef-4250-97fa-b4bcc3484863 · inbound

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly cites this paper.

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly RoboBrain 2.0 Technical Report

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:02.016734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:24:42.800233Z digest=sha256:72981259a9f2fe98cc452e280fbb942bafbfc3dad0473deb7e4ebfff8fd0c7ea

Observation 7c981598-7bfe-4828-8115-cb08b43cab03 · inbound

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly cites this paper.

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly RoboBrain 2.0 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T23:37:01.875768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:37:01.875768Z digest=sha256:a2dbafe191a78a24be592933e30e55340a9a4c0da5af2634e7d665d2cada1cbf

Observation c96ee93e-b8a2-4904-8d92-d2025b1fee69 · inbound

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models cites this paper.

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:28.939398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T04:27:18.284698Z digest=sha256:66029e47eaa65915bc9b0e22de9da40dfcd1afa8918fb168f2a8f96990088710

Observation 8865ec66-acc0-40a1-b058-ada5ceb91e39 · inbound

Long-Horizon Manipulation via Trace-Conditioned VLA Planning cites this paper.

Long-Horizon Manipulation via Trace-Conditioned VLA Planning RoboBrain 2.0 Technical Report

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:38.906730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T21:10:03.554504Z digest=sha256:f0da2e1056c626f3cdb58d30270d5ee8e574466fef3f627c33ffe8e77d64501f

Observation 562395d8-55e6-4094-aced-4014a526ab31 · inbound

Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance cites this paper.

Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance RoboBrain 2.0 Technical Report

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:06.874405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T14:51:10.807902Z digest=sha256:e58acdc3fd97adf14b63b17d5bbba36b22cc0450ee9ddfc34c0f19ee3e4fb7fd

Observation be5652f5-c3a7-465d-8802-17a00bfcf6de · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data RoboBrain 2.0 Technical Report

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:57:33.123620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:864e4392d0d9f7db0ccff7f9b4849c08010d26c012c8ee6ff9dd0d35cb26d11e

Observation cc82a7c8-12e2-42ec-a883-59a68d12ed19 · inbound

Rethinking VLM Representation for VLA Initialization cites this paper.

Rethinking VLM Representation for VLA Initialization RoboBrain 2.0 Technical Report

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.048577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T22:21:29.733181Z digest=sha256:454b011314f3d437bdfdcedc0eed0759a109817cb79d810add713cf5eb905ada

Observation def6574d-f7bc-4516-84fa-b5d4a983d7fc · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision RoboBrain 2.0 Technical Report

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:58.975304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:b5c8d08d6903d34189b075502ad44995eeff7440f8699c62ce1b6fa3430a6882

Observation 789c2ad1-2173-4197-8139-3b88c54c1d28 · inbound

Grounded 3D-Aware Spatial Vision-Language Modeling cites this paper.

Grounded 3D-Aware Spatial Vision-Language Modeling RoboBrain 2.0 Technical Report

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:15.373936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T08:08:36.012761Z digest=sha256:77d3490f1ff9e9c55e6ee133b1d36ef44a8aa8b666dade4dbe32824b4bc8b7c6

Observation 3f16f2a2-b1dc-499f-8ce2-581e53067d84 · inbound

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners cites this paper.

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners RoboBrain 2.0 Technical Report

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:23.011643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T14:19:40.492920Z digest=sha256:440db3333b27b80b0761d5bba13236b8af9e74cca466a611ea24ff3b03d09db9

Observation 6e5147c4-8e9e-4803-9239-d81c15592bd9 · inbound

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators cites this paper.

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators RoboBrain 2.0 Technical Report

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.199501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:25:30.998989Z digest=sha256:80ee1a176e84fb2573862494df8a2da25449503339fdf74d022314d8d13688cc

Observation 321c1fc5-bf7f-4f15-a02b-395bb2b7be90 · inbound

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data cites this paper.

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data RoboBrain 2.0 Technical Report

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.661631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T18:31:03.548002Z digest=sha256:1ae8678ccf225bf98ea4e020e5a36fe26ae809138e1dad9784f6474553d78d82

Observation 1c86762d-7507-4838-a49b-2c4e8bb7adef · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report RoboBrain 2.0 Technical Report

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.986423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:f5ebe69d9813d90c441995da43be3be7efd148b243f00c529b99616d4f8974be

Observation b280ef3b-d1f4-4507-b32d-f76ad2c8f552 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboBrain 2.0 Technical Report

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.253641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:30be33bb3a85eaed6a6b5149bdcdbf19e1b73fd7b106dd1ab18b29d36f14ada7

Observation accaf290-34ed-4811-a996-24dba529a17e · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models RoboBrain 2.0 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:c1f62d2be7d4d395bfd5283ca85bd3bc4e434d78e03c353f0277d5b1d30fb51b

Observation d14bc912-963f-4384-bef1-2007f85d81ef · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.762558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T06:26:32.209719Z digest=sha256:3cf8c39328e84b9a07e1cc8e1e8a223474b796bdeb0e2c816d8cd202a4c20323

Observation 8e09e1a5-fbbc-4c8b-96aa-c59eb4f48b7d · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:49.545830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:49.545830Z digest=sha256:774760b4551a865c6ef1c655d5e1374149b38d9fe3b378450d537ca5e1b2cbd7

Observation 67811327-e148-4b8a-a6c2-c92a09affd0d · inbound

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale cites this paper.

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale RoboBrain 2.0 Technical Report

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.372338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T06:28:14.694168Z digest=sha256:bfe798e5ee0b16e4e54176890f560874b8dfe026a064d6dbbe6af056a996306e

Observation 5afc2cf3-1a49-4f3e-b658-38abde83644e · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboBrain 2.0 Technical Report

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.342268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:25676023c95cbb1a0c5433b9b5504b7b6c13ef80bc10fc9ceaaf42957d082e0a

Observation 7be93e99-e912-46d8-a1e9-be24969da321 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought RoboBrain 2.0 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:27.100594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:27.100594Z digest=sha256:02865d0f6d3e0d58e251d59defc53fb7fa6023d3915cbbfb79ef6617c83985e7

Observation b0a6796d-6f6e-498f-896d-ebb3505f7740 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model RoboBrain 2.0 Technical Report

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.627244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:4751abf9bf84a3bb1fbf22f221bb1ebba58f431fe4ea027b5425964c560d626c

Observation c588d71b-7b20-48a4-92ad-439f99dc3459 · inbound

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models cites this paper.

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:29:51.478385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T07:55:28.884473Z digest=sha256:51b03faa34e137e29dfd092586136ac363ad45962b963fd6fd45ae0809865022

Observation 24cb4fcb-10bf-48fd-8a4a-65e2b9372917 · inbound

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models cites this paper.

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:33:54.753630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T04:41:28.564104Z digest=sha256:9b3234882c73eb37ae96d7e1fd184f0893341306d6f52e3bb96dc45c5a78c582

Observation 4c17d25a-fd0e-4a60-90f7-ddad61284316 · inbound

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy cites this paper.

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy RoboBrain 2.0 Technical Report

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.044905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T04:52:51.524022Z digest=sha256:c3232e48ccc46635d465b5f95a51c13374156b140fd2a5a92d159277cc6eb9ee

Observation 06fed410-4c3c-41b4-9306-01dfa57cda84 · inbound

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation cites this paper.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation RoboBrain 2.0 Technical Report

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:49.418964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T05:05:56.065261Z digest=sha256:5e6c6b3085ee206751be1f7954b53139223a8c73940aee91b3ad97b2d5e12ac5

Observation 897b8aa1-69fd-47f7-8154-e21fb5a12df2 · inbound

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping cites this paper.

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping RoboBrain 2.0 Technical Report

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:17:02.454959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T14:16:39.649823Z digest=sha256:6ec4b00dc23faead6f317973fa0008964c1c2545fdd06f7f5734ddae53d0a9e9

Observation eddaea49-160f-4f46-b2b9-b8d3487ef1b8 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI RoboBrain 2.0 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:f24bc68facc167ae37e76520d81c333e0b3bf56296b338579b4be15245f0ccb8

Observation c4030680-8200-4dd6-a557-2cc703beefc7 · inbound

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition cites this paper.

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition RoboBrain 2.0 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T12:04:50.479862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T12:00:14.336135Z digest=sha256:46d80c88debb2b59ae5d97e684d18697a71780388a17e1a7c2287a221a8544d8

Observation 08bcabca-1051-4af6-b540-b0d21096afe0 · inbound

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition cites this paper.

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition RoboBrain 2.0 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:36.948882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:36.948882Z digest=sha256:f9f62c0dfe1d80e81c3600756b83113aafad40059a24655e101fc73d901bd394

Observation bd05024f-bb2d-4fd8-a1c9-7448c1edee98 · inbound

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation cites this paper.

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation RoboBrain 2.0 Technical Report

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.771049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T02:41:24.989404Z digest=sha256:0737e33cf71d789ff5f59ba080d403c340d56bebf027fd45ccb4f3dbcf2a9f04

Observation 98b34350-2ead-4960-abfc-0934a5747fc6 · inbound

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation cites this paper.

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation RoboBrain 2.0 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:02.584798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:02.584798Z digest=sha256:c446ed361a3cf39902a34bd35fd198cc585859d9f082cdcec245448434bf1e4e

Observation 3b298bf2-7407-4c2f-bb5f-bab3dc6a40ab · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RoboBrain 2.0 Technical Report

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:5b2a5165909cb0884ed276b0daae84b966245dfa191469a4bfdafd20402c7017

Observation 275781c7-e1be-445d-8f89-fd43418dc941 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory RoboBrain 2.0 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:31.686911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:31.686911Z digest=sha256:14f4bd89ef6e4d28abb421806588a8e2fca23db8018035ea79c31d00467d6a2f

Observation 8e10392d-9963-4043-9964-8d0bf5d46972 · inbound

UniETP: Unifying Environments for Generalizable Embodied Task Planning cites this paper.

UniETP: Unifying Environments for Generalizable Embodied Task Planning RoboBrain 2.0 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:20:28.787487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:20:28.787487Z digest=sha256:9b23e69248e804872e8e2d40f9a59f59b613bcd12cb075a980491bf54c140f0d

Observation 0695f7e4-7f7f-49ba-9bf2-497e60fab566 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:41.142227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:41.142227Z digest=sha256:45b69142e237ff1151b32852c5e26789d9854fab1a4648e113262415540888a6

Observation 57720f11-e530-450d-82e3-d0b31f305f76 · inbound

Towards General Language-Conditioned Latent Safety Filters cites this paper.

Towards General Language-Conditioned Latent Safety Filters RoboBrain 2.0 Technical Report

Reference 246

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:36.759144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:36.759144Z digest=sha256:dae3b3d4f7238b2fd9ac17c529abb950f6484c3d9a00dde21037366e5a417629

Observation 29606749-d338-4124-8d6e-841411590730 · inbound

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance cites this paper.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RoboBrain 2.0 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.781637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.781637Z digest=sha256:d1b87c949687b30c0ddc58166e470ba175701f8b99dc2ac28e46328e3b07d4f6

Observation a1f67bc6-4803-4003-9aae-34437e4bcdcd · inbound

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation cites this paper.

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation RoboBrain 2.0 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T19:38:41.751687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:38:41.751687Z digest=sha256:255741b7354a61829d2deef8178dfa76278818c78cef507a981c40cf323ec17e

Observation d0722474-d9b0-4ecd-8677-e418b507e6aa · inbound

Compiling and Benchmarking Task-State Horizons for Embodied Agents cites this paper.

Compiling and Benchmarking Task-State Horizons for Embodied Agents RoboBrain 2.0 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:36:21.485253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:36:21.485253Z digest=sha256:7a1b20189fc17df3ef4736632b3254484f505f9b7305186a5288f7c0cf9b97d6