Pith. sign in

Paper Citation Record · LEDGER

Hume: Introducing System-2 Thinking in Visual-Language-Action Model

As of 18 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 36 inbound Pith citation observations for arXiv:2505.21432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21432 v4

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:55.651566Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:57.174397Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 81007fef-4341-4caf-a768-4a4bb9bd60e8 · outbound

This paper cites Vision-language foundation models as effective robot imitators.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Vision-language foundation models as effective robot imitators

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.176151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.176151Z digest=sha256:7fb966ab1bc0827a5774ff1754cf7277939d827363cf9809d3c16e2c26e58ef3

Observation a679ba6e-d086-424b-9449-b56e638d267f · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.316972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.316972Z digest=sha256:db821433a525edb8e3ed66a016dca356d037427fb719217cb8a26c621aa056a0

Observation fbb5c2de-20a4-4c51-9fb5-d85fe417456e · outbound

This paper cites Fastumi: A scalable and hardware-independent universal manipulation interface with dataset.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Fastumi: A scalable and hardware-independent universal manipulation interface with dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.649706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:46.445264Z digest=sha256:a09bd35ba90a90766ece65d3de912a472e07fcbf261b2f7c40d89ef7c6dece5d

Observation f7a06d05-3a29-4c1f-aa89-74229fc47b72 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.578291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.578291Z digest=sha256:c1dd59b5630f6cc42ee178c739eccafb969c8c25ab6b37408bf572f1d3154454

Observation 72bb9ad0-c3ad-4caf-92d2-0ae9f341887a · outbound

This paper cites Learning 2d invariant affordance knowledge for 3d affordance grounding.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Learning 2d invariant affordance knowledge for 3d affordance grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.457287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:46.673699Z digest=sha256:3a33331d109889beca90ac44bb562b2b3fa58041dded8caf862fd1ddf5a43702

Observation 523ae9b6-ae5a-4258-84ba-9c07faaa7f70 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.813552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.813552Z digest=sha256:14821a34007583e24e0fff526a74be2bf23096b77af45bf5a1e72bee80054687

Observation b16ca1ec-1d7f-44e0-94fd-da39d0af58cc · outbound

This paper cites Improving domain generalization in self-supervised monocular depth estimation via stabilized adversarial training.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Improving domain generalization in self-supervised monocular depth estimation via stabilized adversarial training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.302774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:46.961211Z digest=sha256:6c268ba40cda774be777e8798d66a4ab1fa6e7bfb8c3e00f5a89640e0c2254e8

Observation 4a1439db-2ae7-4131-a03b-830828a63ccd · outbound

This paper cites MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.040696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.040696Z digest=sha256:0dd3af6a4f21a130cf98bec18c551d7ecbb3542eb2d76ed7eed2fd8591f9f2c4

Observation e4c00cb3-77d4-4214-b80e-5fa8af747019 · outbound

This paper cites ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy A Star.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy A Star

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.143573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.143573Z digest=sha256:de25cbf7ed445c438dc0f7c8a701bef2a996d96cd407fc0d274959ad33d7ff99

Observation 5f130c18-80c5-4cd6-9397-eb6338f72955 · outbound

This paper cites Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.253710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.253710Z digest=sha256:2d3a8e9234778783aaea301310c36c0ea1bd85891b3a0ac043281014154a7a7d

Observation 94eaf881-2020-449e-ab3a-2b66d3dc8480 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Gemini Robotics: Bringing AI into the Physical World

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.358067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.358067Z digest=sha256:fcd0af5d62c97acbfd6fb7be8cd018aa302cded839e81b02fc5696ed0844e01e

Observation 7b973670-3451-42d6-ada0-178115d937c5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.463400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.463400Z digest=sha256:dc85e379d128feba32ce0ef4abacfa330d8bfc848adbf035719bab83caffd765

Observation c5d2e319-e561-438f-b073-5c14a9713627 · outbound

This paper cites Thinking, fast and slow.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Thinking, fast and slow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.579910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.579910Z digest=sha256:9826e7d0db936124bdf85075001baad5487274f7bdf455a4816683e73f1b31d8

Observation 7005a5ee-e8a7-4078-b940-90adc3152e70 · outbound

This paper cites COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.705703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.705703Z digest=sha256:5fad326d581470164b1be6784cfed4b0788fe7cf9ec83d21e9f65253b5db5dca

Observation dad5dba8-74ae-4f28-932a-592062ea71e2 · outbound

This paper cites Kinematic- aware prompting for generalizable articulated object manipulation with llms.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Kinematic- aware prompting for generalizable articulated object manipulation with llms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.141533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:47.802501Z digest=sha256:89173ea7884fa12be9cceec02fb8e19355c496afc11ec97d710805a23d6d664f

Observation 8c35e3b9-4873-4117-a821-d10f28fa2b16 · outbound

This paper cites MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.950909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.950909Z digest=sha256:079143b94f32dfda781a82bafe1549f31442b0c34af055d604346f898958eff9

Observation 6dddf33d-15cc-4694-939f-d600362909b7 · outbound

This paper cites Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.084439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.084439Z digest=sha256:2d74fc47dbd7abf9cda8c2fed4393a9bcc0e09a66bbcc0183e145a16f3911d7c

Observation 597e2a01-7034-4650-b63b-32cafef1ea61 · outbound

This paper cites Robotic policy learning via human-assisted action preference optimization.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Robotic policy learning via human-assisted action preference optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.230494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.230494Z digest=sha256:a1fa240404e831eb311385abc69db00b45aa888d085b650e732d0467473f6a07

Observation 92e6060e-9b8b-44a7-a2c4-33d3bfb0e455 · outbound

This paper cites Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.991650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:48.361825Z digest=sha256:afe38eaa8e638e6bed6f8e534e940a6de78ca655e888481272a090567e8a295b

Observation 85cd5166-56d9-467d-ae2d-9aa866a98f5e · outbound

This paper cites Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.487515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.487515Z digest=sha256:83e260a98f06270dece98b7df0c2d1dd9b53d012ae9afcd7b04f148674b2b99d

Observation 1aa81f56-cba4-40c7-b47d-aabcf0641517 · outbound

This paper cites Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.854828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:48.605623Z digest=sha256:bd9b9211543033c5608a7d00e0f9ab3e4153a5d0c5acc03d7a17d733f0b66843

Observation 1af8d243-7c3a-4325-b55d-d0415085ca79 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.688344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:48.729701Z digest=sha256:8aadc58f424e9e5cc61b3fae564c462573c731555804328e76e372961f61de39

Observation 33ca2d67-8cba-48d1-8c09-e441a74ddea4 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.892757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.892757Z digest=sha256:3ea76ef3a95a3e3f9e2239ecd03a85d19cf846f1685f13bf3495ea23d8a747a8

Observation afbe1c2c-f490-4a80-a099-310c6b6a865f · outbound

This paper cites A dual process vla: Efficient robotic manipulation leveraging vlm.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model A dual process vla: Efficient robotic manipulation leveraging vlm

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.459634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:49.081807Z digest=sha256:12ec28c061862f7a9cac58939cc6f644c424f910923453856baa018f0adfe7b6

Observation 0593f14e-de49-4900-97f9-ced694cb1150 · outbound

This paper cites HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.257537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.257537Z digest=sha256:61db60fd164b6674a1372a9a32786f65bd3e84b0606830717046d81ebb93d83b

Observation fbc7f4d3-eef8-4e1c-9eb2-c8f2cbfe781d · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.385522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.385522Z digest=sha256:906e797aacd64d19c6735667c4d34c4508965141e80e5ad65a184d7993355779

Observation 5ca8f132-157a-4cc0-a0b7-b44fbf22ab28 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.511193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.511193Z digest=sha256:e58ab4d020840b43a94c820ead0494b062f8d8fe61f4fb085cf5891bc8308635

Observation 3c1db6a1-da02-49f4-a743-2c29061bc1a0 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.614480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.614480Z digest=sha256:7b31648f228e6f7084c72a56f011f2cecc2878422793fe24d1f7257942e86157

Observation faf03009-fb2c-461a-bd4d-14ef462cf812 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.760140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.760140Z digest=sha256:b3f6eda2058edabcd40a987b9fd87b454f3c080fc20d78b6428fc6072f676bdb

Observation f3b374f1-6e31-40ef-983a-45e2b9a5fa7a · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control, 2025.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Helix: A vision-language-action model for generalist humanoid control, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.206921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:49.915193Z digest=sha256:90b592fd62b1a16d5378975a3cf7d317436ef148b750dcc4b02f57147139e877

Observation 32fb4798-d4e5-4b97-82c7-86ae6160cf74 · outbound

This paper cites Flow Matching for Generative Modeling.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.042676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.042676Z digest=sha256:8a4c00382151d011336a7cb880104fd1195a3ece23c3dd401f17c67600391399

Observation e36e8996-aef0-42d3-9da8-5e096f4a255c · outbound

This paper cites Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.178954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.178954Z digest=sha256:9db6dd11023dd6ce84a93966e6b6e3470091f4c1d68243d26f9b2ad378c749e0

Observation 4e2d2e85-f720-40dd-9f4e-71daf55ff675 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.325026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.325026Z digest=sha256:d060b91e4d7e3332dde5b214e59d16d945f599a255aebc96a00e2b6efac05162

Observation ed099548-4729-4806-a43b-4047d2b92ee5 · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Evaluating real-world robot manipulation policies in simulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.079764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:50.473479Z digest=sha256:3bbcafdc661ec6e0ed88a3c2808974b5978de3fd6807048fced44f9ac5b8cceb

Observation b6980a7b-23da-4fac-9da9-898329f2d91a · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.601029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.601029Z digest=sha256:9164dfacdc90797c95765c2b6149611bd9a3087142637d30b39be35c309383e8

Observation 4661f28f-b4fd-428c-8cab-dd0ea2467d6c · outbound

This paper cites FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.714630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.714630Z digest=sha256:32728d105d433137543bd92a6c6c227841ee3ac5330a165f38a5e161961c2e51

Observation 257b1298-91e4-4d6a-b99e-c8e12cacef87 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.849364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.849364Z digest=sha256:c5e13d3526898cd852e8f4f9708fd2ef8ce0bdbbb0f68edaf292ded29778d43b

Observation 15194271-763a-4a9e-ae20-58bc0171772c · outbound

This paper cites RoboMP2: A robotic multimodal perception-planning framework with multimodal large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RoboMP2: A robotic multimodal perception-planning framework with multimodal large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.951725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:50.961129Z digest=sha256:080963cac2c4725b56b0c1d3601d5f6b0466c9d918c5bc879af5f128fa9a1e1d

Observation 6f71dcb0-3c70-43a6-b824-801f1038533a · outbound

This paper cites Universal Actions for Enhanced Embodied Foundation Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Universal Actions for Enhanced Embodied Foundation Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.178608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.178608Z digest=sha256:7bc5857d1a9f432777b21673bc459033e984fc3d9384318fa66072a0212c414c

Observation 98106cf3-bab8-479b-bef4-aedf443dc267 · outbound

This paper cites Learning causality-inspired representation consistency for video anomaly detection.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Learning causality-inspired representation consistency for video anomaly detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.715736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:51.301356Z digest=sha256:70ac245b9ac45dab3f6dbe49585a91d3f7bd9d8b8f41c7b1ff7f311b71f798f2

Observation 7987215a-b994-45c6-abc3-f9d0c0442d99 · outbound

This paper cites Pali-x: On scaling up a multilingual vision and language model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Pali-x: On scaling up a multilingual vision and language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.525438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:51.416783Z digest=sha256:63e2beac18ce0dd69d239380a762d072840b33eefbe4af5395a1dfa3a64ff891

Observation a686961f-ed9d-4b19-866a-bda44f39afa6 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.312277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:51.548815Z digest=sha256:0443c5c8f2d24a9bd7a7cb064b9220fb30331b1d7003921e605c0fe655ae5e8d

Observation f30ddaf8-0366-41b9-ac3c-a4e833fb91c2 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Open x-embodiment: Robotic learning datasets and rt-x models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.120220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:51.664459Z digest=sha256:06ce7e547d49ddf7d12b50f3b6c06092b91b04302b3b4ff27ba872944563b09b

Observation bb30c695-ea59-47f2-b80d-691d8ebc69ca · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Reflexion: Language agents with verbal reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.772547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.772547Z digest=sha256:4a1392cd7e54fc3e2ae0aae1b54159502fae5e0dd2e40aec07b126cc16e30049

Observation d0c08af9-ddd5-4a01-aed7-1077f92b290c · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Reasoning with Language Model is Planning with World Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.884717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.884717Z digest=sha256:e3c5b6979dc5c6ab6c52804c052e4e028f9c54194816e8226bd13bb48ddaa25e

Observation 76800bd8-a57b-47d3-be37-85e5bc17e353 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.985432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.985432Z digest=sha256:7fba6292fcccd48d793a3723b50e618c79a923924130dac200830ebf3595d229

Observation 232a0add-3c70-40f2-85a8-dc5ba1c1ae8a · outbound

This paper cites Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.124767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.124767Z digest=sha256:087ce2d8b954d1de25a9eaf9ee576e38897ece0f67bdbb45cdb1987f89a30348

Observation c515e4aa-907b-4748-b0fe-ab70894abfb7 · outbound

This paper cites Mllmguard: A multi- dimensional safety evaluation suite for multimodal large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Mllmguard: A multi- dimensional safety evaluation suite for multimodal large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.909259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:52.282259Z digest=sha256:e45e7562a3db81f312e168376b4b39702cb04a1bd90e1da896e335f6ab1d5202

Observation 51c50869-8ef4-42d7-9746-4530cff6eec3 · outbound

This paper cites SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.387699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.387699Z digest=sha256:b119bbbfea1c45624d1350381ad087ca928a885473fe554ac3655414e319a882

Observation b14d58a8-8cfa-42de-b0fc-c8e86a6fd48f · outbound

This paper cites MorphMark: Flexible Adaptive Watermarking for Large Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MorphMark: Flexible Adaptive Watermarking for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.513946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.513946Z digest=sha256:3711c12067e2a0eb3fece27474491c49eb7328aee851d205270eb232e821fa20

Observation 0ade5a7e-5cda-44de-b5bf-de1ce2217f55 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Tree of thoughts: Deliberate problem solving with large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.729436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:52.622187Z digest=sha256:f3f0254aa25bbfd2c71b5e1e76b09c1202bfe2cae0044adfb98dd46f0e6267cf

Observation 52dd79ae-1572-46ca-b6ee-a04dcf031ad7 · outbound

This paper cites AlignBot: Aligning VLM-powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model AlignBot: Aligning VLM-powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.754745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.754745Z digest=sha256:c0e65ff154273a1ec0197b31dd90ce438b8827ce64994a2d97b85d88f66bd0d1

Observation 56722f56-4188-4b25-aede-774f1ca05879 · outbound

This paper cites Sets: Leveraging self-verification and self-correction for improved test-time scaling.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Sets: Leveraging self-verification and self-correction for improved test-time scaling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.896307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.896307Z digest=sha256:45b5823fc9e9c43c78b7ed6345dc38d720d45087fc9a81559b8ec7c94663d1a1

Observation 23c7015d-acb4-4e75-9d04-2bd052aeccba · outbound

This paper cites Interpretable Contrastive Monte Carlo Tree Search Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Interpretable Contrastive Monte Carlo Tree Search Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.995028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.995028Z digest=sha256:59eab4f9764a8773e613422a88ca5a923284b53e9f23c37abb612c6352e6cf48

Observation 51468219-21ee-4e85-a594-c758e0825966 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.140367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.140367Z digest=sha256:56cb460ad20f61b4bd7476efff96b7b9bd8d5bf78146113972e99310424460db

Observation b0acd536-3231-4720-905d-cf6adbc16476 · outbound

This paper cites Cascaded Diffusion Models for High Fidelity Image Generation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cascaded Diffusion Models for High Fidelity Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.244353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.244353Z digest=sha256:d65efe3b7674f5bb6164ab45508d3b1f8bc78ee7d4dd026a48f7cc526e8ffdbf

Observation 55a1e608-4ed0-4fa6-bde0-4ddcb2a9a26a · outbound

This paper cites Revis- iting multi-agent world modeling from a diffusion-inspired perspective.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Revis- iting multi-agent world modeling from a diffusion-inspired perspective

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.327795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.327795Z digest=sha256:59028274d9f2a442d3ad8fd6b012f75f11f8810ba99a6b0832ef267ed00a0eea

Observation f73061fa-010c-4826-9dc2-865bbd20f8fd · outbound

This paper cites f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.228091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:53.420181Z digest=sha256:dca9c8f3ce4d74b666f3d5fd368bb363b083976c39f12b4e2924233aaef2672f

Observation d2f7f3f1-10d6-4d08-ab54-bf8ea252f4b1 · outbound

This paper cites Bring Metric Functions into Diffusion Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Bring Metric Functions into Diffusion Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.005341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:53.554629Z digest=sha256:ccf96ef2d4c1034155667781dc773b018688b2e0e9c779dbc3b8a277879cb4de

Observation 50763b71-7e9d-4562-b70a-c7dee4706c6c · outbound

This paper cites Spectral-cascaded diffusion model for remote sensing image spectral super-resolution.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Spectral-cascaded diffusion model for remote sensing image spectral super-resolution

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.526003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:53.717819Z digest=sha256:7b1bf1fee4bbb85a266daacc902cb624038367cb1fc1955dab831795c847342c

Observation f51f89c3-d1cb-4027-84ce-d0d70f442a8f · outbound

This paper cites High-resolution frame interpolation with patch-based cascaded diffusion.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model High-resolution frame interpolation with patch-based cascaded diffusion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.305583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:53.835285Z digest=sha256:1838a3971363c04f822fbc48dc1253fd3e241ac0672b94663454bf641e0bc05b

Observation c527b747-cc31-4b73-8b80-f6f1e271296e · outbound

This paper cites Cascaded diffusion models for virtual try-on: Improving control and resolution.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cascaded diffusion models for virtual try-on: Improving control and resolution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.127619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:53.990234Z digest=sha256:11f90f4229234aaa1852dfe75998bef0f9078b62b41011cce48b832069ed64bb

Observation 2c28b4f4-4442-4192-947e-bf9bac79cfd2 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.126278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.126278Z digest=sha256:fac5a47558ffa796952df2eaf4d0152886c64636f592b69c70c7886cd0f39575

Observation 05da854f-a948-4814-ad54-9ffdfb1cc9e1 · outbound

This paper cites Pre-training for robots: Offline rl enables learning new tasks from a handful of trials.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Pre-training for robots: Offline rl enables learning new tasks from a handful of trials

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.971727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:54.254443Z digest=sha256:5cc947f6ec217cb9d267b9bc33d9793c410940abf85bfb07c4c5b2bad4dd6b98

Observation 287d1e5a-a6e8-4633-a58d-c5f73ef9088f · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.814329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:54.384981Z digest=sha256:12b27f62a3c0a23ca37741aafebf0c52c5b71d13f31008bbf74d057e69968af8

Observation f5826893-a8b4-4a18-848d-bfba099c12f1 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.509544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.509544Z digest=sha256:631cbc4e9e25eee3eca3c6df2e6a87a2a7e8ef45c619c5fe1211dd2ca9b8ecc0

Observation b73c3ab7-a074-4ed8-bac6-3ee1f980e446 · outbound

This paper cites Octo: An open-source generalist robot policy.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Octo: An open-source generalist robot policy

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.674777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.674777Z digest=sha256:dcac4a063e742b817d66929eec27cb635100174cd4a77c864abce5163d32ec3a

Observation 949e21aa-1443-418a-a674-d4b10016708d · outbound

This paper cites Scaling proprioceptive-visual learning with het- erogeneous pre-trained transformers.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Scaling proprioceptive-visual learning with het- erogeneous pre-trained transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.645007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:54.796683Z digest=sha256:b4f296d0a9c75de24577a0586b365629f3561242eab6ece106e404d0cd5b3c93

Observation fb8bd160-8b34-48c3-9899-1b76a00db3e0 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.889344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.889344Z digest=sha256:5c17bc10095786abbe2cc59d0ce1f184bf3ab3ce52cd05c12666aedbab4f8866

Observation 4ee513b6-aba5-475c-9d35-14fb20a3afc8 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.004692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.004692Z digest=sha256:f7994b5708e9f432658b20d80e0c5042ef9b9b212c9571cb0eab86e47c7ebf44

Observation ad6b665d-616d-44f5-bdca-262e55c9930f · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Diffusion policy: Visuomotor policy learning via action diffusion

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.157444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.157444Z digest=sha256:56ac2946903148661ce855dedf07e7bef53860e9a217ad05650c4f7bd4d019d6

Observation 7ac1ebce-6d32-4038-8412-5bb69cd72bc6 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.223798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.223798Z digest=sha256:a12d71c5d44e7dd399beda1e290f4e513b80344be6f39eb88cdab2cfc597d3cb

Observation ae1bf471-dc67-454c-9dff-58f2d586f78c · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.329461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.329461Z digest=sha256:bfe580b62eb6796bd0302c5ea2dde20d5c46d7b593a92e96993894240f495fe6

Observation 1791c3be-837b-4f62-b6bf-1553d4807b86 · outbound

This paper cites Continuous control with deep reinforcement learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Continuous control with deep reinforcement learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.428373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.428373Z digest=sha256:fcd82593d5f32b8f38883c11ff2e98e0c6779242a7380c2244d6611b52a912d0

Observation 9b18831f-3bec-46ab-82ca-e17095e176ce · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Addressing function approximation error in actor-critic methods

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.417704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:33:55.564589Z digest=sha256:1a6dedc40e941076f75c40a83acc27b8ac39c4bea36802a9893757c8cc9580a9

Observation 5d8e87df-c426-4323-8554-e9c8d0a2fea1 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Soft Actor-Critic Algorithms and Applications

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.651566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.651566Z digest=sha256:04a0ab2e860b75a5c2d13f28804314ea848c96567c99dfc868a3034b04fd2576

Pith citing papers

Observation 7f957329-1602-4060-b1fb-d8301b5ef22a · inbound

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving cites this paper.

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:19:42.874884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T19:19:42.748573Z digest=sha256:289644aad256c9a8794c645075695154a8a691f09f5357cc11293144900bac12

Observation 9943f6b5-3f45-44ca-a277-c0385c11385e · inbound

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface cites this paper.

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.315435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.315435Z digest=sha256:0285b6a31a5c9f687f1d613184ffd8a276453485b38e2652e88402b5dd16a95c

Observation bf1e6f74-a5fa-4853-a89d-9d6deb14871f · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.581802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:0dff30d988b4734f137e49eb6df2f5a8019c9d2ff3cefb68d2ac868cc04e1cac

Observation e6c42ed0-3a1e-41ec-be9f-2ac5e9d0b556 · inbound

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions cites this paper.

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:42:47.069003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T12:42:46.978148Z digest=sha256:6c10a3f6f7a5cda689963cd903f5f97d36d7b03dbd596116d2843d2c364065cb

Observation 3edc250c-f891-48c5-987d-ad183462ea08 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:39.902748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:6c77baf4f5be300ccafb4942c7f9beaf2277aa21c1442c9708220f0c42281942

Observation c5a28893-c809-42c3-bf40-11c38ea3f232 · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:11.780961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:11.780961Z digest=sha256:ac1b3366b4bca92f8c91d0d7aacc2e5e7a724d4e5d0bdc4dfd9c7311a44dbd3a

Observation fa07afbc-06fa-432a-b218-afd5a0eb7310 · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:00:26.042642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:9e67e90de8a788ff8cdf3d1e5b459bcdb273f8b23961d4ea7302016c72a65139

Observation 150d728d-7652-4a3c-bc24-46f9e1ba84ac · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.030889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.030889Z digest=sha256:c656c1897bf866988fa489f2474dffb8f346c6d317a41b142350bf9d2d854bbb

Observation 73a9281f-48a8-4fa0-a5d9-b79e4083dc45 · inbound

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data cites this paper.

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T12:10:54.741769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:10:54.741769Z digest=sha256:b33c4486a4611219a8f081aa903128b35383d202954757951612dd24697fa5b1

Observation e773579a-6a7d-4e0d-8a2f-5b3e8483c3fb · inbound

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models cites this paper.

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:50:37.167968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T12:50:02.670780Z digest=sha256:c74b78cb1aab804407f142c53806f31de70e84016ef5136d52f342927918ae35

Observation 8deee7bd-f6c2-4f84-b333-13298a69cef4 · inbound

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving cites this paper.

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.527563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T11:56:40.836234Z digest=sha256:9b11aa7d251de164226545464bf093ef57182efc88360f2ceb81db7484707cd9

Observation 635e0017-9b9f-4d17-b2f7-0e961982ed91 · inbound

Spatial navigation in preclinical Alzheimer's disease: A review cites this paper.

Spatial navigation in preclinical Alzheimer's disease: A review Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-13T19:53:53.219607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:53:53.219607Z digest=sha256:4b7b95f2d3b8dbfcd5acd2de89b6239f7de3a12540243710997127d2b5e52f87

Observation 5241d33f-930f-4b74-8528-f7da22a0e8d4 · inbound

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models cites this paper.

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:43:19.018696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T21:40:30.771335Z digest=sha256:59456080ab845089d5c2d0f98d6a0738af2b58af6189fdcd6b83bcde120b8fe3

Observation 90f67b5f-aacf-4d18-a407-f2dfa1a89af7 · inbound

Deep Image Clustering Based on Curriculum Learning and Density Information cites this paper.

Deep Image Clustering Based on Curriculum Learning and Density Information Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:03:28.527131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T00:03:05.442554Z digest=sha256:1f67fce3e5693a81b9f16567f4c1aebdf5df1e47eef71295f1e81e2d9f8a1608

Observation 28a5c88f-922d-4a9e-afe3-e79a16620ddd · inbound

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models cites this paper.

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.503826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T16:57:52.595165Z digest=sha256:5182c41954421604fa43a1aef25f7d9322b43d4482eb5bb2ebe2415c8355432f

Observation bce58944-e575-4e19-aeeb-2f96c4ab1f0d · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:04.809662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T15:14:03.865035Z digest=sha256:71fb953b6e1e191c9a688258748e093cdaab36d4a5d65daf8f7aee80f0e509ab

Observation c66e8685-deec-4791-a0d1-cde96d09ed9f · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.064730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T01:10:48.111387Z digest=sha256:e9fbf23139aaf23a142d322c2636df0d82c9741c283320b0f61b8ebed213caab

Observation c8b0042f-638e-4de8-b0bf-1858489ed884 · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.358813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:8268163e15314281456e15537d57e23e74e0aebe04db2d28a719668cf73109f7

Observation 9237db13-56b0-4f9c-8a7b-61c5ff07d408 · inbound

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control cites this paper.

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:30.361693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T03:47:16.205357Z digest=sha256:65920a7d0fd847830807039bd0df8c52e0c365dabb2e16ff63133300fa819e3c

Observation 2913645b-6fa2-4e10-8c4d-453eb1512fbb · inbound

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control cites this paper.

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.759639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T00:18:14.307099Z digest=sha256:87f65466d07a0961c97e37865d81c213e8278e3b4750d92c578466fd15b68043

Observation 30972e20-377c-40dd-b05c-ff820df2d8c4 · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:12:18.267361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:b38dd632d2143565dd7751b47b5370949f3d7aa81462559e500bd8006560308e

Observation 0c5495cd-9e10-45be-8883-67a3eef00e56 · inbound

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA cites this paper.

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.737640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T19:57:25.010339Z digest=sha256:08bf57c982f5768787ebbb42e072149d7fd380fae9a030ecc9e42e4caa23722f

Observation de9b08b3-5e48-45e1-9deb-259b4f49c5da · inbound

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA cites this paper.

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:13:46.004799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:13:46.004799Z digest=sha256:fe26e8a79bbe6ad7ee69589bbfecc3d11b6d85c6baba884aa425a63f3d4ce5a7

Observation 5e3fd18b-156e-424d-98fa-ba4770f928df · inbound

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space cites this paper.

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:58.507228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T00:57:58.312567Z digest=sha256:b9e6cde354c265f072d254adeaf4212647c6f8ab8b0c2e3d8feb6c5003347e53

Observation 53e4a2ea-4c7a-4e32-b782-a1ccb57258a1 · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.647699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:93352ceb44b8eba3223db43bcd0b74dbcce03a137dd75eaaf9e618affe23cbd5

Observation 8cf1726e-96e8-4ac7-8326-8919eb2502a4 · inbound

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation cites this paper.

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-27T04:40:32.814702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-25T19:29:00.117285Z digest=sha256:faac240b09ac6378a3b909fa2bcc62ecd1658d9260058f30a4cbf37ff3581b1f

Observation a96ac403-2fb8-43bc-bf7f-951da40ff893 · inbound

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation cites this paper.

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:51.756964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T04:48:28.868562Z digest=sha256:c5356e5e648090bb54136f59c45acd8d8bbb7ec27c200bf336b798ebf37d6be7

Observation 1c1c081c-dc97-461f-81d4-34b61c42e28c · inbound

Recursive Self-Evolving Agents via Held-Out Selection cites this paper.

Recursive Self-Evolving Agents via Held-Out Selection Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:14:37.572040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:13:15.923479Z digest=sha256:ec93d75afd509ffcea13d91a802d47caad89585a5a3348508d046ce697748ccc

Observation 6d0ce5b9-3a94-40f1-b56c-203336e675e1 · inbound

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning cites this paper.

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.827744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T06:56:45.197407Z digest=sha256:58027722d42e6824c89797d769bb10bf3c112d86bb9f8def607a442e5f48808d

Observation 984e3d47-f327-4463-81e0-474103ed5288 · inbound

ROSA: A Robotics Foundation Model Serving System for Robot Factories cites this paper.

ROSA: A Robotics Foundation Model Serving System for Robot Factories Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:16:52.678721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T11:11:39.584755Z digest=sha256:3ddf2773c68fa89fb1caf8932ebbe5b2cb8694b626b40dd3d430d79bf1062f13

Observation 67504498-1254-428f-91dc-015b7465a512 · inbound

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models cites this paper.

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T00:12:42.173815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:12:42.173815Z digest=sha256:ece88934a9e2b57e71890fa35a5374c2454f5a0886e2b26d2c5373e9c8e141c0

Observation 73f704b0-a5e4-469e-b3e9-0953aa79b3a5 · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T12:10:21.115628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:10:21.115628Z digest=sha256:00b40f10e34835d15ae63908694341685dd2233b467b5efaa39cc8ed4183862b

Observation 5efd68fd-27c8-462e-ab83-5827f5d3a342 · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.336447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.336447Z digest=sha256:00ca2b02ba25fc0cd5249dd62868cb8f261cd508a7ca224ff0320486fc2994ad

Observation 9c83eae8-90ec-454b-8d38-30cba65811db · inbound

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking cites this paper.

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:30:57.030649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:30:57.030649Z digest=sha256:094c7991c514e8b5b4a29dbd009f02fdc62307a1337c80cea4e0cfca709df868

Observation d8f48e30-f03c-4298-aa2d-b3bd694feb13 · inbound

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation cites this paper.

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T19:56:57.191876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:56:57.191876Z digest=sha256:d1417668db711e6c428dead41467690a7379d57620625c1b63ca93338c05c712

Observation c3761f79-4d71-4820-b142-ab1456f4d407 · inbound

Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection cites this paper.

Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:57.174397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:57.174397Z digest=sha256:eda7356e83bd74a09c4aa8117d23d41924d2d994e42013a7ce780aa3b7ea6f2a